# site-reliability-engineering

A software engineering discipline that treats the operation of production software systems as a software problem.

This is one page of public article previews, not the complete archive. Follow Next page to continue. Summaries are not the original full articles.

## Honoring #IconsOfQuality: Mark Hrynczak

DevFeed: [Honoring #IconsOfQuality: Mark Hrynczak](<https://devfeed.tech/articles/honoring-iconsofquality-mark-hrynczak-27000.md>)

Original publisher: [Read original article](<https://www.browserstack.com/blog/honoring-icons-of-quality-mark-hrynczak/>)

Author: Rajrupa Roychowdhury

Published: 2026-09-16T08:20:52Z

Content type: opinion

Language: en

Sources: [BrowserStack Blog](<https://devfeed.tech/sources/browserstack-blog.md>)

Topics: [Testing](<https://devfeed.tech/topics/testing.md>), [Loop Engineering](<https://devfeed.tech/topics/loop-engineering.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [site-reliability-engineering](<https://devfeed.tech/topics/site-reliability-engineering.md>), [Amazon Web Services](<https://devfeed.tech/topics/aws.md>), [atlassian](<https://devfeed.tech/topics/atlassian.md>)

Tags: [agentic](<https://devfeed.tech/tags/agentic.md>), [ai](<https://devfeed.tech/tags/ai.md>), [atlassian](<https://devfeed.tech/tags/atlassian.md>), [aws](<https://devfeed.tech/tags/aws.md>), [icons-of-quality](<https://devfeed.tech/tags/icons-of-quality.md>), [quality](<https://devfeed.tech/tags/quality.md>), [sre](<https://devfeed.tech/tags/sre.md>), [testing](<https://devfeed.tech/tags/testing.md>), [tooling](<https://devfeed.tech/tags/tooling.md>)

### AI overview

BrowserStack profiles Mark Hrynczak, Canva's Head of Quality and QA Director, discussing how distributed quality ownership, agentic testing, and AI-driven decision support can help engineering teams move faster while maintaining reliability.

### Source excerpt

To celebrate the relentless passion and invaluable contributions of leaders in software quality, BrowserStack is proud to honour Icons of Quality.

## GitHub availability report: August 2026

DevFeed: [GitHub availability report: August 2026](<https://devfeed.tech/articles/github-availability-report-august-2026-83.md>)

Original publisher: [Read original article](<https://github.blog/news-insights/company-news/github-availability-report-august-2026/>)

Author: Jakub Oleksy

Published: 2026-09-10T02:05:17Z

Content type: article

Language: en

Sources: [GitHub Engineering](<https://devfeed.tech/sources/github-engineering.md>)

Topics: [GitHub](<https://devfeed.tech/topics/github.md>), [site-reliability-engineering](<https://devfeed.tech/topics/site-reliability-engineering.md>), [migration](<https://devfeed.tech/topics/migration.md>), [MySQL](<https://devfeed.tech/topics/mysql.md>), [GitHub Actions](<https://devfeed.tech/topics/github-actions.md>)

Tags: [availability](<https://devfeed.tech/tags/availability.md>), [azure](<https://devfeed.tech/tags/azure.md>), [company-news](<https://devfeed.tech/tags/company-news.md>), [database](<https://devfeed.tech/tags/database.md>), [github](<https://devfeed.tech/tags/github.md>), [github-actions](<https://devfeed.tech/tags/github-actions.md>), [github-availability-report](<https://devfeed.tech/tags/github-availability-report.md>), [migration](<https://devfeed.tech/tags/migration.md>), [monitoring](<https://devfeed.tech/tags/monitoring.md>), [mysql](<https://devfeed.tech/tags/mysql.md>), [news-insights](<https://devfeed.tech/tags/news-insights.md>), [outage](<https://devfeed.tech/tags/outage.md>), [performance](<https://devfeed.tech/tags/performance.md>), [production](<https://devfeed.tech/tags/production.md>)

### AI overview

GitHub reports five August incidents with degraded service performance and describes capacity, resiliency, database, Azure migration, and GitHub Actions improvements.

### Source excerpt

In August, we experienced five incidents that resulted in degraded performance across GitHub services. The post GitHub availability report: August 2026 appeared first on The GitHub Blog.

## From noise to signal: Monitoring Amazon DocumentDB like a pro

DevFeed: [From noise to signal: Monitoring Amazon DocumentDB like a pro](<https://devfeed.tech/articles/from-noise-to-signal-monitoring-amazon-documentdb-like-a-pro-4699.md>)

Original publisher: [Read original article](<https://aws.amazon.com/blogs/database/from-noise-to-signal-monitoring-amazon-documentdb-like-a-pro/>)

Author: Deepak Deepesh

Published: 2026-09-09T15:23:20Z

Content type: tutorial

Language: en

Sources: [AWS Database Blog](<https://devfeed.tech/sources/aws-database-blog.md>)

Topics: [Amazon DocumentDB](<https://devfeed.tech/topics/amazon-documentdb.md>), [site-reliability-engineering](<https://devfeed.tech/topics/site-reliability-engineering.md>), [Amazon CloudWatch Logs](<https://devfeed.tech/topics/amazon-cloudwatch-logs.md>), [Database](<https://devfeed.tech/topics/database.md>)

Tags: [amazon](<https://devfeed.tech/tags/amazon.md>), [amazon-cloudwatch](<https://devfeed.tech/tags/amazon-cloudwatch.md>), [amazon-cloudwatch-logs](<https://devfeed.tech/tags/amazon-cloudwatch-logs.md>), [amazon-documentdb](<https://devfeed.tech/tags/amazon-documentdb.md>), [best-practices](<https://devfeed.tech/tags/best-practices.md>), [database](<https://devfeed.tech/tags/database.md>), [devops](<https://devfeed.tech/tags/devops.md>), [expert-400](<https://devfeed.tech/tags/expert-400.md>), [health](<https://devfeed.tech/tags/health.md>), [how-to](<https://devfeed.tech/tags/how-to.md>), [logs](<https://devfeed.tech/tags/logs.md>), [metrics](<https://devfeed.tech/tags/metrics.md>), [monitoring](<https://devfeed.tech/tags/monitoring.md>), [on-call](<https://devfeed.tech/tags/on-call.md>), [operations](<https://devfeed.tech/tags/operations.md>), [performance](<https://devfeed.tech/tags/performance.md>), [production](<https://devfeed.tech/tags/production.md>)

### AI overview

A tutorial on monitoring production Amazon DocumentDB clusters with tiered alerts, CloudWatch metrics, Performance Insights, profiler logs, slow-query diagnosis, and garbage-collection health checks.

### Source excerpt

Learn how to build a tiered alerting strategy for Amazon DocumentDB by organizing CloudWatch metrics into Critical, Warning, and Advisory tiers, diagnosing slow queries through three complementary lenses, and monitoring garbage collection health before it becomes a cluster-wide incident.

## How cloud native goes AI native

DevFeed: [How cloud native goes AI native](<https://devfeed.tech/articles/how-cloud-native-goes-ai-native-4600.md>)

Original publisher: [Read original article](<https://www.cncf.io/blog/2026/09/09/how-cloud-native-goes-ai-native/>)

Author: Doron Grinstein, CEO of Control Plane

Published: 2026-09-09T08:13:23Z

Content type: opinion

Language: en

Sources: [Cloud Native Computing Foundation](<https://devfeed.tech/sources/cloud-native-computing-foundation.md>)

Topics: [cloud-infrastructure](<https://devfeed.tech/topics/cloud-infrastructure.md>), [App](<https://devfeed.tech/topics/app.md>), [coding](<https://devfeed.tech/topics/coding.md>), [site-reliability-engineering](<https://devfeed.tech/topics/site-reliability-engineering.md>), [Database](<https://devfeed.tech/topics/database.md>), [Claude](<https://devfeed.tech/topics/claude.md>), [cursor](<https://devfeed.tech/topics/cursor.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [apps](<https://devfeed.tech/tags/apps.md>), [blog](<https://devfeed.tech/tags/blog.md>), [claude](<https://devfeed.tech/tags/claude.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [cursor](<https://devfeed.tech/tags/cursor.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [production](<https://devfeed.tech/tags/production.md>), [sre](<https://devfeed.tech/tags/sre.md>)

### AI overview

The article argues that AI-assisted, "vibe-coded" applications are increasingly easy to start but often fail to reach production because they bypass cloud-native operational practices. It frames closing that production-readiness gap as a key infrastructure challenge.

### Source excerpt

"A sales guy writing code" used to be the lead-up to a joke. But now no one's laughing. Designers used to sit meekly waiting for the high priests of code to make their designs real. Now...

## Try Azure SRE Agent with no always-on charges

DevFeed: [Try Azure SRE Agent with no always-on charges](<https://devfeed.tech/articles/try-azure-sre-agent-with-no-always-on-charges-23836.md>)

Original publisher: [Read original article](<https://devblogs.microsoft.com/blog/try-azure-sre-agent-with-no-always-on-charges/>)

Author: Nir Mashkowski

Published: 2026-08-25T15:00:00Z

Content type: release

Language: en

Sources: [Developer Blogs](<https://devfeed.tech/sources/developer-blogs.md>)

Topics: [Azure](<https://devfeed.tech/topics/azure.md>), [site-reliability-engineering](<https://devfeed.tech/topics/site-reliability-engineering.md>), [telemetry](<https://devfeed.tech/topics/telemetry.md>), [incident management](<https://devfeed.tech/topics/incident-management.md>), [Cloud](<https://devfeed.tech/topics/cloud.md>)

Tags: [announcement](<https://devfeed.tech/tags/announcement.md>), [azure](<https://devfeed.tech/tags/azure.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [incident-management](<https://devfeed.tech/tags/incident-management.md>), [microsoft](<https://devfeed.tech/tags/microsoft.md>), [microsoft-for-developers](<https://devfeed.tech/tags/microsoft-for-developers.md>), [sre](<https://devfeed.tech/tags/sre.md>), [telemetry](<https://devfeed.tech/tags/telemetry.md>)

### AI overview

Microsoft announces a 30-day trial for Azure SRE Agent with no charges for setup time or keeping agents ready. The announcement also covers general availability of VNet integration and public preview of Live Reports. Active Azure Agent Unit charges apply when agents perform work.

### Source excerpt

We are happy to announce a 30-day trial experience for Azure SRE Agent. New customers can create and configure the SRE Agent at their own pace, with no charges for setup time or keeping agents ready. During the trial, you can connect your agents to telemetry, source code, incident management platforms, and other operational tools, [...] The post Try Azure SRE Agent with no always-on charges appeared first on Microsoft for Developers.

## Centralize human and agentic work with Datadog Work Management

DevFeed: [Centralize human and agentic work with Datadog Work Management](<https://devfeed.tech/articles/centralize-human-and-agentic-work-with-datadog-work-management-2319.md>)

Original publisher: [Read original article](<https://www.datadoghq.com/blog/work-management/>)

Author: Roxanne Moslehi

Published: 2026-08-18T00:00:00Z

Content type: article

Language: en

Sources: [Datadog | The Monitor blog](<https://devfeed.tech/sources/datadog-the-monitor-blog.md>)

Topics: [AI Agent](<https://devfeed.tech/topics/ai-agent.md>), [incident](<https://devfeed.tech/topics/incident.md>), [site-reliability-engineering](<https://devfeed.tech/topics/site-reliability-engineering.md>), [Slack](<https://devfeed.tech/topics/slack.md>), [SRE](<https://devfeed.tech/topics/sre.md>), [SIEM, Security](<https://devfeed.tech/topics/siem-security.md>), [GitHub](<https://devfeed.tech/topics/github.md>), [error tracking](<https://devfeed.tech/topics/error-tracking.md>), [dashboards](<https://devfeed.tech/topics/dashboards.md>), [Traces](<https://devfeed.tech/topics/traces.md>)

Tags: [agentic](<https://devfeed.tech/tags/agentic.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-agent](<https://devfeed.tech/tags/ai-agent.md>), [cloud-siem](<https://devfeed.tech/tags/cloud-siem.md>), [dashboards](<https://devfeed.tech/tags/dashboards.md>), [devops](<https://devfeed.tech/tags/devops.md>), [devsecops](<https://devfeed.tech/tags/devsecops.md>), [error-tracking](<https://devfeed.tech/tags/error-tracking.md>), [github](<https://devfeed.tech/tags/github.md>), [incident](<https://devfeed.tech/tags/incident.md>), [management](<https://devfeed.tech/tags/management.md>), [slack](<https://devfeed.tech/tags/slack.md>), [sre](<https://devfeed.tech/tags/sre.md>), [traces](<https://devfeed.tech/tags/traces.md>), [work-management](<https://devfeed.tech/tags/work-management.md>), [workflow-automation](<https://devfeed.tech/tags/workflow-automation.md>)

### AI overview

Datadog Work Management centralizes work created by people, automations, and Datadog AI agents. It preserves context from logs, traces, monitors, alerts, ownership, assignments, approvals, artifacts, and activity while integrating with Datadog and external collaboration systems.

### Source excerpt

Learn how Datadog Work Management helps you coordinate human and AI agent-driven work while preserving context, ownership, and activity across tools.

## AI SRE: Cut MTTR in Half with Autonomous Incident Resolution

DevFeed: [AI SRE: Cut MTTR in Half with Autonomous Incident Resolution](<https://devfeed.tech/articles/ai-sre-cut-mttr-in-half-with-autonomous-incident-resolution-12166.md>)

Original publisher: [Read original article](<https://www.port.io/blog/autonomous-incident-resolution>)

Author: Matar Peles

Published: 2026-08-10T11:34:38Z

Content type: article

Language: en

Sources: [Developer Experience & Platform Engineering Blog | Port](<https://devfeed.tech/sources/developer-experience-platform-engineering-blog-port.md>)

Topics: [site-reliability-engineering](<https://devfeed.tech/topics/site-reliability-engineering.md>), [incident](<https://devfeed.tech/topics/incident.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [monitor](<https://devfeed.tech/topics/monitor.md>), [dashboards](<https://devfeed.tech/topics/dashboards.md>), [Traces](<https://devfeed.tech/topics/traces.md>), [archive search](<https://devfeed.tech/topics/archive-search.md>), [Pull Request](<https://devfeed.tech/topics/pull-request.md>), [codex](<https://devfeed.tech/topics/codex.md>), [cursor](<https://devfeed.tech/topics/cursor.md>)

Tags: [agent](<https://devfeed.tech/tags/agent.md>), [ai](<https://devfeed.tech/tags/ai.md>), [codex](<https://devfeed.tech/tags/codex.md>), [cursor](<https://devfeed.tech/tags/cursor.md>), [dashboards](<https://devfeed.tech/tags/dashboards.md>), [datadog](<https://devfeed.tech/tags/datadog.md>), [incident](<https://devfeed.tech/tags/incident.md>), [monitor](<https://devfeed.tech/tags/monitor.md>), [monitoring](<https://devfeed.tech/tags/monitoring.md>), [on-call](<https://devfeed.tech/tags/on-call.md>), [sre](<https://devfeed.tech/tags/sre.md>), [traces](<https://devfeed.tech/tags/traces.md>), [workflow](<https://devfeed.tech/tags/workflow.md>)

### AI overview

This developer article describes an autonomous incident-resolution workflow built in Port. It argues that incident response time is largely spent reconstructing context from alerts, logs, dashboards, traces, ownership data, recent changes, and runbooks. The proposed agent-based workflow aims to use that full context to triage, diagnose, and fix production incidents, with a reported 50% reduction in MTTR. The supplied text ends during a discussion of why simply routing alerts to coding agents can fail.

### Source excerpt

See how an SRE agent workflow cuts MTTR in half, using full context to triage, diagnose, and fix incidents autonomously, all the way to a RCA.

## Eric Schwartz on what it takes to run an AI SRE at petabyte scale

DevFeed: [Eric Schwartz on what it takes to run an AI SRE at petabyte scale](<https://devfeed.tech/articles/eric-schwartz-on-what-it-takes-to-run-an-ai-sre-at-petabyte-scale-16015.md>)

Original publisher: [Read original article](<https://workos.com/blog/eric-schwartz-traversal-ai-sre-petabyte-scale>)

Author: WorkOS

Published: 2026-08-07T00:00:00Z

Content type: article

Language: en

Sources: [WorkOS Blog](<https://devfeed.tech/sources/workos-blog.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [site-reliability-engineering](<https://devfeed.tech/topics/site-reliability-engineering.md>), [data-platforms](<https://devfeed.tech/topics/data-platforms.md>), [telemetry](<https://devfeed.tech/topics/telemetry.md>), [observability ai agents](<https://devfeed.tech/topics/observability-ai-agents.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [data](<https://devfeed.tech/tags/data.md>), [data-platforms](<https://devfeed.tech/tags/data-platforms.md>), [on-call](<https://devfeed.tech/tags/on-call.md>), [production](<https://devfeed.tech/tags/production.md>), [root-cause-analysis](<https://devfeed.tech/tags/root-cause-analysis.md>), [sre](<https://devfeed.tech/tags/sre.md>), [telemetry](<https://devfeed.tech/tags/telemetry.md>)

### AI overview

An interview with Traversal product manager Eric Schwartz examines how the company operates an AI site reliability engineer for large enterprises. The article explains that petabyte-scale telemetry requires continuously analyzing, compressing, and indexing data ahead of runtime, with an SRE-focused tool and prompt harness. Traversal reports that deployments are generally running in production within a week with minimal tuning, supported by forward-deployed engineering for last-mile optimization.

### Source excerpt

Traversal PM Eric Schwartz on data platforms, routing models by severity, and the permission ladder toward self-driving production, from AI Engineer 2026.

## Cloud provider postmortems: volume vs depth

DevFeed: [Cloud provider postmortems: volume vs depth](<https://devfeed.tech/articles/cloud-provider-postmortems-volume-vs-depth-34008.md>)

Original publisher: [Read original article](<https://sridharrajarao.com/blog/cloud-postmortems-volume-vs-depth/>)

Author: Sridhar Rajarao

Published: 2026-08-05T00:00:00Z

Content type: opinion

Language: en

Sources: [Sridhar Rajarao](<https://devfeed.tech/sources/sridhar-rajarao.md>)

Topics: [Cloud](<https://devfeed.tech/topics/cloud.md>), [incident](<https://devfeed.tech/topics/incident.md>), [SRE](<https://devfeed.tech/topics/sre.md>), [site-reliability-engineering](<https://devfeed.tech/topics/site-reliability-engineering.md>), [engineering-culture](<https://devfeed.tech/topics/engineering-culture.md>)

Tags: [2017](<https://devfeed.tech/tags/2017.md>), [2025](<https://devfeed.tech/tags/2025.md>), [2026](<https://devfeed.tech/tags/2026.md>), [aws](<https://devfeed.tech/tags/aws.md>), [azure](<https://devfeed.tech/tags/azure.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [dynamodb](<https://devfeed.tech/tags/dynamodb.md>), [engineering-culture](<https://devfeed.tech/tags/engineering-culture.md>), [gcp](<https://devfeed.tech/tags/gcp.md>), [incident](<https://devfeed.tech/tags/incident.md>), [incident-management](<https://devfeed.tech/tags/incident-management.md>), [postmortems](<https://devfeed.tech/tags/postmortems.md>), [s3](<https://devfeed.tech/tags/s3.md>), [sre](<https://devfeed.tech/tags/sre.md>), [transparency](<https://devfeed.tech/tags/transparency.md>), [writeup](<https://devfeed.tech/tags/writeup.md>)

### AI overview

The article compares public postmortem practices among Google Cloud, Azure, and AWS. It argues that Google Cloud emphasizes high volume and speed, Azure emphasizes detailed transparency and customer accountability, and AWS publishes fewer writeups with greater depth and industry influence.

### Source excerpt

GCP publishes 100+ postmortems a year. AWS publishes almost none. Azure has become the transparency leader. What each posture reveals about engineering culture, and what SREs should steal from all three.

## Managing standards in a developer portal - a how-to guide

DevFeed: [Managing standards in a developer portal - a how-to guide](<https://devfeed.tech/articles/managing-standards-in-a-developer-portal-a-how-to-guide-12254.md>)

Original publisher: [Read original article](<https://www.port.io/blog/managing-standards-in-a-developer-portal>)

Author: Jenny Salem

Published: 2026-07-30T10:23:42Z

Content type: tutorial

Language: en

Sources: [Developer Experience & Platform Engineering Blog | Port](<https://devfeed.tech/sources/developer-experience-platform-engineering-blog-port.md>)

Topics: [internal developer portal](<https://devfeed.tech/topics/internal-developer-portal.md>), [Automation](<https://devfeed.tech/topics/automation.md>), [site-reliability-engineering](<https://devfeed.tech/topics/site-reliability-engineering.md>), [incident management](<https://devfeed.tech/topics/incident-management.md>), [Microservice](<https://devfeed.tech/topics/microservice.md>), [API](<https://devfeed.tech/topics/api.md>)

Tags: [automation](<https://devfeed.tech/tags/automation.md>), [developer](<https://devfeed.tech/tags/developer.md>), [developer-portal](<https://devfeed.tech/tags/developer-portal.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [guide](<https://devfeed.tech/tags/guide.md>), [how-to](<https://devfeed.tech/tags/how-to.md>), [incident-management](<https://devfeed.tech/tags/incident-management.md>), [production](<https://devfeed.tech/tags/production.md>), [software](<https://devfeed.tech/tags/software.md>), [sre](<https://devfeed.tech/tags/sre.md>), [standards](<https://devfeed.tech/tags/standards.md>), [technical](<https://devfeed.tech/tags/technical.md>)

### AI overview

This how-to guide explains how internal developer portals can manage production-readiness standards. It describes reliability criteria, definition-of-done checklists, standards compliance across deployment stages, and the use of automation and self-service to balance engineering checks with development speed.

### Source excerpt

Production readiness varies across engineering teams, but generally, it's a process defining reliability criteria for software in production.

## You need reliable AI context for your site reliability

DevFeed: [You need reliable AI context for your site reliability](<https://devfeed.tech/articles/you-need-reliable-ai-context-for-your-site-reliability-2196.md>)

Original publisher: [Read original article](<https://stackoverflow.blog/2026/07/28/you-need-reliable-ai-context-for-your-site-reliability/>)

Author: Phoebe Sajor

Published: 2026-07-28T07:40:00Z

Content type: article

Language: en

Sources: [Stack Overflow Blog](<https://devfeed.tech/sources/stack-overflow-blog.md>)

Topics: [site-reliability-engineering](<https://devfeed.tech/topics/site-reliability-engineering.md>), [SRE](<https://devfeed.tech/topics/sre.md>), [AI Agent](<https://devfeed.tech/topics/ai-agent.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Kubernetes](<https://devfeed.tech/topics/kubernetes.md>), [AI Platform](<https://devfeed.tech/topics/ai-platform.md>)

Tags: [agent](<https://devfeed.tech/tags/agent.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-agent](<https://devfeed.tech/tags/ai-agent.md>), [ai-engineering](<https://devfeed.tech/tags/ai-engineering.md>), [autonomous](<https://devfeed.tech/tags/autonomous.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [kubernetes](<https://devfeed.tech/tags/kubernetes.md>), [platform](<https://devfeed.tech/tags/platform.md>), [podcast](<https://devfeed.tech/tags/podcast.md>), [se-stackoverflow](<https://devfeed.tech/tags/se-stackoverflow.md>), [se-tech](<https://devfeed.tech/tags/se-tech.md>), [site-reliability-engineering](<https://devfeed.tech/tags/site-reliability-engineering.md>), [sre](<https://devfeed.tech/tags/sre.md>), [strategy](<https://devfeed.tech/tags/strategy.md>)

### AI overview

A discussion of reliable AI context for modern site-reliability work, including context engineering, Kubernetes-based infrastructure, and the changing role of human SREs in strategy and AI agent management.

### Source excerpt

Ryan is joined by Asaf Savich, Komodor's AI Engineering Group Manager, to discuss why modern reliability work requires navigating massive cross-service context, what good context engineering actually likes when AI is integrated into site reliability, and how the work of human SREs is shifting towards strategy and AI agent management.

## ITIL vs SRE: why the big clouds went their own way

DevFeed: [ITIL vs SRE: why the big clouds went their own way](<https://devfeed.tech/articles/itil-vs-sre-why-the-big-clouds-went-their-own-way-34015.md>)

Original publisher: [Read original article](<https://sridharrajarao.com/blog/itil-vs-sre/>)

Author: Sridhar Rajarao

Published: 2026-07-26T00:00:00Z

Content type: opinion

Language: en

Sources: [Sridhar Rajarao](<https://devfeed.tech/sources/sridhar-rajarao.md>)

Topics: [SRE](<https://devfeed.tech/topics/sre.md>), [site-reliability-engineering](<https://devfeed.tech/topics/site-reliability-engineering.md>), [Development](<https://devfeed.tech/topics/development.md>), [Terraform](<https://devfeed.tech/topics/terraform.md>), [pulumi](<https://devfeed.tech/topics/pulumi.md>), [Availability](<https://devfeed.tech/topics/availability.md>), [feature flags](<https://devfeed.tech/topics/feature-flags.md>), [incident](<https://devfeed.tech/topics/incident.md>)

Tags: [automated](<https://devfeed.tech/tags/automated.md>), [aws](<https://devfeed.tech/tags/aws.md>), [azure](<https://devfeed.tech/tags/azure.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [feature-flags](<https://devfeed.tech/tags/feature-flags.md>), [gcp](<https://devfeed.tech/tags/gcp.md>), [human-review](<https://devfeed.tech/tags/human-review.md>), [hyperscaler](<https://devfeed.tech/tags/hyperscaler.md>), [incident](<https://devfeed.tech/tags/incident.md>), [incident-management](<https://devfeed.tech/tags/incident-management.md>), [infrastructure-as-code](<https://devfeed.tech/tags/infrastructure-as-code.md>), [itil](<https://devfeed.tech/tags/itil.md>), [on-call](<https://devfeed.tech/tags/on-call.md>), [postmortems](<https://devfeed.tech/tags/postmortems.md>), [pulumi](<https://devfeed.tech/tags/pulumi.md>), [release](<https://devfeed.tech/tags/release.md>), [reliability](<https://devfeed.tech/tags/reliability.md>), [service-catalog](<https://devfeed.tech/tags/service-catalog.md>), [sre](<https://devfeed.tech/tags/sre.md>), [terraform](<https://devfeed.tech/tags/terraform.md>)

### AI overview

This opinion article compares ITIL practices with SRE operations at hyperscaler scale. It argues that human change boards, single production instances, developer-to-operations handoffs, documentation-first configuration management, and weekly release windows do not fit environments serving millions of external customers. It describes automated approvals, gradual deployments, service-team ownership, infrastructure as code, continuous release, error budgets, SLOs, and blameless postmortems as alternatives.

### Source excerpt

The big clouds don't run ITIL. Five assumptions ITIL makes that break at hyperscaler scale, and what AWS, Azure, GCP, and OCI use instead.

## Why platform engineering will eat the world

DevFeed: [Why platform engineering will eat the world](<https://devfeed.tech/articles/why-platform-engineering-will-eat-the-world-12277.md>)

Original publisher: [Read original article](<https://platformengineering.org/blog/why-platform-engineering-will-eat-the-world>)

Author: Kaspar von Grünberg

Published: 2026-07-23T08:08:12Z

Content type: opinion

Language: en

Sources: [Platform Engineering Blog](<https://devfeed.tech/sources/platform-engineering-blog.md>)

Topics: [Platform Engineering](<https://devfeed.tech/topics/platform-engineering.md>), [internal developer platform](<https://devfeed.tech/topics/internal-developer-platform.md>), [Developer Platform](<https://devfeed.tech/topics/developer-platform.md>), [software-development](<https://devfeed.tech/topics/software-development.md>), [Software Engineering](<https://devfeed.tech/topics/software-engineering.md>), [Developer experience](<https://devfeed.tech/topics/developer-experience.md>), [site-reliability-engineering](<https://devfeed.tech/topics/site-reliability-engineering.md>), [SRE](<https://devfeed.tech/topics/sre.md>), [observability](<https://devfeed.tech/topics/observability.md>), [Security](<https://devfeed.tech/topics/security.md>)

Tags: [developer-experience](<https://devfeed.tech/tags/developer-experience.md>), [developer-platform](<https://devfeed.tech/tags/developer-platform.md>), [development](<https://devfeed.tech/tags/development.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [internal-developer-platform](<https://devfeed.tech/tags/internal-developer-platform.md>), [operations](<https://devfeed.tech/tags/operations.md>), [platform](<https://devfeed.tech/tags/platform.md>), [platform-engineering](<https://devfeed.tech/tags/platform-engineering.md>), [security](<https://devfeed.tech/tags/security.md>), [software](<https://devfeed.tech/tags/software.md>), [software-development](<https://devfeed.tech/tags/software-development.md>), [sre](<https://devfeed.tech/tags/sre.md>)

### AI overview

The article argues that platform engineering is becoming the dominant model for large-scale software development. It describes platforms as structured digital production lines that reduce time to market, consolidate functions such as operations, databases, security, observability, and SRE, and support greater speed, scale, and security.

### Source excerpt

Platform engineering is transforming software development, streamlining operations, and redefining roles. Embrace the shift or risk becoming obsolete.

## The agent reliability score: What your AI platform must guarantee before agents go live

DevFeed: [The agent reliability score: What your AI platform must guarantee before agents go live](<https://devfeed.tech/articles/the-agent-reliability-score-what-your-ai-platform-must-guarantee-before-agents-go-live-12226.md>)

Original publisher: [Read original article](<https://platformengineering.org/blog/the-agent-reliability-score-what-your-ai-platform-must-guarantee-before-agents-go-live>)

Author: Eugene Sergueev

Published: 2026-07-23T05:40:01Z

Content type: article

Language: en

Sources: [Platform Engineering Blog](<https://devfeed.tech/sources/platform-engineering-blog.md>)

Topics: [AI Platform](<https://devfeed.tech/topics/ai-platform.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Retrieval Augmented Generation (RAG)](<https://devfeed.tech/topics/retrieval-augmented-generation-rag.md>), [Architecture & Design](<https://devfeed.tech/topics/architecture-design.md>), [Security](<https://devfeed.tech/topics/security.md>), [site-reliability-engineering](<https://devfeed.tech/topics/site-reliability-engineering.md>)

Tags: [agent](<https://devfeed.tech/tags/agent.md>), [agents](<https://devfeed.tech/tags/agents.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-agents](<https://devfeed.tech/tags/ai-agents.md>), [article](<https://devfeed.tech/tags/article.md>), [developers](<https://devfeed.tech/tags/developers.md>), [knowledge-base](<https://devfeed.tech/tags/knowledge-base.md>), [model](<https://devfeed.tech/tags/model.md>), [platform](<https://devfeed.tech/tags/platform.md>), [production](<https://devfeed.tech/tags/production.md>), [rag](<https://devfeed.tech/tags/rag.md>), [security](<https://devfeed.tech/tags/security.md>)

### AI overview

The article introduces the Agent Reliability Score, a 28-test framework for evaluating whether an AI platform is ready to deploy agents in production. It argues that agent failures often result from missing platform guarantees around context validation, freshness, grounding, action safety, and monitoring rather than from model limitations.

### Source excerpt

Use the Agent Reliability Score, a 28-test framework, to evaluate your AI platform's readiness. Ensure reliability contracts, context validation, and guardrails before agents go live

## Runbook Best Practices and Automated Incident Response with Harness AI SRE

DevFeed: [Runbook Best Practices and Automated Incident Response with Harness AI SRE](<https://devfeed.tech/articles/runbooks-for-modern-ops-best-practices-ai-sre-13420.md>)

Original publisher: [Read original article](<https://www.harness.io/blog/how-to-build-runbooks-that-work----and-automate-them-with-harness-ai-sre>)

Author: Ryan Taylor

Published: 2026-07-15T00:00:00Z

Content type: tutorial

Language: en

Sources: [Harness Blog](<https://devfeed.tech/sources/harness-blog.md>)

Topics: [site-reliability-engineering](<https://devfeed.tech/topics/site-reliability-engineering.md>), [Incident response](<https://devfeed.tech/topics/incident-response.md>), [Automation](<https://devfeed.tech/topics/automation.md>)

Tags: [automation](<https://devfeed.tech/tags/automation.md>), [best-practices](<https://devfeed.tech/tags/best-practices.md>), [how-to](<https://devfeed.tech/tags/how-to.md>), [incident](<https://devfeed.tech/tags/incident.md>), [incident-response](<https://devfeed.tech/tags/incident-response.md>), [on-call](<https://devfeed.tech/tags/on-call.md>), [operations](<https://devfeed.tech/tags/operations.md>), [sre](<https://devfeed.tech/tags/sre.md>), [sre-automation](<https://devfeed.tech/tags/sre-automation.md>)

### AI overview

This tutorial explains how to create effective software-system runbooks that are actionable, accessible, accurate, authoritative, and adaptable. It also describes how Harness AI SRE can automate runbook execution during incidents, including filing tickets, triggering rollbacks, and posting incident-timeline updates.

### Source excerpt

Learn what makes a runbook effective, how to keep them accurate and actionable, and how Harness AI SRE automates runbook execution during incidents. | Blog

## AI Reliability Engineering

DevFeed: [AI Reliability Engineering](<https://devfeed.tech/articles/ai-reliability-engineering-29073.md>)

Original publisher: [Read original article](<https://blog.alexewerlof.com/p/ai-reliability-engineering>)

Author: Alex Ewerlöf

Published: 2026-07-12T17:34:19Z

Content type: opinion

Language: en

Sources: [Alex Ewerlof Notes](<https://devfeed.tech/sources/alex-ewerlof-notes.md>)

Topics: [AI Engineering](<https://devfeed.tech/topics/ai-engineering.md>), [site-reliability-engineering](<https://devfeed.tech/topics/site-reliability-engineering.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [Generative AI](<https://devfeed.tech/topics/generative-ai.md>), [Software Engineering](<https://devfeed.tech/topics/software-engineering.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-engineering](<https://devfeed.tech/tags/ai-engineering.md>), [generative-ai](<https://devfeed.tech/tags/generative-ai.md>), [llms](<https://devfeed.tech/tags/llms.md>), [reliability-engineering](<https://devfeed.tech/tags/reliability-engineering.md>), [software-engineering](<https://devfeed.tech/tags/software-engineering.md>), [sre](<https://devfeed.tech/tags/sre.md>)

### AI overview

This article examines how site reliability engineering practices can be adapted for AI systems and AI-generated black boxes. It focuses on the need to run these systems predictably, securely, and at scale as large language models and their supporting harnesses become more capable.

### Source excerpt

Why SRE is a key skill in the age of AI-generated black boxes and how to renovate the traditional toolbox for the new era

## Your genie is vanishing: introducing the Opsgenie rescue program

DevFeed: [Your genie is vanishing: introducing the Opsgenie rescue program](<https://devfeed.tech/articles/your-genie-is-vanishing-introducing-the-opsgenie-rescue-program-11859.md>)

Original publisher: [Read original article](<https://incident.io/blog/introducing-the-opsgenie-rescue-program>)

Author: Tom Wentworth

Published: 2026-07-09T13:30:00Z

Content type: article

Language: en

Sources: [The incident.io Blog](<https://devfeed.tech/sources/the-incident-io-blog.md>)

Topics: [migration](<https://devfeed.tech/topics/migration.md>), [atlassian](<https://devfeed.tech/topics/atlassian.md>), [incident](<https://devfeed.tech/topics/incident.md>), [site-reliability-engineering](<https://devfeed.tech/topics/site-reliability-engineering.md>)

Tags: [atlassian](<https://devfeed.tech/tags/atlassian.md>), [incident](<https://devfeed.tech/tags/incident.md>), [incident-channel](<https://devfeed.tech/tags/incident-channel.md>), [incident-management](<https://devfeed.tech/tags/incident-management.md>), [incident-response](<https://devfeed.tech/tags/incident-response.md>), [migration](<https://devfeed.tech/tags/migration.md>), [outage](<https://devfeed.tech/tags/outage.md>), [post-mortem](<https://devfeed.tech/tags/post-mortem.md>), [slack-incident](<https://devfeed.tech/tags/slack-incident.md>), [sre](<https://devfeed.tech/tags/sre.md>)

### AI overview

incident.io introduces the Opsgenie Rescue Program for customers affected by Atlassian's planned shutdown of Opsgenie. The program offers simplified migration assistance and free overlap, allowing customers to run both systems in parallel without paying two vendors while they validate and complete the transition.

### Source excerpt

Today, we're launching the Opsgenie Rescue Program to make that landing soft: simplified migration and free overlap so you never pay two vendors at once.

## A Career Journey from Cryptanalysis Research to SRE and Chief Editor at Microsoft

DevFeed: [A Career Journey from Cryptanalysis Research to SRE and Chief Editor at Microsoft](<https://devfeed.tech/articles/navigating-the-ocean-32261.md>)

Original publisher: [Read original article](<https://medium.com/data-science-at-microsoft/navigating-the-ocean-ef276deeed8c?source=rss----a6e43238cdaf---4>)

Author: Alexandra Savelieva

Published: 2026-07-07T07:16:01Z

Content type: opinion

Language: en

Sources: [Data Science at Microsoft](<https://devfeed.tech/sources/data-science-at-microsoft.md>)

Topics: [Data Science](<https://devfeed.tech/topics/data-science.md>), [Microsoft](<https://devfeed.tech/topics/microsoft.md>), [Cryptography](<https://devfeed.tech/topics/cryptography.md>), [site-reliability-engineering](<https://devfeed.tech/topics/site-reliability-engineering.md>), [Algorithms, Complexity](<https://devfeed.tech/topics/algorithms-complexity.md>), [audit trail](<https://devfeed.tech/topics/audit-trail.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [algorithms](<https://devfeed.tech/tags/algorithms.md>), [audit-trail](<https://devfeed.tech/tags/audit-trail.md>), [careers](<https://devfeed.tech/tags/careers.md>), [crack](<https://devfeed.tech/tags/crack.md>), [data-science](<https://devfeed.tech/tags/data-science.md>), [logs](<https://devfeed.tech/tags/logs.md>), [machine-learning](<https://devfeed.tech/tags/machine-learning.md>), [microsoft](<https://devfeed.tech/tags/microsoft.md>), [paper](<https://devfeed.tech/tags/paper.md>), [sre](<https://devfeed.tech/tags/sre.md>)

### AI overview

The author reflects on becoming chief editor of Microsoft's Data Science + AI online publication and recounts a career spanning cryptanalysis research, Bing Ads R&D, and site reliability engineering.

### Source excerpt

Thoughts on taking the helm of DS@MImage by the author (generated with ChatGPT). This week marks an important chapter of Data Science + AI at Microsoft, as well as in my own professional journey, as I step into the role of chief editor of this online publication after the farewell of its wonderful founder, Casey Doyle. This change is not something that I had planned -- rather it's a combination of unexpected circumstances have come together to make it happen, like many other things that have shaped my career and enabled this opportunity. If you read Casey's farewell article from last week, you may see why a metaphor of navigating the ocean came to mind when I was thinking about what's next for DS@M now that I'm "captaining the ship." The first chapter of my journey took place in 2010 as a Ph.D. intern in Microsoft Research. Under the supervision of Dmitry Khovratovich, I studied block hash functions. The paper that I coauthored ended up making a big splash in cryptanalysis (see Biclique attack -- Wikipedia), and a fun fact is that it took two years and several rejections at conferences and workshops for it to be recognized. Our approach involved surprisingly simple math and deterministic algorithms to crack a problem that was previously considered a "puzzle" requiring some craft with a bit of luck to solve. This was my main takeaway from the internship: twist and dissect the complex problems until they get reduced to an intuitively understood form, so that solving them becomes a matter of applying the right calculus. I returned in October 2012 as a full-time employee in Bing Ads R&D. I was expecting an applied research job and ended up as an SRE (Site Reliability Engineer) in the Audit Trail service working with logs collected for customer ads. It was "type 2" fun work -- absolutely not fun in the moment, but exciting when I look back. It served as a practical crash course that left the "bible" of SRE imprinted in my brain: how to design services for reliability, how t

## Customers over control: how we measure On-call reliability

DevFeed: [Customers over control: how we measure On-call reliability](<https://devfeed.tech/articles/customers-over-control-how-we-measure-on-call-reliability-11739.md>)

Original publisher: [Read original article](<https://incident.io/blog/customers-over-control>)

Author: Mike Fisher

Published: 2026-05-28T16:29:00Z

Content type: article

Language: en

Sources: [The incident.io Blog](<https://devfeed.tech/sources/the-incident-io-blog.md>)

Topics: [SRE](<https://devfeed.tech/topics/sre.md>), [Availability](<https://devfeed.tech/topics/availability.md>), [site-reliability-engineering](<https://devfeed.tech/topics/site-reliability-engineering.md>), [nginx](<https://devfeed.tech/topics/nginx.md>), [Amazon Web Services](<https://devfeed.tech/topics/aws.md>), [API](<https://devfeed.tech/topics/api.md>), [Network](<https://devfeed.tech/topics/network.md>), [incident](<https://devfeed.tech/topics/incident.md>)

Tags: [2025](<https://devfeed.tech/tags/2025.md>), [api](<https://devfeed.tech/tags/api.md>), [availability](<https://devfeed.tech/tags/availability.md>), [aws](<https://devfeed.tech/tags/aws.md>), [customers](<https://devfeed.tech/tags/customers.md>), [http](<https://devfeed.tech/tags/http.md>), [incident](<https://devfeed.tech/tags/incident.md>), [incident-channel](<https://devfeed.tech/tags/incident-channel.md>), [incident-management](<https://devfeed.tech/tags/incident-management.md>), [incident-response](<https://devfeed.tech/tags/incident-response.md>), [load-balancer](<https://devfeed.tech/tags/load-balancer.md>), [metrics](<https://devfeed.tech/tags/metrics.md>), [network](<https://devfeed.tech/tags/network.md>), [on-call](<https://devfeed.tech/tags/on-call.md>), [outage](<https://devfeed.tech/tags/outage.md>), [post-mortem](<https://devfeed.tech/tags/post-mortem.md>), [slack-incident](<https://devfeed.tech/tags/slack-incident.md>), [sre](<https://devfeed.tech/tags/sre.md>)

### AI overview

This article explains how incident.io measures the reliability of its On-call product from the customer's perspective. It focuses on two critical functions, defines SLIs and monthly SLOs, and describes monitoring at the GCP load balancer, alerting, replicated components, and lessons from an AWS outage.

### Source excerpt

Instead of thinking about reliability as an exercise in figuring out what we can control, and ignoring anything beyond that, we think about what we'll be really proud to offer to customers.

## Storage at scale: what I actually watched

DevFeed: [Storage at scale: what I actually watched](<https://devfeed.tech/articles/storage-at-scale-what-i-actually-watched-34024.md>)

Original publisher: [Read original article](<https://sridharrajarao.com/blog/storage-at-scale/>)

Author: Sridhar Rajarao

Published: 2026-05-28T00:00:00Z

Content type: article

Language: en

Sources: [Sridhar Rajarao](<https://devfeed.tech/sources/sridhar-rajarao.md>)

Topics: [dashboards](<https://devfeed.tech/topics/dashboards.md>), [site-reliability-engineering](<https://devfeed.tech/topics/site-reliability-engineering.md>), [Latency](<https://devfeed.tech/topics/latency.md>), [Availability](<https://devfeed.tech/topics/availability.md>), [control-plane](<https://devfeed.tech/topics/control-plane.md>), [Databases](<https://devfeed.tech/topics/databases.md>), [Replication](<https://devfeed.tech/topics/replication.md>)

Tags: [availability](<https://devfeed.tech/tags/availability.md>), [control-plane](<https://devfeed.tech/tags/control-plane.md>), [database](<https://devfeed.tech/tags/database.md>), [latency](<https://devfeed.tech/tags/latency.md>), [metrics](<https://devfeed.tech/tags/metrics.md>), [reliability](<https://devfeed.tech/tags/reliability.md>), [replication](<https://devfeed.tech/tags/replication.md>), [service](<https://devfeed.tech/tags/service.md>), [sre](<https://devfeed.tech/tags/sre.md>), [storage](<https://devfeed.tech/tags/storage.md>)

### AI overview

An SRE practitioner describes seven metrics for judging the health of a large-scale storage service, covering availability, durability, latency, synthetic canaries, hotspots, IOPS, and metadata-database behavior.

### Source excerpt

For eight years I ran SRE for a storage system measured in exabytes. The dashboard I checked every morning shrank to seven numbers. Here they are.

## Three 5xx and one 4xx: the codes I actually care about

DevFeed: [Three 5xx and one 4xx: the codes I actually care about](<https://devfeed.tech/articles/three-5xx-and-one-4xx-the-codes-i-actually-care-about-34014.md>)

Original publisher: [Read original article](<https://sridharrajarao.com/blog/http-codes-at-scale/>)

Author: Sridhar Rajarao

Published: 2026-05-26T00:00:00Z

Content type: tutorial

Language: en

Sources: [Sridhar Rajarao](<https://devfeed.tech/sources/sridhar-rajarao.md>)

Topics: [HTTP](<https://devfeed.tech/topics/http.md>), [site-reliability-engineering](<https://devfeed.tech/topics/site-reliability-engineering.md>), [API](<https://devfeed.tech/topics/api.md>), [nginx](<https://devfeed.tech/topics/nginx.md>), [servers](<https://devfeed.tech/topics/servers.md>)

Tags: [api](<https://devfeed.tech/tags/api.md>), [http](<https://devfeed.tech/tags/http.md>), [load-balancer](<https://devfeed.tech/tags/load-balancer.md>), [nginx](<https://devfeed.tech/tags/nginx.md>), [production](<https://devfeed.tech/tags/production.md>), [reliability](<https://devfeed.tech/tags/reliability.md>), [scale](<https://devfeed.tech/tags/scale.md>), [sre](<https://devfeed.tech/tags/sre.md>)

### AI overview

An SRE-oriented guide to interpreting HTTP status codes 500, 502, 503, and 429 in production. It connects each code with likely causes and recommended operational responses, including checking logs, investigating backend failures, load shedding, and client backoff.

### Source excerpt

How I read 500, 502, 503, and 429 in production at scale, and what each one is really telling you.

## If Hope is Your Strategy, You're Doing On-Call Wrong

DevFeed: [If Hope is Your Strategy, You're Doing On-Call Wrong](<https://devfeed.tech/articles/if-hope-is-your-strategy-you-re-doing-on-call-wrong-17870.md>)

Original publisher: [Read original article](<https://www.codemotion.com/magazine/backend/software-architecture/if-hope-is-your-strategy-youre-doing-on-call-wrong/>)

Author: Natalia de Pablo Garcia

Published: 2026-05-19T10:05:05Z

Content type: article

Language: en

Sources: [Backend Job: skill, salary and insights - Codemotion Magazine](<https://devfeed.tech/sources/backend-job-skill-salary-and-insights-codemotion-magazine.md>)

Topics: [DevOps](<https://devfeed.tech/topics/devops.md>), [Monitoring](<https://devfeed.tech/topics/monitoring.md>), [site-reliability-engineering](<https://devfeed.tech/topics/site-reliability-engineering.md>), [Security](<https://devfeed.tech/topics/security.md>)

Tags: [careers](<https://devfeed.tech/tags/careers.md>), [devops](<https://devfeed.tech/tags/devops.md>), [monitoring](<https://devfeed.tech/tags/monitoring.md>), [on-call](<https://devfeed.tech/tags/on-call.md>), [operational](<https://devfeed.tech/tags/operational.md>), [software-architecture](<https://devfeed.tech/tags/software-architecture.md>)

### AI overview

The article promotes a Codemotion Madrid 2026 talk about improving on-call operations through proactive monitoring, well-defined runbooks, and intelligent alerting. It argues that these practices help engineering teams manage incidents and operate critical systems at scale, particularly in finance.

### Source excerpt

Codemotion Madrid 2026 is fine-tuning every detail to welcome a new edition packed with knowledge, business opportunities, and networking. Among the wide range of topics ahead, DevOps will play a key role at a time when, although everything moves at breakneck speed, security and stability remain non-negotiable. In the talk "If Hope is Your Strategy,... Read more The post If Hope is Your Strategy, You're Doing On-Call Wrong appeared first on Codemotion Magazine.

## What's new in ClickStack - March 2026

DevFeed: [What's new in ClickStack - March 2026](<https://devfeed.tech/articles/what-s-new-in-clickstack-march-2026-5648.md>)

Original publisher: [Read original article](<https://clickhouse.com/blog/whats-new-in-clickstack-march-2026>)

Author: The ClickStack Team

Published: 2026-04-14T16:36:29Z

Content type: release

Language: en

Sources: [ClickHouse Blog](<https://devfeed.tech/sources/clickhouse-blog.md>)

Topics: [clickhouse](<https://devfeed.tech/topics/clickhouse.md>), [observability](<https://devfeed.tech/topics/observability.md>), [dashboards](<https://devfeed.tech/topics/dashboards.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [SQL](<https://devfeed.tech/topics/sql.md>), [Traces](<https://devfeed.tech/topics/traces.md>), [site-reliability-engineering](<https://devfeed.tech/topics/site-reliability-engineering.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [clickhouse](<https://devfeed.tech/tags/clickhouse.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [dashboards](<https://devfeed.tech/tags/dashboards.md>), [logs](<https://devfeed.tech/tags/logs.md>), [metrics](<https://devfeed.tech/tags/metrics.md>), [observability](<https://devfeed.tech/tags/observability.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [release](<https://devfeed.tech/tags/release.md>), [root-cause-analysis](<https://devfeed.tech/tags/root-cause-analysis.md>), [sql](<https://devfeed.tech/tags/sql.md>), [sre](<https://devfeed.tech/tags/sre.md>), [traces](<https://devfeed.tech/tags/traces.md>)

### AI overview

ClickStack's March 2026 release adds smarter Event Deltas for root cause analysis, AI Notebooks for investigating logs, metrics, and traces, expanded SQL chart capabilities, sampled-trace aggregation improvements, persistent local dashboards, and dashboard organization enhancements.

### Source excerpt

ClickStack's March release introduces smarter Event Deltas, AI-powered notebooks, and full SQL chart flexibility, making investigation faster and dashboards more powerful.

## IBM and UC Berkeley Diagnose Why Enterprise Agents Fail Using IT-Bench and MAST

DevFeed: [IBM and UC Berkeley Diagnose Why Enterprise Agents Fail Using IT-Bench and MAST](<https://devfeed.tech/articles/ibm-and-uc-berkeley-diagnose-why-enterprise-agents-fail-using-it-bench-and-mast-7267.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/ibm-research/itbenchandmast>)

Author: Ayhan Sebin; Rohan Arora; Saurabh Jha

Published: 2026-02-18T16:15:45Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [ibm](<https://devfeed.tech/topics/ibm.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [site-reliability-engineering](<https://devfeed.tech/topics/site-reliability-engineering.md>), [incident](<https://devfeed.tech/topics/incident.md>), [Kubernetes](<https://devfeed.tech/topics/kubernetes.md>), [finops](<https://devfeed.tech/topics/finops.md>), [Security](<https://devfeed.tech/topics/security.md>)

Tags: [agentic](<https://devfeed.tech/tags/agentic.md>), [agents](<https://devfeed.tech/tags/agents.md>), [analysis](<https://devfeed.tech/tags/analysis.md>), [automation](<https://devfeed.tech/tags/automation.md>), [benchmarks](<https://devfeed.tech/tags/benchmarks.md>), [enterprise](<https://devfeed.tech/tags/enterprise.md>), [evaluation](<https://devfeed.tech/tags/evaluation.md>), [gpt-oss](<https://devfeed.tech/tags/gpt-oss.md>), [hallucinations](<https://devfeed.tech/tags/hallucinations.md>), [ibm](<https://devfeed.tech/tags/ibm.md>), [incident](<https://devfeed.tech/tags/incident.md>), [kubernetes](<https://devfeed.tech/tags/kubernetes.md>), [llm](<https://devfeed.tech/tags/llm.md>), [logs](<https://devfeed.tech/tags/logs.md>), [loops](<https://devfeed.tech/tags/loops.md>), [metrics](<https://devfeed.tech/tags/metrics.md>), [performance](<https://devfeed.tech/tags/performance.md>), [research](<https://devfeed.tech/tags/research.md>), [systems](<https://devfeed.tech/tags/systems.md>), [tool](<https://devfeed.tech/tags/tool.md>)

### AI overview

IBM Research and UC Berkeley analyze why agentic LLM systems fail in enterprise IT automation using ITBench traces and the MAST failure taxonomy. Their analysis of 310 SRE traces compares Gemini-3-Flash, Kimi-K2, and GPT-OSS-120B, identifying verification errors, cascading failures, premature termination, looping, and hallucinations as major reliability problems.

### Source excerpt

IBM Research and UC Berkeley collaborated to study how agentic LLM systems break in real-world IT automation, for tasks involving incident triage, logs/metrics queries, and Kubernetes actions in long-horizon tool loops. Benchmarks typically reduce performance to a single number, telling you whether an agent failed but never why. To solve this black-box problem, we applied MAST (Multi-Agent System Failure Taxonomy), an emerging practice for diagnosing agentic reliability ).

[Next page](<https://devfeed.tech/topics/site-reliability-engineering.md?cursor=WyIyMDI2LTAyLTE4VDE2OjE1OjQ1KzAwOjAwIiwgImU1NzU2MzE1LTBlMjctNDAwNi1iYmMwLTUxNzUzM2UwOWQ4YyJd>)