# site-reliability-engineering

Published articles for site-reliability-engineering.

This is one page of public article previews, not the complete archive. Follow Next page to continue. Summaries are not the original full articles.

## 1Password increases engineering productivity 21% with Codex

DevFeed: [1Password increases engineering productivity 21% with Codex](<https://devfeed.tech/articles/1password-increases-engineering-productivity-21-with-codex-6264.md>)

Original publisher: [Read original article](<https://openai.com/index/1password>)

Published: 2026-09-08T00:00:00Z

Content type: article

Language: en

Sources: [OpenAI News](<https://devfeed.tech/sources/openai-news.md>)

Topics: [AI-assisted coding](<https://devfeed.tech/topics/ai-assisted-coding.md>)

Tags: [codex](<https://devfeed.tech/tags/codex.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [features](<https://devfeed.tech/tags/features.md>), [internal-tools](<https://devfeed.tech/tags/internal-tools.md>), [productivity](<https://devfeed.tech/tags/productivity.md>), [pull-request](<https://devfeed.tech/tags/pull-request.md>), [security](<https://devfeed.tech/tags/security.md>), [site-reliability-engineering](<https://devfeed.tech/tags/site-reliability-engineering.md>), [software-delivery](<https://devfeed.tech/tags/software-delivery.md>), [user-stories](<https://devfeed.tech/tags/user-stories.md>), [writing-code](<https://devfeed.tech/tags/writing-code.md>)

### AI overview

1Password reports using Codex across its software delivery lifecycle to turn user stories into near-final prototypes and features more quickly. It cites improved engineering productivity and shorter pull request cycle times while maintaining security standards.

### Source excerpt

Engineers at 1Password use Codex to rapidly build new features and internal tools, reaching production-readiness while maintaining rigorous security policies.

## Как мониторить Java-приложения: метрики, алерты и правило 80/20

DevFeed: [Как мониторить Java-приложения: метрики, алерты и правило 80/20](<https://devfeed.tech/articles/java-80-20-24881.md>)

Original publisher: [Read original article](<https://habr.com/ru/companies/yandex/articles/1068874/>)

Author: atushkanova (Яндекс)

Published: 2026-08-12T07:01:43Z

Content type: tutorial

Language: ru

Sources: [Яндекс - Как мы делаем Яндекс / Статьи](<https://devfeed.tech/sources/source.md>)

Topics: [Java](<https://devfeed.tech/topics/java.md>), [яндекс](<https://devfeed.tech/topics/tag-4004cf5948d3.md>)

Tags: [devops](<https://devfeed.tech/tags/devops.md>), [java](<https://devfeed.tech/tags/java.md>), [site-reliability-engineering](<https://devfeed.tech/tags/site-reliability-engineering.md>), [sre](<https://devfeed.tech/tags/sre.md>), [tag-4004cf5948d3](<https://devfeed.tech/tags/tag-4004cf5948d3.md>), [tag-617c72e79812](<https://devfeed.tech/tags/tag-617c72e79812.md>), [tag-622612a5798e](<https://devfeed.tech/tags/tag-622612a5798e.md>), [tag-65aad6d83235](<https://devfeed.tech/tags/tag-65aad6d83235.md>), [tag-73eb9b712998](<https://devfeed.tech/tags/tag-73eb9b712998.md>), [tag-7568ee66754b](<https://devfeed.tech/tags/tag-7568ee66754b.md>), [tag-a6a98345cb30](<https://devfeed.tech/tags/tag-a6a98345cb30.md>), [tag-b92bf5906bbd](<https://devfeed.tech/tags/tag-b92bf5906bbd.md>)

### AI overview

This article explains how to monitor Java applications using a focused set of technical metrics, useful alerts, business metrics, SLOs, and anomaly analysis. It presents an 80/20 approach and warns that excessive metrics and flapping alerts create information noise.

### Source excerpt

Хороший мониторинг помогает быстро понять, что происходит с приложением и куда смотреть в первую очередь. Для этого не нужно пытаться измерить всё: базовый набор технических метрик покрывает большинство типовых проблем, а бизнес-метрики, SLO и анализ аномалий помогают заранее замечать нетипичные отклонения. В Календаре мы называем этот подход правилом 80/20. Всем привет! Меня зовут Настя, я бэкенд-разработчик в Яндекс 360 и отвечаю за надёжность Календаря. В этой статье я покажу, какие метрики стоит взять за основу, как выбирать полезные алерты и чем дополнять базовый набор для оставшихся 20%. Читать далее

## You need reliable AI context for your site reliability

DevFeed: [You need reliable AI context for your site reliability](<https://devfeed.tech/articles/you-need-reliable-ai-context-for-your-site-reliability-2196.md>)

Original publisher: [Read original article](<https://stackoverflow.blog/2026/07/28/you-need-reliable-ai-context-for-your-site-reliability/>)

Author: Phoebe Sajor

Published: 2026-07-28T07:40:00Z

Content type: article

Language: en

Sources: [Stack Overflow Blog](<https://devfeed.tech/sources/stack-overflow-blog.md>)

Topics: [site-reliability-engineering](<https://devfeed.tech/topics/site-reliability-engineering.md>), [SRE](<https://devfeed.tech/topics/sre.md>), [AI Agent](<https://devfeed.tech/topics/ai-agent.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Kubernetes](<https://devfeed.tech/topics/kubernetes.md>), [AI Platform](<https://devfeed.tech/topics/ai-platform.md>)

Tags: [agent](<https://devfeed.tech/tags/agent.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-agent](<https://devfeed.tech/tags/ai-agent.md>), [ai-engineering](<https://devfeed.tech/tags/ai-engineering.md>), [autonomous](<https://devfeed.tech/tags/autonomous.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [kubernetes](<https://devfeed.tech/tags/kubernetes.md>), [platform](<https://devfeed.tech/tags/platform.md>), [podcast](<https://devfeed.tech/tags/podcast.md>), [se-stackoverflow](<https://devfeed.tech/tags/se-stackoverflow.md>), [se-tech](<https://devfeed.tech/tags/se-tech.md>), [site-reliability-engineering](<https://devfeed.tech/tags/site-reliability-engineering.md>), [sre](<https://devfeed.tech/tags/sre.md>), [strategy](<https://devfeed.tech/tags/strategy.md>)

### AI overview

A discussion of reliable AI context for modern site-reliability work, including context engineering, Kubernetes-based infrastructure, and the changing role of human SREs in strategy and AI agent management.

### Source excerpt

Ryan is joined by Asaf Savich, Komodor's AI Engineering Group Manager, to discuss why modern reliability work requires navigating massive cross-service context, what good context engineering actually likes when AI is integrated into site reliability, and how the work of human SREs is shifting towards strategy and AI agent management.

## How Service Level Objectives Align Developers and Product Managers

DevFeed: [How Service Level Objectives Align Developers and Product Managers](<https://devfeed.tech/articles/how-not-to-fight-with-product-managers-as-a-developer-28056.md>)

Original publisher: [Read original article](<https://tech.trivago.com/post/2026-02-02-how-not-to-fight-with-product-managers-as-a-developer/>)

Author: Anis Khan Site Reliability Engineering is my role Cost optimization is my goal GitHub profile Linkedin profile

Published: 2026-02-02T00:00:00Z

Content type: opinion

Language: en

Sources: [Trivago](<https://devfeed.tech/sources/trivago.md>)

Topics: [Availability](<https://devfeed.tech/topics/availability.md>), [User Experience](<https://devfeed.tech/topics/user-experience.md>), [data](<https://devfeed.tech/topics/data.md>)

Tags: [availability](<https://devfeed.tech/tags/availability.md>), [developer](<https://devfeed.tech/tags/developer.md>), [engineering-culture](<https://devfeed.tech/tags/engineering-culture.md>), [maintainability](<https://devfeed.tech/tags/maintainability.md>), [opinions](<https://devfeed.tech/tags/opinions.md>), [outages](<https://devfeed.tech/tags/outages.md>), [performance](<https://devfeed.tech/tags/performance.md>), [site-reliability-engineering](<https://devfeed.tech/tags/site-reliability-engineering.md>), [technical](<https://devfeed.tech/tags/technical.md>), [uptime](<https://devfeed.tech/tags/uptime.md>)

### AI overview

This article argues that Service Level Objectives (SLOs) can reduce conflict between developers and product managers by creating a shared, data-driven agreement about reliability and user experience. It explains how a 99.9% availability target creates a 43.2-minute monthly downtime budget and shows how that budget can guide decisions about releasing features versus restoring stability.

### Source excerpt

It's a scenario developers relate to a little too well. The product manager always comes with more and more feature requests. They also want to release fast by giving a tight deadline. While you...

## How Faster Screen Loading Affected Product Metrics in the hh.ru Mobile App

DevFeed: [How Faster Screen Loading Affected Product Metrics in the hh.ru Mobile App](<https://devfeed.tech/articles/hh-30696.md>)

Original publisher: [Read original article](<https://habr.com/ru/companies/hh/articles/977376/>)

Author: alektas (hh.ru)

Published: 2025-12-17T05:50:44Z

Content type: article

Language: ru

Sources: [HeadHunter RU](<https://devfeed.tech/sources/headhunter-ru.md>)

Topics: [SRE](<https://devfeed.tech/topics/sre.md>), [Android](<https://devfeed.tech/topics/android.md>), [iOS](<https://devfeed.tech/topics/ios.md>)

Tags: [ab-d64cb2e3f30f](<https://devfeed.tech/tags/ab-d64cb2e3f30f.md>), [android](<https://devfeed.tech/tags/android.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [hh-ru](<https://devfeed.tech/tags/hh-ru.md>), [ios](<https://devfeed.tech/tags/ios.md>), [site-reliability](<https://devfeed.tech/tags/site-reliability.md>), [site-reliability-engineering](<https://devfeed.tech/tags/site-reliability-engineering.md>), [sre](<https://devfeed.tech/tags/sre.md>), [tag-b92bf5906bbd](<https://devfeed.tech/tags/tag-b92bf5906bbd.md>), [tag-d1a8ff602d37](<https://devfeed.tech/tags/tag-d1a8ff602d37.md>), [tag-f538878e20ff](<https://devfeed.tech/tags/tag-f538878e20ff.md>), [tag-fa773991fb10](<https://devfeed.tech/tags/tag-fa773991fb10.md>), [ux](<https://devfeed.tech/tags/ux.md>)

### AI overview

An hh.ru development team shares the preparation and interpretation of an A/B experiment that optimized one mobile-app screen, accelerated content loading, and examined the effect on product metrics. The article also discusses applying SRE practices to mobile development.

### Source excerpt

Привет! Меня зовут Саша Тотилас и я руковожу командой разработки в hh.ru. Хочу поделиться с Хабром результатами A/B-эксперимента: при оптимизации одного из экранов нашего приложения мы ускорили загрузку контента и выяснили, как это влияет на продуктовые метрики, а также собрали интересные инсайты. Я не буду глубоко погружаться в технические детали, а сосредоточусь на подготовке эксперимента и интерпретации результатов. Статья будет полезна не только для мобильных разработчиков, но и для аналитиков и продактов. Читать далее

## Life of SRE as a Salesperson

DevFeed: [Life of SRE as a Salesperson](<https://devfeed.tech/articles/life-of-sre-as-a-salesperson-28050.md>)

Original publisher: [Read original article](<https://tech.trivago.com/post/2024-12-20-life-of-sre-as-a-salesperson/>)

Author: Anis Khan Site Reliability Engineering is my role Cost optimization is my goal GitHub profile Linkedin profile

Published: 2025-03-17T00:00:00Z

Content type: opinion

Language: en

Sources: [Trivago](<https://devfeed.tech/sources/trivago.md>)

Topics: [site-reliability-engineering](<https://devfeed.tech/topics/site-reliability-engineering.md>), [site-reliability-engineer](<https://devfeed.tech/topics/site-reliability-engineer.md>), [Platform Engineering](<https://devfeed.tech/topics/platform-engineering.md>), [Developer experience](<https://devfeed.tech/topics/developer-experience.md>), [observability](<https://devfeed.tech/topics/observability.md>), [cloud-infrastructure](<https://devfeed.tech/topics/cloud-infrastructure.md>)

Tags: [developer](<https://devfeed.tech/tags/developer.md>), [developer-experience](<https://devfeed.tech/tags/developer-experience.md>), [development](<https://devfeed.tech/tags/development.md>), [engineering-culture](<https://devfeed.tech/tags/engineering-culture.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [observability](<https://devfeed.tech/tags/observability.md>), [platform-engineering](<https://devfeed.tech/tags/platform-engineering.md>), [site-reliability-engineering](<https://devfeed.tech/tags/site-reliability-engineering.md>), [sre](<https://devfeed.tech/tags/sre.md>)

### AI overview

A Site Reliability Engineer at trivago describes acting as a salesperson or influencer for technology teams by helping product teams adopt practices, tools, and solutions developed with Platform Engineering, Developer Experience, and Observability teams. The article also outlines trivago's SRE mission and squad structure.

### Source excerpt

If you are a Developer or a Product person, you might have this feeling of achievement when you work on a specific product. When it's launched successfully in the market, you s...

## Test Failures Should Be Actionable

DevFeed: [Test Failures Should Be Actionable](<https://devfeed.tech/articles/test-failures-should-be-actionable-23857.md>)

Original publisher: [Read original article](<http://testing.googleblog.com/2024/05/test-failures-should-be-actionable.html>)

Author: Google Testing Bloggers (noreply@blogger.com)

Published: 2024-05-06T13:26:00Z

Content type: article

Language: en

Sources: [Google Testing Blog](<https://devfeed.tech/sources/google-testing-blog.md>)

Topics: [Testing](<https://devfeed.tech/topics/testing.md>), [Unit testing](<https://devfeed.tech/topics/unit-testing.md>), [Software Engineering](<https://devfeed.tech/topics/software-engineering.md>), [site-reliability-engineering](<https://devfeed.tech/topics/site-reliability-engineering.md>), [C++](<https://devfeed.tech/topics/c-plus-plus.md>), [Pytest](<https://devfeed.tech/topics/pytest.md>)

Tags: [article](<https://devfeed.tech/tags/article.md>), [best-practices](<https://devfeed.tech/tags/best-practices.md>), [c-plus-plus](<https://devfeed.tech/tags/c-plus-plus.md>), [pytest](<https://devfeed.tech/tags/pytest.md>), [site-reliability-engineering](<https://devfeed.tech/tags/site-reliability-engineering.md>), [software-engineering](<https://devfeed.tech/tags/software-engineering.md>), [testing](<https://devfeed.tech/tags/testing.md>), [titus-winters](<https://devfeed.tech/tags/titus-winters.md>), [tott](<https://devfeed.tech/tags/tott.md>), [unit-testing](<https://devfeed.tech/tags/unit-testing.md>)

### AI overview

The article argues that unit test failures should be actionable: developers should be able to start investigating using only the test name and failure messages. It recommends precise invariants and assertion-library matchers, illustrating the point with a C++ status-check example.

### Source excerpt

This article was adapted from a Google Testing on the Toilet (TotT) episode. You can download a printer-friendly version of this TotT episode and post it in your office. By Titus Winters There are a lot of rules and best practices around unit testing. There are many posts on this blog; there is deeper material in the Software Engineering at Google book; there is specific guidance for every major language; there is guidance on test frameworks, test naming, and dozens of other test-related topics. Isn't this excessive? Good unit tests contain several important properties, but you could focus on a key principle: Test failures should be actionable. When a test fails, you should be able to begin investigation with nothing more than the test's name and its failure messages--no need to add more information and rerun the test. Effective use of unit test frameworks and assertion libraries (JUnit, Truth, pytest, GoogleTest, etc.) serves two important purposes. Firstly, the more precisely we express the invariants we are testing, the more informative and less brittle our tests will be. Secondly, when those invariants don't hold and the tests fail, the failure info should be immediately actionable. This meshes well with Site Reliability Engineering guidance on alerting. Consider this example of a C++ unit test of a function returning an absl::Status (an Abseil type that returns either an "OK" status or one of a number of different error codes): EXPECT_TRUE(LoadMetadata().ok()); EXPECT_OK(LoadMetadata()); Sample failure output load_metadata_test.cc:42: Failure Value of: LoadMetadata().ok() Expected: true Actual: false load_metadata_test.cc:42: Failure Value of: LoadMetadata() Expected: is OK Actual: NOT_FOUND: /path/to/metadata.bin If the test on the left fails, you have to investigate why the test failed; the test on the right immediately gives you all the available detail, in this case because of a more precise GoogleTest matcher. Here are some other posts on this blog that emp

## What I Learned in One Year as an SRE Trainee

DevFeed: [What I Learned in One Year as an SRE Trainee](<https://devfeed.tech/articles/what-i-learned-in-one-year-as-an-sre-trainee-2156.md>)

Original publisher: [Read original article](<https://developers.soundcloud.com/blog//sre-trainee>)

Published: 2023-01-06T00:00:00Z

Content type: article

Language: en

Sources: [SoundCloud Backstage Blog](<https://devfeed.tech/sources/soundcloud-backstage-blog.md>)

Topics: [SRE](<https://devfeed.tech/topics/sre.md>), [Production Engineering](<https://devfeed.tech/topics/production-engineering.md>), [Infrastructure as code](<https://devfeed.tech/topics/infrastructure-as-code.md>), [Monitoring](<https://devfeed.tech/topics/monitoring.md>), [Incident response](<https://devfeed.tech/topics/incident-response.md>), [Kubernetes](<https://devfeed.tech/topics/kubernetes.md>), [migration](<https://devfeed.tech/topics/migration.md>), [datacenter](<https://devfeed.tech/topics/datacenter.md>), [Cloud](<https://devfeed.tech/topics/cloud.md>), [Google](<https://devfeed.tech/topics/google.md>)

Tags: [cloud](<https://devfeed.tech/tags/cloud.md>), [data-center](<https://devfeed.tech/tags/data-center.md>), [engineering-culture](<https://devfeed.tech/tags/engineering-culture.md>), [google-cloud](<https://devfeed.tech/tags/google-cloud.md>), [incident-response](<https://devfeed.tech/tags/incident-response.md>), [infrastructure-as-code](<https://devfeed.tech/tags/infrastructure-as-code.md>), [knowledge-sharing](<https://devfeed.tech/tags/knowledge-sharing.md>), [kubernetes](<https://devfeed.tech/tags/kubernetes.md>), [learning](<https://devfeed.tech/tags/learning.md>), [migration](<https://devfeed.tech/tags/migration.md>), [monitoring](<https://devfeed.tech/tags/monitoring.md>), [music](<https://devfeed.tech/tags/music.md>), [production-engineering](<https://devfeed.tech/tags/production-engineering.md>), [site-reliability-engineering](<https://devfeed.tech/tags/site-reliability-engineering.md>), [sre](<https://devfeed.tech/tags/sre.md>)

### AI overview

A Site Reliability Engineering trainee at SoundCloud reflects on a first year spent learning how SRE applies software engineering to operations, reliability, and scalability. The article discusses the broad scope of the role, including Infrastructure as Code, Monitoring, Incident Response training, Kubernetes upgrades, infrastructure decommissioning, and a service migration from a data center to Google Cloud.

### Source excerpt

I recently celebrated my one year anniversary as a Site Reliability Engineering (SRE) trainee at SoundCloud. Looking back, I had very little...

## Alerting on SLOs like Pros

DevFeed: [Alerting on SLOs like Pros](<https://devfeed.tech/articles/alerting-on-slos-like-pros-1984.md>)

Original publisher: [Read original article](<https://developers.soundcloud.com/blog//alerting-on-slos>)

Published: 2019-06-04T00:00:00Z

Content type: article

Language: en

Sources: [SoundCloud Backstage Blog](<https://devfeed.tech/sources/soundcloud-backstage-blog.md>)

Topics: [SRE](<https://devfeed.tech/topics/sre.md>), [Monitoring](<https://devfeed.tech/topics/monitoring.md>), [Prometheus](<https://devfeed.tech/topics/prometheus.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [Google](<https://devfeed.tech/topics/google.md>)

Tags: [blog-post](<https://devfeed.tech/tags/blog-post.md>), [development](<https://devfeed.tech/tags/development.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [monitoring](<https://devfeed.tech/tags/monitoring.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [production](<https://devfeed.tech/tags/production.md>), [prometheus](<https://devfeed.tech/tags/prometheus.md>), [site-reliability-engineering](<https://devfeed.tech/tags/site-reliability-engineering.md>), [sre](<https://devfeed.tech/tags/sre.md>)

### AI overview

This article explains how alerting on service-level objectives (SLOs) and error-budget burn can produce meaningful, actionable alerts while reducing unnecessary pages for on-call engineers. It describes SoundCloud's SRE approach and its use of Prometheus for monitoring.

### Source excerpt

If there is anything like a silver bullet for creating meaningful and actionable alerts with a high signal-to-noise ratio, it is alerting based on service-level objectives (SLOs). Fulfilling a well-defined SLO is the very definition of meeting your users' expectations. Conversely, a certain level of service errors is OK as long as you stay within the SLO -- in other words, if the SLO grants you an error budget. Burning through this error budget too quickly is the ultimate signal that some rectifying action is needed. The faster the budget is burned, the more urgent it is that engineers get involved. This post describes how we implemented this concept at SoundCloud, enabling us to fulfill our SLOs without flooding our engineers on call with an unsustainable amount of pages.