# Post Mortem

A structured process following a technical incident that reviews what occurred, why it happened, how it was resolved, and how to prevent recurrence.

This is one page of public article previews, not the complete archive. Follow Next page to continue. Summaries are not the original full articles.

## OpenAI and Hugging Face partner to address security incident during model evaluation

DevFeed: [OpenAI and Hugging Face partner to address security incident during model evaluation](<https://devfeed.tech/articles/openai-and-hugging-face-partner-to-address-security-incident-during-model-evaluation-6466.md>)

Original publisher: [Read original article](<https://openai.com/index/hugging-face-model-evaluation-security-incident>)

Published: 2026-07-21T07:00:00Z

Content type: article

Language: en

Sources: [OpenAI News](<https://devfeed.tech/sources/openai-news.md>)

Topics: [incident](<https://devfeed.tech/topics/incident.md>), [hugging face](<https://devfeed.tech/topics/hugging-face.md>), [Security](<https://devfeed.tech/topics/security.md>), [Vulnerabilities](<https://devfeed.tech/topics/vulnerabilities.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [OpenAI](<https://devfeed.tech/topics/openai.md>), [LLM evaluation / benchmarking](<https://devfeed.tech/topics/llm-evaluation-benchmarking.md>), [Incident response](<https://devfeed.tech/topics/incident-response.md>), [Post Mortem](<https://devfeed.tech/topics/post-mortem.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [evaluation](<https://devfeed.tech/tags/evaluation.md>), [hugging-face](<https://devfeed.tech/tags/hugging-face.md>), [incident](<https://devfeed.tech/tags/incident.md>), [incident-response](<https://devfeed.tech/tags/incident-response.md>), [openai](<https://devfeed.tech/tags/openai.md>), [post-mortem](<https://devfeed.tech/tags/post-mortem.md>), [security](<https://devfeed.tech/tags/security.md>), [vulnerabilities](<https://devfeed.tech/tags/vulnerabilities.md>)

### AI overview

OpenAI and Hugging Face share findings from a security incident that occurred during AI model evaluation. The updates describe a platform-level compromise, exploitation of a previously unknown Artifactory vulnerability, exposed credentials used by models, and ongoing third-party assessment and incident-response work.

### Source excerpt

OpenAI and Hugging Face share early findings from a security incident during AI model evaluation, highlighting advanced cyber capabilities and lessons for defenders.

## Post-mortem GPU crash debugging with LLMs

DevFeed: [Post-mortem GPU crash debugging with LLMs](<https://devfeed.tech/articles/post-mortem-gpu-crash-debugging-with-llms-15044.md>)

Original publisher: [Read original article](<https://gpuopen.com/learn/post-mortem-gpu-crash-debugging-with-llms/>)

Author: Amit Ben-Moshe; Amit Mulay

Published: 2026-07-14T12:00:00Z

Content type: article

Language: en

Sources: [AMD GPUOpen](<https://devfeed.tech/sources/amd-gpuopen.md>)

Topics: [GPU](<https://devfeed.tech/topics/gpu.md>), [debugging](<https://devfeed.tech/topics/debugging.md>), [Model Context Protocol](<https://devfeed.tech/topics/model-context-protocol.md>), [MCP Server](<https://devfeed.tech/topics/mcp-server.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [Code](<https://devfeed.tech/topics/code.md>), [Post Mortem](<https://devfeed.tech/topics/post-mortem.md>)

Tags: [announcement](<https://devfeed.tech/tags/announcement.md>), [coding-agents](<https://devfeed.tech/tags/coding-agents.md>), [debugging](<https://devfeed.tech/tags/debugging.md>), [getting-started](<https://devfeed.tech/tags/getting-started.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [gpu-open-tools](<https://devfeed.tech/tags/gpu-open-tools.md>), [gpuopen-third-party](<https://devfeed.tech/tags/gpuopen-third-party.md>), [gpuopen-tools](<https://devfeed.tech/tags/gpuopen-tools.md>), [graphics-apis](<https://devfeed.tech/tags/graphics-apis.md>), [llms](<https://devfeed.tech/tags/llms.md>), [mcp](<https://devfeed.tech/tags/mcp.md>), [mcp-server](<https://devfeed.tech/tags/mcp-server.md>), [memory](<https://devfeed.tech/tags/memory.md>), [microsoft](<https://devfeed.tech/tags/microsoft.md>), [ml](<https://devfeed.tech/tags/ml.md>), [news](<https://devfeed.tech/tags/news.md>), [product-release](<https://devfeed.tech/tags/product-release.md>), [quick-start](<https://devfeed.tech/tags/quick-start.md>), [radeon-developer-tool-suite](<https://devfeed.tech/tags/radeon-developer-tool-suite.md>), [radeon-gpu-detective](<https://devfeed.tech/tags/radeon-gpu-detective.md>), [rdts](<https://devfeed.tech/tags/rdts.md>), [rgd](<https://devfeed.tech/tags/rgd.md>), [root-cause-analysis](<https://devfeed.tech/tags/root-cause-analysis.md>), [technical-article](<https://devfeed.tech/tags/technical-article.md>), [technical-articles](<https://devfeed.tech/tags/technical-articles.md>), [third-party](<https://devfeed.tech/tags/third-party.md>), [tools](<https://devfeed.tech/tags/tools.md>), [user-guides-manuals](<https://devfeed.tech/tags/user-guides-manuals.md>)

### AI overview

This article introduces the open-source AMD Radeon GPU Detective MCP Server, which gives LLMs structured access to GPU crash dumps and application source code for post-mortem debugging. It describes a workflow in which an LLM investigates crash evidence and suggests source-code fixes through a single natural-language prompt.

### Source excerpt

The new AMD RGD MCP Server connects LLM agents to AMD's GPU crash analysis pipeline, turning a single prompt into root-cause analysis and source-code fix suggestions.

## Post Mortem: HTTP Request Smuggling Vulnerability

DevFeed: [Post Mortem: HTTP Request Smuggling Vulnerability](<https://devfeed.tech/articles/post-mortem-http-request-smuggling-vulnerability-22335.md>)

Original publisher: [Read original article](<https://crystal-lang.org/2026/05/26/http-request-smuggling-vulnerability-in-http-server/>)

Author: Julien Portalier

Published: 2026-05-26T00:00:00Z

Content type: article

Language: en

Sources: [Crystal](<https://devfeed.tech/sources/crystal.md>)

Topics: [Post Mortem](<https://devfeed.tech/topics/post-mortem.md>), [vulnerability](<https://devfeed.tech/topics/vulnerability.md>), [Crystal](<https://devfeed.tech/topics/crystal.md>), [HTTP](<https://devfeed.tech/topics/http.md>), [Security](<https://devfeed.tech/topics/security.md>), [Exploit](<https://devfeed.tech/topics/exploit.md>)

Tags: [2026](<https://devfeed.tech/tags/2026.md>), [access-control](<https://devfeed.tech/tags/access-control.md>), [http](<https://devfeed.tech/tags/http.md>), [post-mortem](<https://devfeed.tech/tags/post-mortem.md>), [rate-limiting](<https://devfeed.tech/tags/rate-limiting.md>), [security](<https://devfeed.tech/tags/security.md>), [server](<https://devfeed.tech/tags/server.md>), [vulnerability](<https://devfeed.tech/tags/vulnerability.md>)

### AI overview

This post-mortem describes an HTTP request smuggling vulnerability in Crystal's HTTP server. The parser mishandled requests containing conflicting framing headers, allowing request injection between a reverse proxy and a Crystal server under specific proxy conditions. The issue was patched in Crystal 1.20.0 and 1.19.2.

### Source excerpt

On 12 April 2026, we received a vulnerability report regarding an HTTP request smuggling vulnerability in HTTP::Server.

## Adaptable apps on ChromeOS: a post-mortem

DevFeed: [Adaptable apps on ChromeOS: a post-mortem](<https://devfeed.tech/articles/adaptable-apps-on-chromeos-a-post-mortem-26155.md>)

Original publisher: [Read original article](<https://dev.to/tkuenneth/adaptable-apps-on-chromeos-a-post-mortem-2gl1>)

Author: Thomas Künneth

Published: 2026-05-23T11:26:02Z

Content type: article

Language: en

Sources: [Thomas Künneth](<https://devfeed.tech/sources/thomas-kunneth.md>)

Topics: [Post Mortem](<https://devfeed.tech/topics/post-mortem.md>), [chromebook](<https://devfeed.tech/topics/chromebook.md>), [Android](<https://devfeed.tech/topics/android.md>), [App](<https://devfeed.tech/topics/app.md>), [Development](<https://devfeed.tech/topics/development.md>)

Tags: [android](<https://devfeed.tech/tags/android.md>), [app](<https://devfeed.tech/tags/app.md>), [chromeos](<https://devfeed.tech/tags/chromeos.md>), [coding](<https://devfeed.tech/tags/coding.md>), [community](<https://devfeed.tech/tags/community.md>), [dev](<https://devfeed.tech/tags/dev.md>), [development](<https://devfeed.tech/tags/development.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [inclusive](<https://devfeed.tech/tags/inclusive.md>), [post-mortem](<https://devfeed.tech/tags/post-mortem.md>), [software](<https://devfeed.tech/tags/software.md>), [uidesign](<https://devfeed.tech/tags/uidesign.md>)

### AI overview

A post-mortem of attempts to remove ChromeOS's resize confirmation for the Be nice Android app in ARC windows. The article explains how the warning appears in windowed laptop mode, why manifest changes did not remove it on Play Store builds, and how tablet mode differs.

### Source excerpt

In my previous article Building a custom launcher for ChromeOS I described how Be nice runs on Chromebooks: not as a real default home app, because default home settings (Settings.ACTION_HOME_SETTINGS) usually are not available on ChromeOS, but as a normal Android app in an ARC window. I pretended the app is the launcher in code (detectIsHomeApp() returns true on ChromeOS), worked around split-screen bugs that leave the app a black rectangle, and gave up transparent wallpaper for an opaque scaffold because ARC does not show the ChromeOS desktop behind the window the way a phone shows its wallpaper. What that article obviously could not cover is the fight that came right after. For about a week in May 2026 I tried to get rid of a system confirmation ChromeOS shows when users open the preset-size menu in the ARC window title bar and choose Resizable, at least on Play Store builds. During my experiments, manifest changes often looked fine from Android Studio; however, on Play, the dialog stayed. I tried adding XML, asking AI assistants, including Google's own models, for the magic flag. Here's what I learned along the way. What users see ChromeOS does not present Play Store Android apps the same way in every posture. On a convertible or clamshell Chromebook in laptop use, ARC usually places apps in individual windows (chromeos.dev). The title bar often labels the current layout (Phone, Tablet, or Resizable) and opens a menu to switch presets; choosing Resizable is what triggers the warning below, not dragging the window frame by itself (in fact, resizing the window by dragging one of the window edges does not work until it is allowed through the dialog). In tablet mode the picture is different: apps commonly launch and stay full screen, which matches how ChromeOS has long treated touch-first use on detachables (Chrome Unboxed on immersive mode). The resize presets and the warning below matter most when you are in a windowed, laptop-style layout, not when an app is alre

## What does using AI for post-mortems actually mean?

DevFeed: [What does using AI for post-mortems actually mean?](<https://devfeed.tech/articles/what-does-using-ai-for-post-mortems-actually-mean-12068.md>)

Original publisher: [Read original article](<https://incident.io/blog/what-does-using-ai-for-post-mortems-actually-mean>)

Author: incident.io

Published: 2026-04-23T14:29:00Z

Content type: opinion

Language: en

Sources: [The incident.io Blog](<https://devfeed.tech/sources/the-incident-io-blog.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Post Mortem](<https://devfeed.tech/topics/post-mortem.md>), [Incident response](<https://devfeed.tech/topics/incident-response.md>), [incident](<https://devfeed.tech/topics/incident.md>), [Compression](<https://devfeed.tech/topics/compression.md>), [Slack](<https://devfeed.tech/topics/slack.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [channel](<https://devfeed.tech/tags/channel.md>), [compression](<https://devfeed.tech/tags/compression.md>), [docs](<https://devfeed.tech/tags/docs.md>), [incident](<https://devfeed.tech/tags/incident.md>), [incident-channel](<https://devfeed.tech/tags/incident-channel.md>), [incident-management](<https://devfeed.tech/tags/incident-management.md>), [incident-response](<https://devfeed.tech/tags/incident-response.md>), [outage](<https://devfeed.tech/tags/outage.md>), [post](<https://devfeed.tech/tags/post.md>), [post-mortem](<https://devfeed.tech/tags/post-mortem.md>), [slack](<https://devfeed.tech/tags/slack.md>), [slack-incident](<https://devfeed.tech/tags/slack-incident.md>), [thread](<https://devfeed.tech/tags/thread.md>)

### AI overview

The article argues that AI can make post-mortems faster by gathering Slack threads, timelines, pull requests, and custom fields into a useful starting point. However, it warns that polished AI-generated documents can become useless if they replace the team's work of understanding what happened, why it happened, and what was learned. It distinguishes compression, which AI can automate effectively, from synthesis, which requires genuine human reasoning.

### Source excerpt

Everyone is using AI to help with post-mortems now. We've built AI into our own post-mortem experience, pulling your Slack thread, timeline, PRs, and custom fields together and giving your team a meaningful starting point in seconds. But "AI for post-mortems" can mean very different things. Post-mortems are just one stage of the broader incident response lifecycle, and AI's role varies across all of it.

## Why post-mortem action items die

DevFeed: [Why post-mortem action items die](<https://devfeed.tech/articles/why-post-mortem-action-items-die-12088.md>)

Original publisher: [Read original article](<https://incident.io/blog/why-post-mortem-action-items-die>)

Author: Kate Bernacchi-Sass

Published: 2026-04-16T13:55:19Z

Content type: opinion

Language: en

Sources: [The incident.io Blog](<https://devfeed.tech/sources/the-incident-io-blog.md>)

Topics: [Post Mortem](<https://devfeed.tech/topics/post-mortem.md>), [incident](<https://devfeed.tech/topics/incident.md>), [Incident response](<https://devfeed.tech/topics/incident-response.md>), [Monitoring](<https://devfeed.tech/topics/monitoring.md>)

Tags: [incident](<https://devfeed.tech/tags/incident.md>), [incident-channel](<https://devfeed.tech/tags/incident-channel.md>), [incident-management](<https://devfeed.tech/tags/incident-management.md>), [incident-response](<https://devfeed.tech/tags/incident-response.md>), [monitoring](<https://devfeed.tech/tags/monitoring.md>), [on-call](<https://devfeed.tech/tags/on-call.md>), [outage](<https://devfeed.tech/tags/outage.md>), [post-mortem](<https://devfeed.tech/tags/post-mortem.md>), [slack-incident](<https://devfeed.tech/tags/slack-incident.md>), [technical](<https://devfeed.tech/tags/technical.md>)

### AI overview

The article explains why post-mortem action items often fail after incident debriefs. It identifies unclear ownership, storage outside the team's normal work tools, vague directions instead of verifiable tasks, and lack of follow-up checks as common failure modes.

### Source excerpt

You can run the best debrief of your life. Honest timeline, blameless tone, real insights. People leave the room nodding. And then nothing happens. Here's how to fix that.

## Wayland set the Linux Desktop back by 10 years

DevFeed: [Wayland set the Linux Desktop back by 10 years](<https://devfeed.tech/articles/wayland-set-the-linux-desktop-back-by-10-years-34275.md>)

Original publisher: [Read original article](<https://omar.yt/wayland-set-the-linux-desktop-back-by-10-years>)

Author: Omar Roth (blog@omar.yt)

Published: 2026-03-18T11:40:32Z

Content type: opinion

Language: en

Sources: [Omar Roth](<https://devfeed.tech/sources/omar-roth.md>)

Topics: [Wayland](<https://devfeed.tech/topics/wayland.md>), [Linux](<https://devfeed.tech/topics/linux.md>), [Post Mortem](<https://devfeed.tech/topics/post-mortem.md>)

Tags: [developer](<https://devfeed.tech/tags/developer.md>), [linux](<https://devfeed.tech/tags/linux.md>), [post-mortem](<https://devfeed.tech/tags/post-mortem.md>), [wayland](<https://devfeed.tech/tags/wayland.md>)

### AI overview

An opinionated engineering post-mortem argues that Wayland misdirected developer effort and harmed Linux desktop users. It frames the discussion around the limitations of existing systems, why they may not be fixable, and the goals and expected timeline of a new project.

### Source excerpt

Wayland has been a broad misdirection and misallocation of time and developer resources at the expense of users. If you're not familiar with this space, hopefully it will still be interesting as an engineering post-mortem on taking on new greenfield projects. Namely: What are the issues with what exists, why can they not be fixed, what do we hope to achieve with a new project, and how long do we expect it to take?

## We rebuilt our post-mortems from the ground up

DevFeed: [We rebuilt our post-mortems from the ground up](<https://devfeed.tech/articles/we-rebuilt-our-post-mortems-from-the-ground-up-11961.md>)

Original publisher: [Read original article](<https://incident.io/blog/post-mortems-launch>)

Author: Pete Hamilton

Published: 2026-03-17T13:30:00Z

Content type: release

Language: en

Sources: [The incident.io Blog](<https://devfeed.tech/sources/the-incident-io-blog.md>)

Topics: [Post Mortem](<https://devfeed.tech/topics/post-mortem.md>), [incident](<https://devfeed.tech/topics/incident.md>), [Slack](<https://devfeed.tech/topics/slack.md>), [real-time](<https://devfeed.tech/topics/real-time.md>), [Tooling](<https://devfeed.tech/topics/tooling.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [incident](<https://devfeed.tech/tags/incident.md>), [incident-channel](<https://devfeed.tech/tags/incident-channel.md>), [incident-management](<https://devfeed.tech/tags/incident-management.md>), [incident-response](<https://devfeed.tech/tags/incident-response.md>), [new-feature](<https://devfeed.tech/tags/new-feature.md>), [outage](<https://devfeed.tech/tags/outage.md>), [post](<https://devfeed.tech/tags/post.md>), [post-mortem](<https://devfeed.tech/tags/post-mortem.md>), [real-time](<https://devfeed.tech/tags/real-time.md>), [slack](<https://devfeed.tech/tags/slack.md>), [slack-incident](<https://devfeed.tech/tags/slack-incident.md>), [tooling](<https://devfeed.tech/tags/tooling.md>)

### AI overview

incident.io announces a rebuilt post-mortems experience designed to reduce the effort of drafting, maintaining, collaborating on, and managing post-mortems. The update brings incident data into a synced editor, supports AI-assisted post-mortem workflows, enables real-time collaborative editing and threaded comments, and adds tooling for tracking post-mortem status, overdue work, and readership at scale.

### Source excerpt

Today we're launching our new post-mortems experience, and I want to walk you through what we've done and why.

## The post-mortem problem

DevFeed: [The post-mortem problem](<https://devfeed.tech/articles/the-post-mortem-problem-12034.md>)

Original publisher: [Read original article](<https://incident.io/blog/the-post-mortem-problem>)

Author: incident.io

Published: 2026-03-04T19:36:00Z

Content type: article

Language: en

Sources: [The incident.io Blog](<https://devfeed.tech/sources/the-incident-io-blog.md>)

Topics: [Post Mortem](<https://devfeed.tech/topics/post-mortem.md>), [Software Engineering](<https://devfeed.tech/topics/software-engineering.md>), [incident](<https://devfeed.tech/topics/incident.md>)

Tags: [culture](<https://devfeed.tech/tags/culture.md>), [incident](<https://devfeed.tech/tags/incident.md>), [incident-channel](<https://devfeed.tech/tags/incident-channel.md>), [incident-management](<https://devfeed.tech/tags/incident-management.md>), [incident-response](<https://devfeed.tech/tags/incident-response.md>), [outage](<https://devfeed.tech/tags/outage.md>), [post-mortem](<https://devfeed.tech/tags/post-mortem.md>), [slack-incident](<https://devfeed.tech/tags/slack-incident.md>), [software-engineering](<https://devfeed.tech/tags/software-engineering.md>), [writing](<https://devfeed.tech/tags/writing.md>)

### AI overview

The article argues that post-mortems often fail because they become compliance artifacts instead of useful parts of the incident response lifecycle. It examines timing, overly demanding templates, the difficulty of post-mortem writing, and organizational culture, emphasizing that effective documents preserve context, explain technical events clearly, and support learning.

### Source excerpt

Post-mortems are one of the most consistently underperforming rituals in software engineering. Most teams do them. Most teams know theirs aren't working. And most teams reach for the same diagnosis: the templates are too long, nobody has time, nobody reads them anyway.

## IA en sysadmin: ¿oportunidad o amenaza?

DevFeed: [IA en sysadmin: ¿oportunidad o amenaza?](<https://devfeed.tech/articles/ia-en-sysadmin-oportunidad-o-amenaza-34063.md>)

Original publisher: [Read original article](<https://tengoping.com/blog/ia-administracion-sistemas-oportunidad-amenaza/>)

Author: Alois

Published: 2026-01-13T00:00:00Z

Content type: opinion

Language: es

Sources: [tengoping.com](<https://devfeed.tech/sources/tengoping-com.md>)

Topics: [AIOps](<https://devfeed.tech/topics/aiops.md>), [datacenter](<https://devfeed.tech/topics/datacenter.md>), [Infrastructure as code](<https://devfeed.tech/topics/infrastructure-as-code.md>), [Terraform](<https://devfeed.tech/topics/terraform.md>), [Ansible](<https://devfeed.tech/topics/ansible.md>), [Command-line interface](<https://devfeed.tech/topics/cli.md>), [Post Mortem](<https://devfeed.tech/topics/post-mortem.md>), [telemetry](<https://devfeed.tech/topics/telemetry.md>)

Tags: [aiops](<https://devfeed.tech/tags/aiops.md>), [ansible](<https://devfeed.tech/tags/ansible.md>), [datacenter](<https://devfeed.tech/tags/datacenter.md>), [iac](<https://devfeed.tech/tags/iac.md>), [logs](<https://devfeed.tech/tags/logs.md>), [post-mortem](<https://devfeed.tech/tags/post-mortem.md>), [sysadmin](<https://devfeed.tech/tags/sysadmin.md>), [terminal](<https://devfeed.tech/tags/terminal.md>), [terraform](<https://devfeed.tech/tags/terraform.md>)

### AI overview

This Spanish-language article examines how AI can assist systems administrators with log triage, anomaly detection, IaC manifest generation, incident summaries, runbook drafts, and command suggestions. It argues that AI reduces mechanical work but does not replace human diagnosis, review, or operational judgment.

### Source excerpt

Reflexión sobre dónde ayuda la IA al sysadmin, qué riesgos reales conlleva en producción y cómo adaptarse sin perder criterio.

## Beyond Detection: Building a Resilient Software Supply Chain (Lessons from the Shai-Hulud Post-Mortem)

DevFeed: [Beyond Detection: Building a Resilient Software Supply Chain (Lessons from the Shai-Hulud Post-Mortem)](<https://devfeed.tech/articles/beyond-detection-building-a-resilient-software-supply-chain-lessons-from-the-shai-hulud-post-mortem-8099.md>)

Original publisher: [Read original article](<https://snyk.io/blog/shai-hulud-post-mortem/>)

Author: Liran Tal

Published: 2026-01-08T05:00:00Z

Content type: article

Language: en

Sources: [Blog RSS Feed | Snyk](<https://devfeed.tech/sources/blog-rss-feed-snyk.md>)

Topics: [Post Mortem](<https://devfeed.tech/topics/post-mortem.md>), [supply-chain-security](<https://devfeed.tech/topics/supply-chain-security.md>), [incident](<https://devfeed.tech/topics/incident.md>), [npm](<https://devfeed.tech/topics/npm.md>), [Security](<https://devfeed.tech/topics/security.md>), [Malware](<https://devfeed.tech/topics/malware.md>), [AI-assisted coding](<https://devfeed.tech/topics/ai-assisted-coding.md>), [ci](<https://devfeed.tech/topics/ci.md>), [GitHub Copilot](<https://devfeed.tech/topics/github-copilot.md>), [Vulnerabilities](<https://devfeed.tech/topics/vulnerabilities.md>)

Tags: [ai-coding](<https://devfeed.tech/tags/ai-coding.md>), [americas](<https://devfeed.tech/tags/americas.md>), [awareness](<https://devfeed.tech/tags/awareness.md>), [blog](<https://devfeed.tech/tags/blog.md>), [ci](<https://devfeed.tech/tags/ci.md>), [ci-cd](<https://devfeed.tech/tags/ci-cd.md>), [copilot](<https://devfeed.tech/tags/copilot.md>), [developer](<https://devfeed.tech/tags/developer.md>), [devops](<https://devfeed.tech/tags/devops.md>), [github](<https://devfeed.tech/tags/github.md>), [github-copilot](<https://devfeed.tech/tags/github-copilot.md>), [google](<https://devfeed.tech/tags/google.md>), [incident](<https://devfeed.tech/tags/incident.md>), [interest](<https://devfeed.tech/tags/interest.md>), [javascript](<https://devfeed.tech/tags/javascript.md>), [malware](<https://devfeed.tech/tags/malware.md>), [node-js](<https://devfeed.tech/tags/node-js.md>), [post-mortem](<https://devfeed.tech/tags/post-mortem.md>), [related-content](<https://devfeed.tech/tags/related-content.md>), [security](<https://devfeed.tech/tags/security.md>), [security-labs](<https://devfeed.tech/tags/security-labs.md>), [snyk-apprisk](<https://devfeed.tech/tags/snyk-apprisk.md>), [snyk-code](<https://devfeed.tech/tags/snyk-code.md>), [snyk-open-source](<https://devfeed.tech/tags/snyk-open-source.md>), [snyk-platform](<https://devfeed.tech/tags/snyk-platform.md>), [snyk-security-intel](<https://devfeed.tech/tags/snyk-security-intel.md>), [supply-chain](<https://devfeed.tech/tags/supply-chain.md>), [tech](<https://devfeed.tech/tags/tech.md>), [vulnerabilities](<https://devfeed.tech/tags/vulnerabilities.md>)

### AI overview

The article examines the Shai-Hulud npm supply chain incident and argues that organizations should move beyond reactive scanning toward layered prevention, real-time intelligence, and automated action. It highlights safer dependency upgrades, a 21-day cooldown strategy, and security guardrails embedded in AI coding workflows.

### Source excerpt

The Shai-Hulud npm incident exposed the limitations of reactive security in modern software supply chains. To survive the next major attack, organizations must shift toward a multi-layered strategy of proactive prevention, real-time intelligence, and automated action.

## Reliability lessons from the 2025 Cloudflare outage

DevFeed: [Reliability lessons from the 2025 Cloudflare outage](<https://devfeed.tech/articles/reliability-lessons-from-the-2025-cloudflare-outage-11697.md>)

Original publisher: [Read original article](<https://www.gremlin.com/blog/reliability-lessons-from-the-2025-cloudflare-outage>)

Author: Andre Newman

Published: 2025-11-20T00:00:00Z

Content type: article

Language: en

Sources: [Gremlin Blog](<https://devfeed.tech/sources/gremlin-blog.md>)

Topics: [Cloudflare](<https://devfeed.tech/topics/cloudflare.md>), [Bot Management](<https://devfeed.tech/topics/bot-management.md>), [configuration](<https://devfeed.tech/topics/configuration.md>), [Chaos Engineering](<https://devfeed.tech/topics/chaos-engineering.md>), [Post Mortem](<https://devfeed.tech/topics/post-mortem.md>), [Network](<https://devfeed.tech/topics/network.md>), [Workers](<https://devfeed.tech/topics/workers.md>)

Tags: [2025](<https://devfeed.tech/tags/2025.md>), [after-the-retrospective](<https://devfeed.tech/tags/after-the-retrospective.md>), [bot-management](<https://devfeed.tech/tags/bot-management.md>), [chaos-engineering](<https://devfeed.tech/tags/chaos-engineering.md>), [cloudflare](<https://devfeed.tech/tags/cloudflare.md>), [configuration](<https://devfeed.tech/tags/configuration.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [http](<https://devfeed.tech/tags/http.md>), [internet](<https://devfeed.tech/tags/internet.md>), [management](<https://devfeed.tech/tags/management.md>), [outage](<https://devfeed.tech/tags/outage.md>), [outages](<https://devfeed.tech/tags/outages.md>), [post-mortem](<https://devfeed.tech/tags/post-mortem.md>), [workers](<https://devfeed.tech/tags/workers.md>)

### AI overview

The article examines the November 2025 Cloudflare outage, explaining how an oversized Bot Management configuration caused HTTP 5XX errors and cascading failures across dependent services. It highlights configuration propagation, service dependencies, and chaos engineering as reliability considerations.

### Source excerpt

In November 2025, a misconfigured Cloudflare service led to a partial outage. Learn what happened, and what you can do to reduce the impact of similar outages.

## Auto Update Post-Mortem

DevFeed: [Auto Update Post-Mortem](<https://devfeed.tech/articles/auto-update-post-mortem-13424.md>)

Original publisher: [Read original article](<https://zed.dev/blog/auto-update-post-mortem>)

Author: Conrad Irwin

Published: 2025-09-12T00:00:00Z

Content type: opinion

Language: en

Sources: [Zed Industries - Blog](<https://devfeed.tech/sources/zed-industries-blog.md>)

Topics: [Post Mortem](<https://devfeed.tech/topics/post-mortem.md>), [bug](<https://devfeed.tech/topics/bug.md>), [JSON](<https://devfeed.tech/topics/json.md>), [Testing](<https://devfeed.tech/topics/testing.md>), [Slack](<https://devfeed.tech/topics/slack.md>)

Tags: [bug](<https://devfeed.tech/tags/bug.md>), [json](<https://devfeed.tech/tags/json.md>), [post-mortem](<https://devfeed.tech/tags/post-mortem.md>), [slack](<https://devfeed.tech/tags/slack.md>), [testing](<https://devfeed.tech/tags/testing.md>)

### AI overview

Zed explains how a settings-parsing change broke automatic updates in production for versions 0.203.0-0.203.4 and 0.204.0. The article identifies the fallback bug, lists fixed versions and recovery steps, and discusses testing and release-process improvements.

### Source excerpt

How we broke auto-update in production and how we're addressing it.

## Delayed Start Compute Operations - Triggering Event

DevFeed: [Delayed Start Compute Operations - Triggering Event](<https://devfeed.tech/articles/delayed-start-compute-operations-triggering-event-5183.md>)

Original publisher: [Read original article](<https://neon.com/blog/delayed-start-compute-operations-triggering-event>)

Author: Mihai Bojin

Published: 2025-05-30T20:24:56Z

Content type: article

Language: en

Sources: [Blog -- Neon Docs](<https://devfeed.tech/sources/blog-neon-docs.md>)

Topics: [Post Mortem](<https://devfeed.tech/topics/post-mortem.md>), [Databases](<https://devfeed.tech/topics/databases.md>), [SQL](<https://devfeed.tech/topics/sql.md>), [Serverless](<https://devfeed.tech/topics/serverless.md>), [Structured-data](<https://devfeed.tech/topics/structured-data.md>)

Tags: [2025](<https://devfeed.tech/tags/2025.md>), [database](<https://devfeed.tech/tags/database.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [golang](<https://devfeed.tech/tags/golang.md>), [ip](<https://devfeed.tech/tags/ip.md>), [monitor](<https://devfeed.tech/tags/monitor.md>), [outage](<https://devfeed.tech/tags/outage.md>), [performance](<https://devfeed.tech/tags/performance.md>), [post-mortem](<https://devfeed.tech/tags/post-mortem.md>), [postgres](<https://devfeed.tech/tags/postgres.md>), [scale](<https://devfeed.tech/tags/scale.md>), [serverless](<https://devfeed.tech/tags/serverless.md>), [serverless-architecture](<https://devfeed.tech/tags/serverless-architecture.md>), [sql](<https://devfeed.tech/tags/sql.md>)

### AI overview

This post-mortem summary explains how a Postgres query plan change in Neon's Activity Monitor caused database CPU saturation. The resulting slowdown prevented the control plane from suspending idle Compute VMs, increasing concurrent VMs and ultimately causing IP exhaustion during the 2025-05-16 outage.

### Source excerpt

For further details, read the top-level Post-Mortem. Summary The Neon Control Plane service is backed by a Postgres database. A scheduled job in the Control plane, Activity Monitor, is responsible for identifying Computes that are ready to be suspended. A Postgres query executed...

## Our simple-to-use incident post-mortem template

DevFeed: [Our simple-to-use incident post-mortem template](<https://devfeed.tech/articles/our-simple-to-use-incident-post-mortem-template-11833.md>)

Original publisher: [Read original article](<https://incident.io/blog/incident-post-mortem-template>)

Author: Chris Evans

Published: 2025-04-16T19:03:00Z

Content type: article

Language: en

Sources: [The incident.io Blog](<https://devfeed.tech/sources/the-incident-io-blog.md>)

Topics: [Post Mortem](<https://devfeed.tech/topics/post-mortem.md>), [Template](<https://devfeed.tech/topics/template.md>), [incident](<https://devfeed.tech/topics/incident.md>), [Incident response](<https://devfeed.tech/topics/incident-response.md>), [Vulnerabilities](<https://devfeed.tech/topics/vulnerabilities.md>), [data](<https://devfeed.tech/topics/data.md>)

Tags: [article](<https://devfeed.tech/tags/article.md>), [data](<https://devfeed.tech/tags/data.md>), [incident](<https://devfeed.tech/tags/incident.md>), [incident-channel](<https://devfeed.tech/tags/incident-channel.md>), [incident-management](<https://devfeed.tech/tags/incident-management.md>), [incident-response](<https://devfeed.tech/tags/incident-response.md>), [metrics](<https://devfeed.tech/tags/metrics.md>), [outage](<https://devfeed.tech/tags/outage.md>), [post-mortem](<https://devfeed.tech/tags/post-mortem.md>), [slack-incident](<https://devfeed.tech/tags/slack-incident.md>), [vulnerabilities](<https://devfeed.tech/tags/vulnerabilities.md>)

### AI overview

This article presents an incident post-mortem template for documenting technical and other incidents. It explains how to combine incident data, human interpretation, metrics, and concrete follow-up actions to understand causes, learn from the response, and prevent similar incidents.

### Source excerpt

Incident post-mortems are a crucial document that cannot be glossed over. In this article, you'll find our go-to post-mortem template that you can use in your own organization.

## Checkpoint - March 2025

DevFeed: [Checkpoint - March 2025](<https://devfeed.tech/articles/checkpoint-march-2025-17146.md>)

Original publisher: [Read original article](<https://blog.ethereum.org/en/2025/03/25/acdcheckpoint-001>)

Author: Nixo Rokish

Published: 2025-03-25T00:00:00Z

Content type: news

Language: en

Sources: [Ethereum Foundation Blog](<https://devfeed.tech/sources/ethereum-foundation-blog.md>)

Topics: [Ethereum](<https://devfeed.tech/topics/ethereum.md>), [Development](<https://devfeed.tech/topics/development.md>), [incident](<https://devfeed.tech/topics/incident.md>), [Post Mortem](<https://devfeed.tech/topics/post-mortem.md>), [configuration](<https://devfeed.tech/topics/configuration.md>), [Network](<https://devfeed.tech/topics/network.md>)

Tags: [configuration](<https://devfeed.tech/tags/configuration.md>), [ethereum](<https://devfeed.tech/tags/ethereum.md>), [feature](<https://devfeed.tech/tags/feature.md>), [features](<https://devfeed.tech/tags/features.md>), [incident](<https://devfeed.tech/tags/incident.md>), [network](<https://devfeed.tech/tags/network.md>), [post-mortem](<https://devfeed.tech/tags/post-mortem.md>), [research-development](<https://devfeed.tech/tags/research-development.md>)

### AI overview

This checkpoint summarizes Ethereum core developers' recent work on the Pectra testnet upgrades and preparation for the Pectra mainnet fork. It covers configuration issues on the Holešky and Sepolia testnets, the launch of Hoodi for further testing, and conditions for selecting a mainnet activation date.

### Source excerpt

Ethereum's weekly All Core Developer calls are a lot to keep up with, so this "Checkpoint" series aims for brief high-level updates with a target cadence of every 4-5 calls, depending on what's happening in core development. See the initial update here. All subsequent updates will be hosted here on...

## Where does the time go after you resolve an incident?

DevFeed: [Where does the time go after you resolve an incident?](<https://devfeed.tech/articles/where-does-the-time-go-after-you-resolve-an-incident-12077.md>)

Original publisher: [Read original article](<https://incident.io/blog/where-does-the-time-go-after-you-resolve-an-incident>)

Author: Eryn Carman

Published: 2024-07-29T20:51:00Z

Content type: article

Language: en

Sources: [The incident.io Blog](<https://devfeed.tech/sources/the-incident-io-blog.md>)

Topics: [incident](<https://devfeed.tech/topics/incident.md>), [Incident response](<https://devfeed.tech/topics/incident-response.md>), [Post Mortem](<https://devfeed.tech/topics/post-mortem.md>)

Tags: [incident](<https://devfeed.tech/tags/incident.md>), [incident-channel](<https://devfeed.tech/tags/incident-channel.md>), [incident-management](<https://devfeed.tech/tags/incident-management.md>), [incident-response](<https://devfeed.tech/tags/incident-response.md>), [outage](<https://devfeed.tech/tags/outage.md>), [post-mortem](<https://devfeed.tech/tags/post-mortem.md>), [slack-incident](<https://devfeed.tech/tags/slack-incident.md>), [technical](<https://devfeed.tech/tags/technical.md>)

### AI overview

An analysis of 13,000 incidents and 14,000 follow-up action items finds that the median time to complete post-incident follow-ups is seven days. Smaller companies tend to finish faster, while medium-sized companies are slowest, largely because of coordination and process challenges.

### Source excerpt

We were curious: where does the time go after an incident is resolved? To find out, we analyzed the post-incident process of 13,000 incidents and 14,000 follow-ups action items.

## Running better incidents from start to finish with Viktor Stanchev of Anchorage Digital

DevFeed: [Running better incidents from start to finish with Viktor Stanchev of Anchorage Digital](<https://devfeed.tech/articles/running-better-incidents-from-start-to-finish-with-viktor-stanchev-of-anchorage-digital-12028.md>)

Original publisher: [Read original article](<https://incident.io/blog/the-debrief-episode-twenty-two>)

Published: 2024-04-29T20:30:37Z

Content type: article

Language: en

Sources: [The incident.io Blog](<https://devfeed.tech/sources/the-incident-io-blog.md>)

Topics: [Incident response](<https://devfeed.tech/topics/incident-response.md>), [incident](<https://devfeed.tech/topics/incident.md>), [incident management](<https://devfeed.tech/topics/incident-management.md>), [Post Mortem](<https://devfeed.tech/topics/post-mortem.md>)

Tags: [incident](<https://devfeed.tech/tags/incident.md>), [incident-channel](<https://devfeed.tech/tags/incident-channel.md>), [incident-management](<https://devfeed.tech/tags/incident-management.md>), [incident-response](<https://devfeed.tech/tags/incident-response.md>), [outage](<https://devfeed.tech/tags/outage.md>), [post-mortem](<https://devfeed.tech/tags/post-mortem.md>), [slack-incident](<https://devfeed.tech/tags/slack-incident.md>)

### AI overview

An episode of The Debrief features Viktor Stanchev of Anchorage Digital discussing practical ways to improve incident response. The conversation covers incident declaration, communication during incidents, and post-mortems, with advice intended for both experienced and newer teams.

### Source excerpt

In this episode of The Debrief, we sit down with Viktor Stanchev of Anchorage Digital to get some actionable advice for running better incidents.

## Post-incident flows: Bringing consistency to your post incident processes

DevFeed: [Post-incident flows: Bringing consistency to your post incident processes](<https://devfeed.tech/articles/post-incident-flows-bringing-consistency-to-your-post-incident-processes-11856.md>)

Original publisher: [Read original article](<https://incident.io/blog/introducing-learning-flows>)

Author: Luis Gonzalez

Published: 2023-10-16T12:51:41Z

Content type: article

Language: en

Sources: [The incident.io Blog](<https://devfeed.tech/sources/the-incident-io-blog.md>)

Topics: [incident](<https://devfeed.tech/topics/incident.md>), [Incident response](<https://devfeed.tech/topics/incident-response.md>), [Post Mortem](<https://devfeed.tech/topics/post-mortem.md>), [Resilience](<https://devfeed.tech/topics/resilience.md>)

Tags: [bug](<https://devfeed.tech/tags/bug.md>), [building](<https://devfeed.tech/tags/building.md>), [incident](<https://devfeed.tech/tags/incident.md>), [incident-channel](<https://devfeed.tech/tags/incident-channel.md>), [incident-management](<https://devfeed.tech/tags/incident-management.md>), [incident-response](<https://devfeed.tech/tags/incident-response.md>), [learning](<https://devfeed.tech/tags/learning.md>), [outage](<https://devfeed.tech/tags/outage.md>), [post](<https://devfeed.tech/tags/post.md>), [post-mortem](<https://devfeed.tech/tags/post-mortem.md>), [resilience](<https://devfeed.tech/tags/resilience.md>), [slack-incident](<https://devfeed.tech/tags/slack-incident.md>), [writing](<https://devfeed.tech/tags/writing.md>)

### AI overview

The article introduces Post-incident Flows, a configurable checklist-based process for standardizing post-incident work. Teams can define tasks such as documenting timelines and summaries, tagging incidents, writing post-mortems, scheduling debriefs, and creating follow-up actions. The goal is to improve learning, communication, and resilience after incidents.

### Source excerpt

With Post-incident flows, you can create a checklist of tasks and actions to ensure consistency across your teams and facilitate deeper learning.

## A guide to post-mortem meetings and how we run them at incident.io

DevFeed: [A guide to post-mortem meetings and how we run them at incident.io](<https://devfeed.tech/articles/a-guide-to-post-mortem-meetings-and-how-we-run-them-at-incident-io-11571.md>)

Original publisher: [Read original article](<https://incident.io/blog/a-guide-to-post-mortem-meetings>)

Author: incident.io

Published: 2023-10-11T13:45:05Z

Content type: tutorial

Language: en

Sources: [The incident.io Blog](<https://devfeed.tech/sources/the-incident-io-blog.md>)

Topics: [Post Mortem](<https://devfeed.tech/topics/post-mortem.md>), [incident](<https://devfeed.tech/topics/incident.md>), [Incident response](<https://devfeed.tech/topics/incident-response.md>), [Learning](<https://devfeed.tech/topics/learning.md>)

Tags: [best-practices](<https://devfeed.tech/tags/best-practices.md>), [culture](<https://devfeed.tech/tags/culture.md>), [incident](<https://devfeed.tech/tags/incident.md>), [incident-channel](<https://devfeed.tech/tags/incident-channel.md>), [incident-management](<https://devfeed.tech/tags/incident-management.md>), [incident-response](<https://devfeed.tech/tags/incident-response.md>), [learning](<https://devfeed.tech/tags/learning.md>), [outage](<https://devfeed.tech/tags/outage.md>), [post-mortem](<https://devfeed.tech/tags/post-mortem.md>), [slack-incident](<https://devfeed.tech/tags/slack-incident.md>)

### AI overview

This guide explains how to prepare for, run, and follow up on post-mortem meetings after incidents. It emphasizes structured discussion, blamelessness, open participation, and addressing contributing factors and impacts to support long-term improvement.

### Source excerpt

Post-mortem meetings can play a crucial role in fostering an environment of continuous learning. Here's how we do them!

## Whose fault was it anyway? On blameless post-mortems

DevFeed: [Whose fault was it anyway? On blameless post-mortems](<https://devfeed.tech/articles/whose-fault-was-it-anyway-on-blameless-post-mortems-11688.md>)

Original publisher: [Read original article](<https://incident.io/blog/blameless-post-mortems>)

Author: incident.io

Published: 2023-10-04T19:54:46Z

Content type: article

Language: en

Sources: [The incident.io Blog](<https://devfeed.tech/sources/the-incident-io-blog.md>)

Topics: [Post Mortem](<https://devfeed.tech/topics/post-mortem.md>), [incident](<https://devfeed.tech/topics/incident.md>), [incident management](<https://devfeed.tech/topics/incident-management.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>)

Tags: [analysis](<https://devfeed.tech/tags/analysis.md>), [incident](<https://devfeed.tech/tags/incident.md>), [incident-channel](<https://devfeed.tech/tags/incident-channel.md>), [incident-management](<https://devfeed.tech/tags/incident-management.md>), [incident-response](<https://devfeed.tech/tags/incident-response.md>), [learning](<https://devfeed.tech/tags/learning.md>), [management](<https://devfeed.tech/tags/management.md>), [outage](<https://devfeed.tech/tags/outage.md>), [post](<https://devfeed.tech/tags/post.md>), [post-mortem](<https://devfeed.tech/tags/post-mortem.md>), [production](<https://devfeed.tech/tags/production.md>), [root-cause-analysis](<https://devfeed.tech/tags/root-cause-analysis.md>), [safety](<https://devfeed.tech/tags/safety.md>), [slack-incident](<https://devfeed.tech/tags/slack-incident.md>)

### AI overview

The article examines blameless post-mortems, arguing that incident analysis should avoid scapegoating while still preserving accountability and enabling meaningful learning. It discusses psychological safety, the limits of focusing on a single root cause, and the risk that AI-generated summaries may remove human accountability entirely.

### Source excerpt

While blameless post-mortems are a great idea on the surface, if taken to the extreme, they can muddy how much you actually learn from incidents.

## Better learning from incidents: A guide to incident post-mortem documents

DevFeed: [Better learning from incidents: A guide to incident post-mortem documents](<https://devfeed.tech/articles/better-learning-from-incidents-a-guide-to-incident-post-mortem-documents-12069.md>)

Original publisher: [Read original article](<https://incident.io/blog/what-is-a-post-mortem-document>)

Author: Luis Gonzalez

Published: 2023-09-27T18:45:18Z

Content type: article

Language: en

Sources: [The incident.io Blog](<https://devfeed.tech/sources/the-incident-io-blog.md>)

Topics: [Post Mortem](<https://devfeed.tech/topics/post-mortem.md>), [incident](<https://devfeed.tech/topics/incident.md>), [Incident response](<https://devfeed.tech/topics/incident-response.md>)

Tags: [guide](<https://devfeed.tech/tags/guide.md>), [incident](<https://devfeed.tech/tags/incident.md>), [incident-channel](<https://devfeed.tech/tags/incident-channel.md>), [incident-management](<https://devfeed.tech/tags/incident-management.md>), [incident-response](<https://devfeed.tech/tags/incident-response.md>), [learning](<https://devfeed.tech/tags/learning.md>), [outage](<https://devfeed.tech/tags/outage.md>), [post-mortem](<https://devfeed.tech/tags/post-mortem.md>), [report](<https://devfeed.tech/tags/report.md>), [retrospective](<https://devfeed.tech/tags/retrospective.md>), [slack-incident](<https://devfeed.tech/tags/slack-incident.md>)

### AI overview

This guide explains post-mortem documents as a way to capture information after an incident, understand contributing factors and risks, and plan actions to prevent or reduce the impact of similar incidents. It distinguishes documents from post-mortem meetings, discusses disagreement about when they are appropriate, and describes blameless post-mortems as an approach that avoids scapegoating. The article also notes alternative terms such as incident debrief documents, Incident Retrospectives, and After Action Reports.

### Source excerpt

Post-mortem documents are a great way to facilitate learning after incidents are resolved.

## Reflecting on one of the biggest incidents in our history

DevFeed: [Reflecting on one of the biggest incidents in our history](<https://devfeed.tech/articles/reflecting-on-one-of-the-biggest-incidents-in-our-history-11865.md>)

Original publisher: [Read original article](<https://incident.io/blog/kubecon-post-mortem>)

Author: Luis Gonzalez

Published: 2023-05-01T00:00:00Z

Content type: article

Language: en

Sources: [The incident.io Blog](<https://devfeed.tech/sources/the-incident-io-blog.md>)

Topics: [incident](<https://devfeed.tech/topics/incident.md>), [Post Mortem](<https://devfeed.tech/topics/post-mortem.md>)

Tags: [conference](<https://devfeed.tech/tags/conference.md>), [incident](<https://devfeed.tech/tags/incident.md>), [incident-channel](<https://devfeed.tech/tags/incident-channel.md>), [incident-management](<https://devfeed.tech/tags/incident-management.md>), [incident-response](<https://devfeed.tech/tags/incident-response.md>), [kubecon](<https://devfeed.tech/tags/kubecon.md>), [outage](<https://devfeed.tech/tags/outage.md>), [post-mortem](<https://devfeed.tech/tags/post-mortem.md>), [recognition](<https://devfeed.tech/tags/recognition.md>), [slack-incident](<https://devfeed.tech/tags/slack-incident.md>), [storage](<https://devfeed.tech/tags/storage.md>)

### AI overview

The article humorously presents running out of T-shirts during KubeCon as a major incident. It recounts how unexpectedly high demand, amplified by recognition of the shirts as standout conference swag, quickly exhausted the team's supply.

### Source excerpt

During KubeCon, we experienced a significant incident and rallied to resolve it. Here's the story behind it.

## Keep the monolith, but split the workloads

DevFeed: [Keep the monolith, but split the workloads](<https://devfeed.tech/articles/keep-the-monolith-but-split-the-workloads-11877.md>)

Original publisher: [Read original article](<https://incident.io/blog/monolith>)

Author: Lawrence Jones

Published: 2023-04-12T00:00:00Z

Content type: article

Language: en

Sources: [The incident.io Blog](<https://devfeed.tech/sources/the-incident-io-blog.md>)

Topics: [Architecture & Design](<https://devfeed.tech/topics/architecture-design.md>), [Scalability](<https://devfeed.tech/topics/scalability.md>), [Microservices](<https://devfeed.tech/topics/microservices.md>), [Deployment](<https://devfeed.tech/topics/deployment.md>), [incident](<https://devfeed.tech/topics/incident.md>), [Heroku](<https://devfeed.tech/topics/heroku.md>), [Post Mortem](<https://devfeed.tech/topics/post-mortem.md>), [Ruby](<https://devfeed.tech/topics/ruby.md>), [Database](<https://devfeed.tech/topics/database.md>)

Tags: [architecture](<https://devfeed.tech/tags/architecture.md>), [deployment](<https://devfeed.tech/tags/deployment.md>), [incident](<https://devfeed.tech/tags/incident.md>), [incident-channel](<https://devfeed.tech/tags/incident-channel.md>), [incident-management](<https://devfeed.tech/tags/incident-management.md>), [incident-response](<https://devfeed.tech/tags/incident-response.md>), [microservices](<https://devfeed.tech/tags/microservices.md>), [observability](<https://devfeed.tech/tags/observability.md>), [outage](<https://devfeed.tech/tags/outage.md>), [post-mortem](<https://devfeed.tech/tags/post-mortem.md>), [ruby](<https://devfeed.tech/tags/ruby.md>), [scalability](<https://devfeed.tech/tags/scalability.md>), [slack-incident](<https://devfeed.tech/tags/slack-incident.md>), [systems](<https://devfeed.tech/tags/systems.md>)

### AI overview

The article argues that teams can preserve the simplicity of a monolith while addressing scaling problems by splitting workloads into separate deployment tiers. Drawing on an outage caused by repeated application crashes, it contrasts this approach with microservices, which can improve reliability and scalability but introduce distributed-system and operational complexity. The described implementation used three independent Heroku dyno tiers, each handling a distinct workload.

### Source excerpt

Everybody loves a monolith, but you can hit issues as you scale. Learn how splitting workloads can improve your monolithic architecture's performance and scalability, and understand the trade-offs between monolithic systems and microservices.

[Next page](<https://devfeed.tech/topics/post-mortem.md?cursor=WyIyMDIzLTA0LTEyVDAwOjAwOjAwKzAwOjAwIiwgIjAwNTM5YjE3LWE2YjQtNDg2ZS04ZWFlLTcyMWM4ZmFkMTM3OSJd>)