# outages

Published articles for outages.

This is one page of public article previews, not the complete archive. Follow Next page to continue. Summaries are not the original full articles.

## How Marathon Incidents Expose Organizational Fragility and Require Structured Incident Response

DevFeed: [How Marathon Incidents Expose Organizational Fragility and Require Structured Incident Response](<https://devfeed.tech/articles/presentation-when-incidents-refuse-to-end-41300.md>)

Original publisher: [Read original article](<https://www.infoq.com/presentations/stream-incidents/>)

Author: Vanessa Huerta Granda

Published: 2026-09-17T09:30:00Z

Content type: article

Language: en

Sources: [InfoQ](<https://devfeed.tech/sources/infoq.md>)

Topics: [incident](<https://devfeed.tech/topics/incident.md>), [Incident response](<https://devfeed.tech/topics/incident-response.md>)

Tags: [devops](<https://devfeed.tech/tags/devops.md>), [incident](<https://devfeed.tech/tags/incident.md>), [incident-response](<https://devfeed.tech/tags/incident-response.md>), [infoq](<https://devfeed.tech/tags/infoq.md>), [limits](<https://devfeed.tech/tags/limits.md>), [organizational](<https://devfeed.tech/tags/organizational.md>), [outages](<https://devfeed.tech/tags/outages.md>), [presentation](<https://devfeed.tech/tags/presentation.md>), [qcon-san-francisco-2026](<https://devfeed.tech/tags/qcon-san-francisco-2026.md>), [qcon-software-development-conference](<https://devfeed.tech/tags/qcon-software-development-conference.md>), [real-world](<https://devfeed.tech/tags/real-world.md>), [stream-incidents](<https://devfeed.tech/tags/stream-incidents.md>), [structured](<https://devfeed.tech/tags/structured.md>), [system](<https://devfeed.tech/tags/system.md>), [transcripts](<https://devfeed.tech/tags/transcripts.md>)

### AI overview

Vanessa Huerta Granda discusses how prolonged incidents reveal gaps between planned work and work as performed. Drawing on real-world scenarios, the presentation examines organizational fragility, human limits, system interdependencies, and the need for structured, humane, cross-functional incident response.

### Source excerpt

Vanessa Huerta Granda explains how marathon incidents expose the gap between work as imagined and work as done. Drawing from real-world scenarios, she shares how complex outages reveal organizational fragility, human limits, and system interdependencies--and why incident response requires structured endurance, humane rotations, and holistic cross-functional coordination. By Vanessa Huerta Granda

## Solana: Building, Proving and Earning Trust in Public

DevFeed: [Solana: Building, Proving and Earning Trust in Public](<https://devfeed.tech/articles/solana-building-proving-and-earning-trust-in-public-17467.md>)

Original publisher: [Read original article](<https://solana.com/news/solana-building-trust-in-public>)

Author: Jacob Creech

Published: 2026-09-14T11:00:00Z

Content type: article

Language: en

Sources: [Solana News Feed](<https://devfeed.tech/sources/solana-news-feed.md>)

Topics: [Network](<https://devfeed.tech/topics/network.md>), [Critical Infrastructure](<https://devfeed.tech/topics/critical-infrastructure.md>), [incident](<https://devfeed.tech/topics/incident.md>), [networking](<https://devfeed.tech/topics/networking.md>), [Transactions](<https://devfeed.tech/topics/transactions.md>), [Amazon Web Services](<https://devfeed.tech/topics/aws.md>)

Tags: [2026](<https://devfeed.tech/tags/2026.md>), [aws](<https://devfeed.tech/tags/aws.md>), [blockchain](<https://devfeed.tech/tags/blockchain.md>), [blockchain-technology](<https://devfeed.tech/tags/blockchain-technology.md>), [building](<https://devfeed.tech/tags/building.md>), [critical-infrastructure](<https://devfeed.tech/tags/critical-infrastructure.md>), [crypto](<https://devfeed.tech/tags/crypto.md>), [crypto-news](<https://devfeed.tech/tags/crypto-news.md>), [cryptocurrency](<https://devfeed.tech/tags/cryptocurrency.md>), [defi](<https://devfeed.tech/tags/defi.md>), [ecosystem](<https://devfeed.tech/tags/ecosystem.md>), [incident](<https://devfeed.tech/tags/incident.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [network](<https://devfeed.tech/tags/network.md>), [nfts](<https://devfeed.tech/tags/nfts.md>), [outages](<https://devfeed.tech/tags/outages.md>), [podcasts](<https://devfeed.tech/tags/podcasts.md>), [routing](<https://devfeed.tech/tags/routing.md>), [solana](<https://devfeed.tech/tags/solana.md>), [solana-ecosystem](<https://devfeed.tech/tags/solana-ecosystem.md>), [technology](<https://devfeed.tech/tags/technology.md>), [transactions](<https://devfeed.tech/tags/transactions.md>), [uptime](<https://devfeed.tech/tags/uptime.md>), [web3](<https://devfeed.tech/tags/web3.md>)

### AI overview

Solana describes how repeated public testing, transparent incident reporting, and corrective engineering have strengthened trust in the network. It highlights a 2026 routing failure that took nearly 29% of network stake offline while blocks and transactions continued, followed by recovery of the infrastructure provider in just over 30 minutes. The article also cites validator capacity improvements, stake-weighted quality of service, and a redesigned transaction scheduler as examples of resilience work.

### Source excerpt

Solana has maintained 100% uptime since February 2024, including when a routing failure took nearly 29% of network stake offline.

## Proxmox Enterprise Support Goes 24/7 on October 19; North American Subsidiary Opens in Kingston

DevFeed: [Proxmox Enterprise Support Goes 24/7 on October 19; North American Subsidiary Opens in Kingston](<https://devfeed.tech/articles/proxmox-enterprise-support-goes-24-7-on-october-19-north-american-subsidiary-opens-in-kingston-12373.md>)

Original publisher: [Read original article](<https://www.storagereview.com/news/proxmox-enterprise-support-goes-24-7-on-october-19-north-american-subsidiary-opens-in-kingston>)

Author: Lyle Smith

Published: 2026-09-03T14:57:12Z

Content type: news

Language: en

Sources: [StorageReview.com](<https://devfeed.tech/sources/storagereview-com.md>)

Topics: [Proxmox](<https://devfeed.tech/topics/proxmox.md>), [virtualization](<https://devfeed.tech/topics/virtualization.md>), [Cloud](<https://devfeed.tech/topics/cloud.md>)

Tags: [2026](<https://devfeed.tech/tags/2026.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [customers](<https://devfeed.tech/tags/customers.md>), [enterprise](<https://devfeed.tech/tags/enterprise.md>), [outages](<https://devfeed.tech/tags/outages.md>), [production](<https://devfeed.tech/tags/production.md>), [proxmox](<https://devfeed.tech/tags/proxmox.md>), [software](<https://devfeed.tech/tags/software.md>), [ssh](<https://devfeed.tech/tags/ssh.md>), [support](<https://devfeed.tech/tags/support.md>)

### AI overview

Proxmox will begin offering 24/7 global enterprise support on October 19, 2026, covering Proxmox Virtual Environment, Proxmox Backup Server, and Proxmox Datacenter Manager. It is also establishing Proxmox North America Inc. in Kingston, Ontario, to support customers and partners in the United States and Canada. Premium subscribers will receive around-the-clock coverage, while Standard subscribers are scheduled for a later Q4 2026 onboarding window.

### Source excerpt

Proxmox is making a major expansion to its enterprise operations, moving to 24/7 global support beginning October 19, 2026, while establishing a new North American subsidiary in Kingston, Ontario. The new support covers Proxmox Virtual Environment, Proxmox Backup Server, and Proxmox Datacenter Manager, giving enterprise customers direct access to Proxmox engineers at any time of The post Proxmox Enterprise Support Goes 24/7 on October 19; North American Subsidiary Opens in Kingston appeared first on StorageReview.com.

## The invisible heartbeat of our networks

DevFeed: [The invisible heartbeat of our networks](<https://devfeed.tech/articles/the-invisible-heartbeat-of-our-networks-10855.md>)

Original publisher: [Read original article](<https://blog.apnic.net/2026/09/01/the-invisible-heartbeat-of-our-networks/>)

Author: Luca Cicchelli

Published: 2026-08-31T23:03:47Z

Content type: article

Language: en

Sources: [APNIC Blog](<https://devfeed.tech/sources/apnic-blog.md>)

Topics: [Networks](<https://devfeed.tech/topics/networks.md>), [distributed-systems](<https://devfeed.tech/topics/distributed-systems.md>), [Critical Infrastructure](<https://devfeed.tech/topics/critical-infrastructure.md>), [Security](<https://devfeed.tech/topics/security.md>), [5G](<https://devfeed.tech/topics/5g.md>), [Cybersecurity](<https://devfeed.tech/topics/cybersecurity.md>), [cloud-infrastructure](<https://devfeed.tech/topics/cloud-infrastructure.md>), [AI Infrastructure](<https://devfeed.tech/topics/ai-infrastructure.md>)

Tags: [5g](<https://devfeed.tech/tags/5g.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-infrastructure](<https://devfeed.tech/tags/ai-infrastructure.md>), [critical-infrastructure](<https://devfeed.tech/tags/critical-infrastructure.md>), [cybersecurity](<https://devfeed.tech/tags/cybersecurity.md>), [data](<https://devfeed.tech/tags/data.md>), [distributed-systems](<https://devfeed.tech/tags/distributed-systems.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [gnss](<https://devfeed.tech/tags/gnss.md>), [gps](<https://devfeed.tech/tags/gps.md>), [guest-post](<https://devfeed.tech/tags/guest-post.md>), [hardware](<https://devfeed.tech/tags/hardware.md>), [network](<https://devfeed.tech/tags/network.md>), [networking](<https://devfeed.tech/tags/networking.md>), [networks](<https://devfeed.tech/tags/networks.md>), [outage](<https://devfeed.tech/tags/outage.md>), [outages](<https://devfeed.tech/tags/outages.md>), [post](<https://devfeed.tech/tags/post.md>), [servers](<https://devfeed.tech/tags/servers.md>), [systems](<https://devfeed.tech/tags/systems.md>), [tech-matters](<https://devfeed.tech/tags/tech-matters.md>), [time](<https://devfeed.tech/tags/time.md>), [transactions](<https://devfeed.tech/tags/transactions.md>)

### AI overview

The article examines how accurate time synchronization underpins modern networks and distributed systems. It uses the Telstra outage to show how misaligned time servers can disrupt communications and rail services, and discusses resilient time sources, terrestrial backups, local atomic clocks, and synchronization requirements for finance, 5G, cloud, and AI infrastructure.

### Source excerpt

Guest Post: The recent Telstra outage highlighted a critical but often overlooked dependency in modern networks: Accurate time synchronization.

## How MongoDB Atlas and Temporal support reliable production RAG and AI agents

DevFeed: [How MongoDB Atlas and Temporal support reliable production RAG and AI agents](<https://devfeed.tech/articles/durable-rag-and-agents-mongodb-and-temporal-doing-it-better-together-35920.md>)

Original publisher: [Read original article](<https://temporal.io/blog/mongodb-temporal-partnership-rag-agents>)

Author: Suresh Ramappa

Published: 2026-08-13T00:00:00Z

Content type: opinion

Language: en

Sources: [Temporal Blog](<https://devfeed.tech/sources/temporal-blog.md>)

Topics: [MongoDB](<https://devfeed.tech/topics/mongodb.md>), [Retrieval Augmented Generation (RAG)](<https://devfeed.tech/topics/retrieval-augmented-generation-rag.md>), [AI Infrastructure](<https://devfeed.tech/topics/ai-infrastructure.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [reliability](<https://devfeed.tech/topics/reliability.md>)

Tags: [agentic](<https://devfeed.tech/tags/agentic.md>), [agents](<https://devfeed.tech/tags/agents.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-infrastructure](<https://devfeed.tech/tags/ai-infrastructure.md>), [api](<https://devfeed.tech/tags/api.md>), [data](<https://devfeed.tech/tags/data.md>), [embeddings](<https://devfeed.tech/tags/embeddings.md>), [long-running](<https://devfeed.tech/tags/long-running.md>), [mongodb](<https://devfeed.tech/tags/mongodb.md>), [outages](<https://devfeed.tech/tags/outages.md>), [pipelines](<https://devfeed.tech/tags/pipelines.md>), [production](<https://devfeed.tech/tags/production.md>), [rag](<https://devfeed.tech/tags/rag.md>), [recovery](<https://devfeed.tech/tags/recovery.md>), [reliability](<https://devfeed.tech/tags/reliability.md>), [retries](<https://devfeed.tech/tags/retries.md>), [retrieval](<https://devfeed.tech/tags/retrieval.md>), [temporal-voices](<https://devfeed.tech/tags/temporal-voices.md>)

### AI overview

The article argues that MongoDB Atlas and Temporal address different reliability needs in production RAG and AI agent systems. Atlas provides operational data, embeddings, vector search, and agent memory in one platform, while Temporal provides durable execution for crash recovery, retries, and long-running ingestion and agent workflows.

### Source excerpt

Why MongoDB Atlas and Temporal are better together for AI: one data platform, one durable execution layer, for RAG and agents in prod.

## What your AI SRE can't see (and what you can do about it)

DevFeed: [What your AI SRE can't see (and what you can do about it)](<https://devfeed.tech/articles/what-your-ai-sre-can-t-see-and-what-you-can-do-about-it-11736.md>)

Original publisher: [Read original article](<https://www.gremlin.com/blog/what-your-ai-sre-cant-see-and-what-you-can-do-about-it>)

Author: Ryan Detwiller

Published: 2026-08-06T00:00:00Z

Content type: opinion

Language: en

Sources: [Gremlin Blog](<https://devfeed.tech/sources/gremlin-blog.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [SRE](<https://devfeed.tech/topics/sre.md>), [incident](<https://devfeed.tech/topics/incident.md>), [Availability](<https://devfeed.tech/topics/availability.md>), [telemetry](<https://devfeed.tech/topics/telemetry.md>), [Latency](<https://devfeed.tech/topics/latency.md>), [Memory Leaks](<https://devfeed.tech/topics/memory-leaks.md>), [Deployment](<https://devfeed.tech/topics/deployment.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [availability](<https://devfeed.tech/tags/availability.md>), [config](<https://devfeed.tech/tags/config.md>), [dependency](<https://devfeed.tech/tags/dependency.md>), [deployment](<https://devfeed.tech/tags/deployment.md>), [latency](<https://devfeed.tech/tags/latency.md>), [load-balancer](<https://devfeed.tech/tags/load-balancer.md>), [memory](<https://devfeed.tech/tags/memory.md>), [outage](<https://devfeed.tech/tags/outage.md>), [outages](<https://devfeed.tech/tags/outages.md>), [sre](<https://devfeed.tech/tags/sre.md>)

### AI overview

The article argues that AI SRE tools can speed up triage, reduce alert fatigue, and automate frontline incident response, but they do not solve all reliability problems. It identifies gaps including acting only after failures begin and being unable to predict sudden failures without detectable warning signals.

### Source excerpt

AI SRE is having a moment. And let's be honest: faster triage, less alert fatigue, and automated frontline response are wins for understaffed teams. But there are still five gaps in their capabilities, and if you don't understand those gaps before you deploy, you'll find out during an outage.

## Inside a Shutdown: BGP Evidence from an Operator Who Was There

DevFeed: [Inside a Shutdown: BGP Evidence from an Operator Who Was There](<https://devfeed.tech/articles/inside-a-shutdown-bgp-evidence-from-an-operator-who-was-there-11448.md>)

Original publisher: [Read original article](<https://labs.ripe.net/author/mdkamruzzaman-khan-2/inside-a-shutdown-bgp-evidence-from-an-operator-who-was-there/>)

Author: Md.Kamruzzaman Khan

Published: 2026-07-22T11:29:16Z

Content type: article

Language: en

Sources: [RIPE Labs](<https://devfeed.tech/sources/ripe-labs.md>)

Topics: [BGP](<https://devfeed.tech/topics/bgp.md>), [Routing (disambiguation)](<https://devfeed.tech/topics/routing.md>), [networking](<https://devfeed.tech/topics/networking.md>), [Cloudflare](<https://devfeed.tech/topics/cloudflare.md>), [Internet](<https://devfeed.tech/topics/internet.md>)

Tags: [api](<https://devfeed.tech/tags/api.md>), [article](<https://devfeed.tech/tags/article.md>), [bgp](<https://devfeed.tech/tags/bgp.md>), [cloudflare](<https://devfeed.tech/tags/cloudflare.md>), [competition](<https://devfeed.tech/tags/competition.md>), [internet](<https://devfeed.tech/tags/internet.md>), [network](<https://devfeed.tech/tags/network.md>), [outage](<https://devfeed.tech/tags/outage.md>), [outages](<https://devfeed.tech/tags/outages.md>), [radar](<https://devfeed.tech/tags/radar.md>), [research](<https://devfeed.tech/tags/research.md>), [ris](<https://devfeed.tech/tags/ris.md>), [routing](<https://devfeed.tech/tags/routing.md>)

### AI overview

This article examines Bangladesh's July 2024 internet shutdown through public BGP routing data and firsthand observations from one affected network. It compares operator experiences with measurements from RIPE RIS, Cloudflare Radar, IODA, and OONI, while explaining what the data can and cannot establish about the event.

### Source excerpt

When Bangladesh vanished from the global routing table in July 2024, public routing data captured the event as it unfolded. Drawing on first-hand observations from inside one affected network, this article explores what that data can and cannot tell us.

## Characterizing Metastable Faults and Failures

DevFeed: [Characterizing Metastable Faults and Failures](<https://devfeed.tech/articles/characterizing-metastable-faults-and-failures-41850.md>)

Original publisher: [Read original article](<https://muratbuffalo.blogspot.com/2026/07/characterizing-metastable-faults-and.html>)

Author: Murat (noreply@blogger.com)

Published: 2026-07-22T06:21:52Z

Content type: opinion

Language: en

Sources: [Metadata](<https://devfeed.tech/sources/metadata.md>)

Topics: [distributed-systems](<https://devfeed.tech/topics/distributed-systems.md>), [systems](<https://devfeed.tech/topics/systems.md>), [scheduling](<https://devfeed.tech/topics/scheduling.md>)

Tags: [distributed-systems](<https://devfeed.tech/tags/distributed-systems.md>), [fault-tolerance](<https://devfeed.tech/tags/fault-tolerance.md>), [metastability](<https://devfeed.tech/tags/metastability.md>), [outages](<https://devfeed.tech/tags/outages.md>), [paper-review](<https://devfeed.tech/tags/paper-review.md>), [queues](<https://devfeed.tech/tags/queues.md>), [retries](<https://devfeed.tech/tags/retries.md>), [scheduling](<https://devfeed.tech/tags/scheduling.md>), [stabilization](<https://devfeed.tech/tags/stabilization.md>)

### AI overview

A review of a June 2026 paper that develops an analytical account of metastable faults and failures in distributed systems. The paper connects these failures to composition among self-stabilizing systems and scheduling, while the review questions its treatment of prior work and its use of local pairwise reasoning.

### Source excerpt

Metastability has been studied in previous work as a self-sustaining degradation in goodput that persists even after the trigger is gone. The degraded state loiters on entirely due to the system's own internal feedback loops (retries, queues), and there is no simple reset button to press in distributed systems. So this is not a rare exotic problem. Since production systems would have already been hardened to handle the obvious failures, what remains is these hard-to-detect emergent failures. The "Metastable Failures in the Wild" paper (OSDI'22) reports 22 incidents across 11 organizations. Four of the 15 major AWS outages in a decade were metastable failures, with durations ranging from 1.5 to 73 hours. This paper (June 2026) argues that the systems community has treated metastability phenomenologically, which led people to chase symptoms rather than causes. The paper sets out to give the first analytical causal account of these failures. This framing leads to two connections I found delightful. The first casts a metastable failure as a sin of composition among self-stabilizing systems. The second ties the healing mechanism to scheduling. Since I have worked on self-stabilizing systems for a long time (between 1998-2010), and thought hard about how they compose, these two connections really excite me. I do have some reservations, though, which I will get to in my review. The paper overlooks prior work on the composition of stabilizing systems. It also pulls a bit of a sleight-of-hand in its formalization to argue that the metastable fault tolerance (MFT) design can be done via local pairwise reasoning between components. I don't buy that as I explain below. These reservations do not dampen how much I enjoyed this paper. This is an idea paper, and it has been a while since I saw one of these in distributed systems. Moreover, the author list includes two of my all-time favorite distributed systems researchers: Robbert Van Renesse and Lorenzo Alvisi. So let's dive in.

## Eliminate Reliability Blind Spots in AWS, Azure, and GCP

DevFeed: [Eliminate Reliability Blind Spots in AWS, Azure, and GCP](<https://devfeed.tech/articles/eliminate-reliability-blind-spots-in-aws-azure-and-gcp-11565.md>)

Original publisher: [Read original article](<https://www.gremlin.com/blog/eliminate-reliability-blind-spots-detected-risks-aws-azure-gcp>)

Author: Andre Newman

Published: 2026-07-14T00:00:00Z

Content type: article

Language: en

Sources: [Gremlin Blog](<https://devfeed.tech/sources/gremlin-blog.md>)

Topics: [Cloud](<https://devfeed.tech/topics/cloud.md>), [Kubernetes](<https://devfeed.tech/topics/kubernetes.md>), [Resilience](<https://devfeed.tech/topics/resilience.md>), [Amazon Web Services](<https://devfeed.tech/topics/aws.md>), [Azure](<https://devfeed.tech/topics/azure.md>), [Chaos Engineering](<https://devfeed.tech/topics/chaos-engineering.md>)

Tags: [availability](<https://devfeed.tech/tags/availability.md>), [aws](<https://devfeed.tech/tags/aws.md>), [azure](<https://devfeed.tech/tags/azure.md>), [containers](<https://devfeed.tech/tags/containers.md>), [deployment](<https://devfeed.tech/tags/deployment.md>), [experimentation](<https://devfeed.tech/tags/experimentation.md>), [features](<https://devfeed.tech/tags/features.md>), [gcp](<https://devfeed.tech/tags/gcp.md>), [gremlin](<https://devfeed.tech/tags/gremlin.md>), [kubernetes](<https://devfeed.tech/tags/kubernetes.md>), [orchestration](<https://devfeed.tech/tags/orchestration.md>), [outages](<https://devfeed.tech/tags/outages.md>), [resilience](<https://devfeed.tech/tags/resilience.md>), [uptime](<https://devfeed.tech/tags/uptime.md>)

### AI overview

Gremlin's expanded Detected Risks feature automatically identifies high-priority reliability risks across AWS, Azure, GCP, and Kubernetes environments. It is designed to reveal issues such as misconfigured deployments, crash-looping containers, missing readiness probes, and poor pod distribution before they cause outages.

### Source excerpt

Discover reliability risks without running a single test. See how Gremlin identifies high-priority reliability risks across AWS, Azure, and GCP to prevent outages.

## Network Rack for Small Office (Build Guide)

DevFeed: [Network Rack for Small Office (Build Guide)](<https://devfeed.tech/articles/network-rack-for-small-office-build-guide-20879.md>)

Original publisher: [Read original article](<https://linuxblog.io/network-rack-for-small-office/>)

Author: Hayden James

Published: 2026-06-22T12:58:14Z

Content type: tutorial

Language: en

Sources: [Hayden James](<https://devfeed.tech/sources/hayden-james.md>)

Topics: [Network](<https://devfeed.tech/topics/network.md>), [Hardware](<https://devfeed.tech/topics/hardware.md>), [Firewall](<https://devfeed.tech/topics/firewall.md>)

Tags: [battery](<https://devfeed.tech/tags/battery.md>), [blog](<https://devfeed.tech/tags/blog.md>), [community-forums](<https://devfeed.tech/tags/community-forums.md>), [general](<https://devfeed.tech/tags/general.md>), [guide](<https://devfeed.tech/tags/guide.md>), [hardware](<https://devfeed.tech/tags/hardware.md>), [home-office](<https://devfeed.tech/tags/home-office.md>), [homelab](<https://devfeed.tech/tags/homelab.md>), [network](<https://devfeed.tech/tags/network.md>), [outages](<https://devfeed.tech/tags/outages.md>), [router](<https://devfeed.tech/tags/router.md>), [sysadmins](<https://devfeed.tech/tags/sysadmins.md>), [unifi](<https://devfeed.tech/tags/unifi.md>)

### AI overview

A build guide describes equipping a small professional office with a wheeled 20U network rack, UniFi networking equipment, UPS units, modular patching, cable management, and 10G connectivity. The design emphasizes keeping key services online during power interruptions and allowing future expansion.

### Source excerpt

Last week I walked into an empty 5 x 4 ft storage room. By the end of the week, that blank space held a fully functional network core designed to keep a small professional office online through power cuts, ISP hiccups, and future growth. Continue reading...

## How Coverwatch uses Temporal to orchestrate AI-powered insurance workflows

DevFeed: [How Coverwatch uses Temporal to orchestrate AI-powered insurance workflows](<https://devfeed.tech/articles/how-coverwatch-uses-temporal-to-orchestrate-ai-powered-insurance-workflows-35850.md>)

Original publisher: [Read original article](<https://temporal.io/blog/how-coverwatch-uses-temporal-to-orchestrate-ai-powered-insurance-workflows>)

Author: Wilmer Yan

Published: 2026-06-09T00:00:00Z

Content type: article

Language: en

Sources: [Temporal Blog](<https://devfeed.tech/sources/temporal-blog.md>)

Topics: [Orchestration](<https://devfeed.tech/topics/orchestration.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Automation](<https://devfeed.tech/topics/automation.md>), [human review](<https://devfeed.tech/topics/human-review.md>), [API](<https://devfeed.tech/topics/api.md>)

Tags: [agents](<https://devfeed.tech/tags/agents.md>), [ai](<https://devfeed.tech/tags/ai.md>), [automation](<https://devfeed.tech/tags/automation.md>), [community](<https://devfeed.tech/tags/community.md>), [human-review](<https://devfeed.tech/tags/human-review.md>), [outages](<https://devfeed.tech/tags/outages.md>), [temporal](<https://devfeed.tech/tags/temporal.md>), [workflows](<https://devfeed.tech/tags/workflows.md>)

### AI overview

Coverwatch uses Temporal to coordinate long-running commercial insurance workflows involving AI agents, broker review, carrier communication, and operational follow-up. The platform is designed to preserve workflow state across delays, outages, and external-service issues.

### Source excerpt

Coverwatch runs its AI-powered insurance workflows on Temporal, keeping agents, broker review, and carrier submissions durable through outages.

## Building Resilient Fintech Infrastructure for Scale

DevFeed: [Building Resilient Fintech Infrastructure for Scale](<https://devfeed.tech/articles/why-does-fintech-break-at-scale-build-for-resilience-23785.md>)

Original publisher: [Read original article](<https://cockroachlabs.com/blog/fintech-infrastructure-resilience-at-scale>)

Author: David Weiss

Published: 2026-04-02T00:00:00Z

Content type: article

Language: en

Sources: [Cockroach Labs](<https://devfeed.tech/sources/cockroach-labs.md>)

Topics: [Resilience](<https://devfeed.tech/topics/resilience.md>), [Data Infrastructure](<https://devfeed.tech/topics/data-infrastructure.md>), [Database](<https://devfeed.tech/topics/database.md>), [CockroachDB](<https://devfeed.tech/topics/cockroachdb.md>), [legacy](<https://devfeed.tech/topics/legacy.md>)

Tags: [cockroach-labs](<https://devfeed.tech/tags/cockroach-labs.md>), [cockroachdb](<https://devfeed.tech/tags/cockroachdb.md>), [compliance](<https://devfeed.tech/tags/compliance.md>), [fintech](<https://devfeed.tech/tags/fintech.md>), [fraud-detection](<https://devfeed.tech/tags/fraud-detection.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [legacy](<https://devfeed.tech/tags/legacy.md>), [outages](<https://devfeed.tech/tags/outages.md>), [postgresql](<https://devfeed.tech/tags/postgresql.md>), [resilience](<https://devfeed.tech/tags/resilience.md>), [uptime](<https://devfeed.tech/tags/uptime.md>)

### AI overview

The article argues that fintech and quant firms need resilient data infrastructure to handle growth, real-time payments, instant settlement, AI-driven fraud detection, and cross-border compliance. It describes the limits of legacy database architectures and presents SumUp's migration from PostgreSQL to CockroachDB as an example of improved resilience and near-zero downtime.

### Source excerpt

The fintech companies and quant firms that define the next decade aren't just building better products. They're succeeding with more resilient fintech infrastructure.

## The hidden reliability risks in your agentic AI workflows

DevFeed: [The hidden reliability risks in your agentic AI workflows](<https://devfeed.tech/articles/the-hidden-reliability-risks-in-your-agentic-ai-workflows-11721.md>)

Original publisher: [Read original article](<https://www.gremlin.com/blog/the-hidden-reliability-risks-in-your-agentic-ai-workflows>)

Author: Andre Newman

Published: 2026-03-17T00:00:00Z

Content type: article

Language: en

Sources: [Gremlin Blog](<https://devfeed.tech/sources/gremlin-blog.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [distributed-systems](<https://devfeed.tech/topics/distributed-systems.md>), [Latency](<https://devfeed.tech/topics/latency.md>), [Network](<https://devfeed.tech/topics/network.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [Low Latency](<https://devfeed.tech/topics/low-latency.md>), [Retrieval Augmented Generation (RAG)](<https://devfeed.tech/topics/retrieval-augmented-generation-rag.md>), [systems](<https://devfeed.tech/topics/systems.md>), [Orchestration](<https://devfeed.tech/topics/orchestration.md>), [AWS Transform](<https://devfeed.tech/topics/aws-transform.md>)

Tags: [agentic-ai](<https://devfeed.tech/tags/agentic-ai.md>), [ai](<https://devfeed.tech/tags/ai.md>), [complexity](<https://devfeed.tech/tags/complexity.md>), [dependency](<https://devfeed.tech/tags/dependency.md>), [llm](<https://devfeed.tech/tags/llm.md>), [network](<https://devfeed.tech/tags/network.md>), [orchestration](<https://devfeed.tech/tags/orchestration.md>), [outages](<https://devfeed.tech/tags/outages.md>), [production](<https://devfeed.tech/tags/production.md>), [rag](<https://devfeed.tech/tags/rag.md>), [systems](<https://devfeed.tech/tags/systems.md>), [testing](<https://devfeed.tech/tags/testing.md>), [workflows](<https://devfeed.tech/tags/workflows.md>)

### AI overview

This article examines reliability risks in agentic AI workflows, focusing on unstable network interactions, non-deterministic behavior, and the complexity of third-party dependencies. It explains how latency, outages, retrieval-augmented generation systems, and multi-agent orchestration can create cascading failures, and advocates proactive reliability testing.

### Source excerpt

Prevent AI outages before they impact production. Learn how to discover and mitigate reliability risks in agentic AI systems like network stability, non-deterministic behavior, and third-party dependencies.

## Designing Resilient APIs: Failure-Handling Patterns for Distributed Systems

DevFeed: [Designing Resilient APIs: Failure-Handling Patterns for Distributed Systems](<https://devfeed.tech/articles/designing-resilient-apis-failure-handling-patterns-for-distributed-systems-39563.md>)

Original publisher: [Read original article](<https://ankit-rana.com/logs/11-resilient-api-design-patterns/>)

Author: hello@ankit-rana.com

Published: 2026-03-16T00:00:00Z

Content type: tutorial

Language: en

Sources: [Ankit Rana | Mechanical Sympathy](<https://devfeed.tech/sources/ankit-rana-mechanical-sympathy.md>)

Topics: [distributed-systems](<https://devfeed.tech/topics/distributed-systems.md>), [Resilience](<https://devfeed.tech/topics/resilience.md>), [API](<https://devfeed.tech/topics/api.md>), [Architecture & Design](<https://devfeed.tech/topics/architecture-design.md>)

Tags: [api-design](<https://devfeed.tech/tags/api-design.md>), [apis](<https://devfeed.tech/tags/apis.md>), [async](<https://devfeed.tech/tags/async.md>), [distributed-systems](<https://devfeed.tech/tags/distributed-systems.md>), [fault-tolerance](<https://devfeed.tech/tags/fault-tolerance.md>), [idempotency](<https://devfeed.tech/tags/idempotency.md>), [jitter](<https://devfeed.tech/tags/jitter.md>), [observability](<https://devfeed.tech/tags/observability.md>), [outages](<https://devfeed.tech/tags/outages.md>), [reliability](<https://devfeed.tech/tags/reliability.md>), [resilience](<https://devfeed.tech/tags/resilience.md>), [retries](<https://devfeed.tech/tags/retries.md>), [systems](<https://devfeed.tech/tags/systems.md>)

### AI overview

A practical guide to designing resilient APIs in distributed systems. It covers fail-fast validation, graceful degradation, idempotency for safe retries, and bounded retry policies using exponential backoff, jitter, attempt limits, and total time budgets.

### Source excerpt

Resilience is failing in controlled ways rather than not failing. Validate and fail fast at the boundary, degrade gracefully by serving cached or reduced responses, require an idempotency key for side-effecting operations, and bound retries with exponential backoff, jitter, an attempt cap, and a total time budget. Unbounded retries amplify outages.

## McDonald's Pilots Real-Time Automation for Ice Cream Product Availability

DevFeed: [McDonald's Pilots Real-Time Automation for Ice Cream Product Availability](<https://devfeed.tech/articles/from-ideation-to-automation-the-scoop-on-outages-23978.md>)

Original publisher: [Read original article](<https://medium.com/mcdonalds-technical-blog/from-ideation-to-automation-the-scoop-on-outages-1ad0eab5cee1?source=rss----3bac42476d27---4>)

Author: Global Technology

Published: 2026-03-12T16:13:47Z

Content type: article

Language: en

Sources: [McDonald's Technical Blog - Medium](<https://devfeed.tech/sources/mcdonald-s-technical-blog-medium.md>)

Topics: [Automation](<https://devfeed.tech/topics/automation.md>), [Availability](<https://devfeed.tech/topics/availability.md>), [Internet of things](<https://devfeed.tech/topics/iot.md>), [ordering](<https://devfeed.tech/topics/ordering.md>), [data](<https://devfeed.tech/topics/data.md>)

Tags: [automation](<https://devfeed.tech/tags/automation.md>), [availability](<https://devfeed.tech/tags/availability.md>), [customer-experience](<https://devfeed.tech/tags/customer-experience.md>), [digital-transformation](<https://devfeed.tech/tags/digital-transformation.md>), [efficiency](<https://devfeed.tech/tags/efficiency.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [internet-of-things](<https://devfeed.tech/tags/internet-of-things.md>), [iot](<https://devfeed.tech/tags/iot.md>), [mobile](<https://devfeed.tech/tags/mobile.md>), [ordering](<https://devfeed.tech/tags/ordering.md>), [outages](<https://devfeed.tech/tags/outages.md>), [technology](<https://devfeed.tech/tags/technology.md>)

### AI overview

McDonald's piloted a real-time automation pipeline that connects ice cream machines to Sesame POS. At a Global Tech pilot restaurant, the system automatically removes products during machine downtime and re-enables them when the machine becomes operational, updating kiosks, mobile apps, and POS systems. The solution has not yet been rolled out broadly.

### Source excerpt

A process that once required a multi-click manual process is now fully automated in real time, with a Global Tech pilot restaurant testing the solution to boost efficiency and enhance the customer experience. by: Chloe Tominac, Manager, Engineering Tech Lead & Lauren Adamonis, Manager, Engineering Tech Lead Quick Bytes: Crew members had to manually mark ice cream items unavailable through a multi-click process, and later use the same process to restore items, which often led to missed updates and customer frustration A real-time automation pipeline was piloted to connect the ice cream machine to Sesame POS, instantly updating product availability The solution launched in McDonald's Global Tech pilot restaurant, improving restaurant efficiency and ensuring customers see accurate menus across all ordering channels In restaurant operations, every second counts -- especially when equipment goes offline. That's why McDonald's tech teams set out to automate the ice cream product outage process, transforming a manual workflow into a seamless, real-time system as part of our ongoing digital transformation. To explore how this could work in a live environment, we put the solution to the test in one of our Global Tech pilot restaurants. This pilot is helping us learn how real-time equipment data can improve restaurant efficiency and customer experience. While it's not yet rolled out broadly, the insights from this test are shaping how we think about scaling automation across restaurant operations. Previously, a multi-click process was required to take ice cream items off the menu during machine downtime. Without automatic recovery, items often remained unavailable even after the machine was back online until a crew member restored them, resulting in customers hearing that the ice cream machine was "broken." Automation now ensures menu items are immediately re-enabled as soon as the machine is operational. Now, thanks to a collaboration between the Internet of Things (IoT) team

## Temporal's growth, developer trust, and durable execution for AI systems

DevFeed: [Temporal's growth, developer trust, and durable execution for AI systems](<https://devfeed.tech/articles/a-year-of-change-a-foundation-for-what-s-next-35698.md>)

Original publisher: [Read original article](<https://temporal.io/blog/a-year-of-change-a-foundation-for-whats-next>)

Author: Samar Abbas

Published: 2026-03-05T00:00:00Z

Content type: opinion

Language: en

Sources: [Temporal Blog](<https://devfeed.tech/sources/temporal-blog.md>)

Topics: [systems](<https://devfeed.tech/topics/systems.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [trust](<https://devfeed.tech/topics/trust.md>), [Cloud](<https://devfeed.tech/topics/cloud.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [architecture](<https://devfeed.tech/tags/architecture.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [community](<https://devfeed.tech/tags/community.md>), [developer](<https://devfeed.tech/tags/developer.md>), [enterprise](<https://devfeed.tech/tags/enterprise.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [outages](<https://devfeed.tech/tags/outages.md>), [production](<https://devfeed.tech/tags/production.md>), [scale](<https://devfeed.tech/tags/scale.md>), [trust](<https://devfeed.tech/tags/trust.md>), [workflows](<https://devfeed.tech/tags/workflows.md>)

### AI overview

Temporal's leadership reflects on a year of rapid growth, increased production adoption, and the expanding importance of durable execution for AI and other stateful, long-running systems. The article links developer trust to broader enterprise adoption and describes Temporal's cloud-agnostic and open-source options.

### Source excerpt

Temporal grew 380% YoY, handling AI workflows and durable execution at scale. See how developer trust drives enterprise adoption and what's next.

## Scaling Payments in the Age of Real-Time Fraud: Building a Resilient Foundation for Fintech

DevFeed: [Scaling Payments in the Age of Real-Time Fraud: Building a Resilient Foundation for Fintech](<https://devfeed.tech/articles/scaling-payments-in-the-age-of-real-time-fraud-building-a-resilient-foundation-for-fintech-23814.md>)

Original publisher: [Read original article](<https://cockroachlabs.com/blog/scaling-payments-real-time-fraud-fintech>)

Author: Becca Weng

Published: 2026-02-25T00:00:00Z

Content type: article

Language: en

Sources: [Cockroach Labs](<https://devfeed.tech/sources/cockroach-labs.md>)

Topics: [systems](<https://devfeed.tech/topics/systems.md>), [Architecture & Design](<https://devfeed.tech/topics/architecture-design.md>), [Resilience](<https://devfeed.tech/topics/resilience.md>), [Availability](<https://devfeed.tech/topics/availability.md>), [consistency](<https://devfeed.tech/topics/consistency.md>), [Low Latency](<https://devfeed.tech/topics/low-latency.md>), [data](<https://devfeed.tech/topics/data.md>)

Tags: [fintech](<https://devfeed.tech/tags/fintech.md>), [high-availability](<https://devfeed.tech/tags/high-availability.md>), [latency](<https://devfeed.tech/tags/latency.md>), [outages](<https://devfeed.tech/tags/outages.md>), [payments](<https://devfeed.tech/tags/payments.md>), [pci-dss](<https://devfeed.tech/tags/pci-dss.md>), [real-time](<https://devfeed.tech/tags/real-time.md>), [security](<https://devfeed.tech/tags/security.md>), [sensitive-data](<https://devfeed.tech/tags/sensitive-data.md>)

### AI overview

This article explains why fintech payment systems require resilient architectures that preserve atomic data updates, maintain low latency during demand spikes, support global regulatory and data-residency requirements, and provide high availability and security. It also introduces real-time fraud detection as an added source of complexity.

### Source excerpt

Fintech is one of the most competitive and highly regulated industries in the world. Whether you're a payment processor, digital-first bank, trading platform, or wallet provider, your infrastructure is directly tied to customer trust. In this environment, outages aren't just technical incidents; a single missed transaction, delayed authorization, or moment of downtime can erode customer confidence instantly.

## What Database Modernization Means in the Cloud Era

DevFeed: [What Database Modernization Means in the Cloud Era](<https://devfeed.tech/articles/what-database-modernization-means-in-the-cloud-era-23758.md>)

Original publisher: [Read original article](<https://cockroachlabs.com/blog/cloud-database-modernization>)

Author: David Weiss

Published: 2026-02-09T00:00:00Z

Content type: article

Language: en

Sources: [Cockroach Labs](<https://devfeed.tech/sources/cockroach-labs.md>)

Topics: [Databases](<https://devfeed.tech/topics/databases.md>), [Cloud](<https://devfeed.tech/topics/cloud.md>), [cloud-infrastructure](<https://devfeed.tech/topics/cloud-infrastructure.md>), [Resilience](<https://devfeed.tech/topics/resilience.md>), [legacy](<https://devfeed.tech/topics/legacy.md>), [SQL](<https://devfeed.tech/topics/sql.md>), [consistency](<https://devfeed.tech/topics/consistency.md>)

Tags: [2026](<https://devfeed.tech/tags/2026.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [cloud-infrastructure](<https://devfeed.tech/tags/cloud-infrastructure.md>), [cloud-native](<https://devfeed.tech/tags/cloud-native.md>), [consistency](<https://devfeed.tech/tags/consistency.md>), [databases](<https://devfeed.tech/tags/databases.md>), [distributed](<https://devfeed.tech/tags/distributed.md>), [failover](<https://devfeed.tech/tags/failover.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [legacy](<https://devfeed.tech/tags/legacy.md>), [modernization](<https://devfeed.tech/tags/modernization.md>), [operational](<https://devfeed.tech/tags/operational.md>), [outages](<https://devfeed.tech/tags/outages.md>), [patterns](<https://devfeed.tech/tags/patterns.md>), [performance](<https://devfeed.tech/tags/performance.md>), [resilience](<https://devfeed.tech/tags/resilience.md>), [scale](<https://devfeed.tech/tags/scale.md>), [schema](<https://devfeed.tech/tags/schema.md>), [sql](<https://devfeed.tech/tags/sql.md>), [systems](<https://devfeed.tech/tags/systems.md>)

### AI overview

This article explains database modernization in cloud environments, emphasizing architectures that scale horizontally, tolerate infrastructure failures, support global access, preserve transactional consistency, and reduce operational risk. It presents distributed SQL as an approach for incremental modernization.

### Source excerpt

Database modernization is a priority for many enterprises in 2026, as pressure builds for applications to grow more distributed, failure-tolerant, and globally accessible. Increasing cloud deployments may seem like a quick cure...

## Fraud Doesn't Sleep--Your Infrastructure Can't Either

DevFeed: [Fraud Doesn't Sleep--Your Infrastructure Can't Either](<https://devfeed.tech/articles/fraud-doesn-t-sleep-your-infrastructure-can-t-either-23786.md>)

Original publisher: [Read original article](<https://cockroachlabs.com/blog/fraud-doesnt-sleep-infrastructure-can't>)

Author: Harsh Shah

Published: 2026-02-04T00:00:00Z

Content type: opinion

Language: en

Sources: [Cockroach Labs](<https://devfeed.tech/sources/cockroach-labs.md>)

Topics: [incident](<https://devfeed.tech/topics/incident.md>), [Cockroach Labs](<https://devfeed.tech/topics/cockroach-labs.md>), [distributed-systems](<https://devfeed.tech/topics/distributed-systems.md>), [CockroachDB](<https://devfeed.tech/topics/cockroachdb.md>), [Latency](<https://devfeed.tech/topics/latency.md>), [Resilience](<https://devfeed.tech/topics/resilience.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Amazon SageMaker](<https://devfeed.tech/topics/amazon-sagemaker.md>), [Amazon Bedrock](<https://devfeed.tech/topics/amazon-bedrock.md>), [Embeddings](<https://devfeed.tech/topics/embeddings.md>), [event driven](<https://devfeed.tech/topics/event-driven.md>), [Transactions](<https://devfeed.tech/topics/transactions.md>)

Tags: [aws](<https://devfeed.tech/tags/aws.md>), [azure](<https://devfeed.tech/tags/azure.md>), [cockroachdb](<https://devfeed.tech/tags/cockroachdb.md>), [embeddings](<https://devfeed.tech/tags/embeddings.md>), [event-driven](<https://devfeed.tech/tags/event-driven.md>), [incident](<https://devfeed.tech/tags/incident.md>), [latency](<https://devfeed.tech/tags/latency.md>), [microsoft](<https://devfeed.tech/tags/microsoft.md>), [outages](<https://devfeed.tech/tags/outages.md>), [real-time](<https://devfeed.tech/tags/real-time.md>), [resilience](<https://devfeed.tech/tags/resilience.md>), [sql](<https://devfeed.tech/tags/sql.md>), [transactions](<https://devfeed.tech/tags/transactions.md>)

### AI overview

The article argues that fraud defenses should remain resilient during cloud-provider and regional outages. It presents CockroachDB with AWS AI services, including Amazon SageMaker and Amazon Bedrock, as components of a globally distributed, low-latency, event-driven fraud decisioning architecture.

### Source excerpt

Two major cloud outages in two weeks made one thing painfully clear: your fraud defenses can't depend on any single region or provider being perfect all the time. Microsoft's Azure Front Door misconfiguration rippled through widely used services and status systems, just days after a separate AWS incident disrupted thousands of apps globally. These weren't niche blips--they were broad shocks to the digital economy.

## How Service Level Objectives Align Developers and Product Managers

DevFeed: [How Service Level Objectives Align Developers and Product Managers](<https://devfeed.tech/articles/how-not-to-fight-with-product-managers-as-a-developer-28056.md>)

Original publisher: [Read original article](<https://tech.trivago.com/post/2026-02-02-how-not-to-fight-with-product-managers-as-a-developer/>)

Author: Anis Khan Site Reliability Engineering is my role Cost optimization is my goal GitHub profile Linkedin profile

Published: 2026-02-02T00:00:00Z

Content type: opinion

Language: en

Sources: [Trivago](<https://devfeed.tech/sources/trivago.md>)

Topics: [Availability](<https://devfeed.tech/topics/availability.md>), [User Experience](<https://devfeed.tech/topics/user-experience.md>), [data](<https://devfeed.tech/topics/data.md>)

Tags: [availability](<https://devfeed.tech/tags/availability.md>), [developer](<https://devfeed.tech/tags/developer.md>), [engineering-culture](<https://devfeed.tech/tags/engineering-culture.md>), [maintainability](<https://devfeed.tech/tags/maintainability.md>), [opinions](<https://devfeed.tech/tags/opinions.md>), [outages](<https://devfeed.tech/tags/outages.md>), [performance](<https://devfeed.tech/tags/performance.md>), [site-reliability-engineering](<https://devfeed.tech/tags/site-reliability-engineering.md>), [technical](<https://devfeed.tech/tags/technical.md>), [uptime](<https://devfeed.tech/tags/uptime.md>)

### AI overview

This article argues that Service Level Objectives (SLOs) can reduce conflict between developers and product managers by creating a shared, data-driven agreement about reliability and user experience. It explains how a 99.9% availability target creates a 43.2-minute monthly downtime budget and shows how that budget can guide decisions about releasing features versus restoring stability.

### Source excerpt

It's a scenario developers relate to a little too well. The product manager always comes with more and more feature requests. They also want to release fast by giving a tight deadline. While you...

## From a single point of failure to a cell-based architecture: How we scaled Mercado Envíos' stock...

DevFeed: [From a single point of failure to a cell-based architecture: How we scaled Mercado Envíos' stock...](<https://devfeed.tech/articles/from-a-single-point-of-failure-to-a-cell-based-architecture-how-we-scaled-mercado-envios-stock-22551.md>)

Original publisher: [Read original article](<https://medium.com/mercadolibre-tech/from-a-single-point-of-failure-to-a-cell-based-architecture-how-we-scaled-mercado-env%C3%ADos-stock-528f581fb71b?source=rss----5011f85401f0---4>)

Author: Rafael Silvestri

Published: 2025-12-29T21:09:09Z

Content type: article

Language: en

Sources: [Mercado Libre Tech](<https://devfeed.tech/sources/mercado-libre-tech.md>)

Topics: [distributed-systems](<https://devfeed.tech/topics/distributed-systems.md>), [Architecture & Design](<https://devfeed.tech/topics/architecture-design.md>), [Scalability](<https://devfeed.tech/topics/scalability.md>), [migration](<https://devfeed.tech/topics/migration.md>), [MySQL](<https://devfeed.tech/topics/mysql.md>), [Resilience](<https://devfeed.tech/topics/resilience.md>)

Tags: [architecture](<https://devfeed.tech/tags/architecture.md>), [cell-based-architecture](<https://devfeed.tech/tags/cell-based-architecture.md>), [database](<https://devfeed.tech/tags/database.md>), [database-scalability](<https://devfeed.tech/tags/database-scalability.md>), [distributed-systems](<https://devfeed.tech/tags/distributed-systems.md>), [fury](<https://devfeed.tech/tags/fury.md>), [internaldeveloperplatform](<https://devfeed.tech/tags/internaldeveloperplatform.md>), [migration](<https://devfeed.tech/tags/migration.md>), [outages](<https://devfeed.tech/tags/outages.md>), [resilience](<https://devfeed.tech/tags/resilience.md>), [scalability](<https://devfeed.tech/tags/scalability.md>), [software-architecture](<https://devfeed.tech/tags/software-architecture.md>)

### AI overview

This article explains how Mercado Libre migrated Mercado Envíos' consolidated MySQL-based stock system into independent cells using a cell-based architecture and Fury. The migration was intended to isolate failures, reduce systemic risk, and provide more predictable scalability after the regional database reached its vertical-scaling and operational limits.

### Source excerpt

From a single point of failure to a cell-based architecture: How we scaled Mercado Envíos' stock system How we migrated Mercado Libre's second-largest MySQL instance into independent cells using Fury, reducing risk and achieving predictable scalability in Fulfillment. Abstract What happens when the database supporting a continent's logistics reaches its limits? We reached that point when a critical regional database, used for inventory operations across Latin America (LATAM), could no longer scale vertically. This article describes how we transitioned from that consolidated model to a cell-based architecture -- isolating failures, reducing systemic risk, and improving operational predictability -- while keeping logistics running throughout the migration. When a core system reaches its breaking point Software architecture uses patterns to prevent local failures from causing global outages. One of them is the cell-based architecture, which is conceptually similar to the naval bulkhead mechanism. Ships use watertight compartments, or bulkheads, to divide the hull into separate sections. If one compartment floods, the others remain sealed and the ship keeps moving. In distributed systems, we apply the same idea: each cell operates autonomously, with its own compute, database, and traffic. If one cell fails, the rest continue serving requests. This isolation reduces the blast radius and increases resilience. This pattern became essential at Mercado Libre when the stock system powering Mercado Envíos reached its operational limit. Every inbound, outbound, reservation, and logistics movement depended on a single database that could no longer scale. By early 2024, the question was clear: What do you do when vertical scaling is no longer an option? The problem: One database serving all of LATAM Our initial architecture was simple: multiple stock services connected to a single MySQL cluster. This cluster managed: Stock availability per Fulfillment Center Reservations for Fulfil

## Gremlin Release Roundup 2025: Reliability across AI, on-prem, and applications

DevFeed: [Gremlin Release Roundup 2025: Reliability across AI, on-prem, and applications](<https://devfeed.tech/articles/gremlin-release-roundup-2025-reliability-across-ai-on-prem-and-applications-11684.md>)

Original publisher: [Read original article](<https://www.gremlin.com/blog/release-roundup-2025>)

Author: Andre Newman

Published: 2025-12-15T00:00:00Z

Content type: release

Language: en

Sources: [Gremlin Blog](<https://devfeed.tech/sources/gremlin-blog.md>)

Topics: [on-prem](<https://devfeed.tech/topics/on-prem.md>), [Failure Flags](<https://devfeed.tech/topics/failure-flags.md>), [experiments](<https://devfeed.tech/topics/experiments.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [MCP Server](<https://devfeed.tech/topics/mcp-server.md>), [telemetry](<https://devfeed.tech/topics/telemetry.md>), [SRE](<https://devfeed.tech/topics/sre.md>)

Tags: [2025](<https://devfeed.tech/tags/2025.md>), [ai](<https://devfeed.tech/tags/ai.md>), [analysis](<https://devfeed.tech/tags/analysis.md>), [experiments](<https://devfeed.tech/tags/experiments.md>), [failure-flags](<https://devfeed.tech/tags/failure-flags.md>), [features](<https://devfeed.tech/tags/features.md>), [gremlin](<https://devfeed.tech/tags/gremlin.md>), [mcp](<https://devfeed.tech/tags/mcp.md>), [mcp-server](<https://devfeed.tech/tags/mcp-server.md>), [no-code](<https://devfeed.tech/tags/no-code.md>), [on-prem](<https://devfeed.tech/tags/on-prem.md>), [outages](<https://devfeed.tech/tags/outages.md>), [release](<https://devfeed.tech/tags/release.md>), [releases](<https://devfeed.tech/tags/releases.md>), [sre](<https://devfeed.tech/tags/sre.md>), [telemetry](<https://devfeed.tech/tags/telemetry.md>)

### AI overview

Gremlin's 2025 release roundup describes improvements aimed at preventing outages and making reliability testing easier. Highlights include Reliability Intelligence for analyzing failures and recommending remediations, a Gremlin MCP server for querying environments through an LLM, new experiments and Failure Flags capabilities, expanded platform support, streamlined onboarding, and web UI refinements.

### Source excerpt

This year's release roundup covers our new on-prem offering, new Failure Flags features, intelligent analysis of failed experiments, and much more.

## How the EU Data Act Is Making Multi-Cloud Affordable

DevFeed: [How the EU Data Act Is Making Multi-Cloud Affordable](<https://devfeed.tech/articles/how-the-eu-data-act-is-making-multi-cloud-affordable-23730.md>)

Original publisher: [Read original article](<https://cockroachlabs.com/blog/affordable-multi-cloud>)

Author: Ankur Raina

Published: 2025-12-09T00:00:00Z

Content type: opinion

Language: en

Sources: [Cockroach Labs](<https://devfeed.tech/sources/cockroach-labs.md>)

Topics: [Cloud](<https://devfeed.tech/topics/cloud.md>), [data](<https://devfeed.tech/topics/data.md>), [Amazon Web Services](<https://devfeed.tech/topics/aws.md>), [Google](<https://devfeed.tech/topics/google.md>)

Tags: [availability](<https://devfeed.tech/tags/availability.md>), [aws](<https://devfeed.tech/tags/aws.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [data](<https://devfeed.tech/tags/data.md>), [datacenter](<https://devfeed.tech/tags/datacenter.md>), [eu](<https://devfeed.tech/tags/eu.md>), [google](<https://devfeed.tech/tags/google.md>), [incident](<https://devfeed.tech/tags/incident.md>), [microsoft](<https://devfeed.tech/tags/microsoft.md>), [multi-cloud](<https://devfeed.tech/tags/multi-cloud.md>), [outages](<https://devfeed.tech/tags/outages.md>), [services](<https://devfeed.tech/tags/services.md>)

### AI overview

The article argues that relying on a single cloud provider creates outage risks and that spreading data across clouds can improve resilience. It discusses cost and complexity barriers, including data egress fees, and says the EU Data Act may make switching providers easier.

### Source excerpt

Every major cloud hyperscaler is having global impact outages. The 2024 Microsoft incident happened due to a bug in Crowdstrike. Google Cloud failed this summer because of a small configuration issue. AWS's most popular region US-East-1 was impacted in October 2025 due to a small change in one of their services. In each case, thousands of organizations worldwide saw their applications crash.

## Reliability lessons from the 2025 Cloudflare outage

DevFeed: [Reliability lessons from the 2025 Cloudflare outage](<https://devfeed.tech/articles/reliability-lessons-from-the-2025-cloudflare-outage-11697.md>)

Original publisher: [Read original article](<https://www.gremlin.com/blog/reliability-lessons-from-the-2025-cloudflare-outage>)

Author: Andre Newman

Published: 2025-11-20T00:00:00Z

Content type: article

Language: en

Sources: [Gremlin Blog](<https://devfeed.tech/sources/gremlin-blog.md>)

Topics: [Cloudflare](<https://devfeed.tech/topics/cloudflare.md>), [Bot Management](<https://devfeed.tech/topics/bot-management.md>), [configuration](<https://devfeed.tech/topics/configuration.md>), [Chaos Engineering](<https://devfeed.tech/topics/chaos-engineering.md>), [Post Mortem](<https://devfeed.tech/topics/post-mortem.md>), [Network](<https://devfeed.tech/topics/network.md>), [Workers](<https://devfeed.tech/topics/workers.md>)

Tags: [2025](<https://devfeed.tech/tags/2025.md>), [after-the-retrospective](<https://devfeed.tech/tags/after-the-retrospective.md>), [bot-management](<https://devfeed.tech/tags/bot-management.md>), [chaos-engineering](<https://devfeed.tech/tags/chaos-engineering.md>), [cloudflare](<https://devfeed.tech/tags/cloudflare.md>), [configuration](<https://devfeed.tech/tags/configuration.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [http](<https://devfeed.tech/tags/http.md>), [internet](<https://devfeed.tech/tags/internet.md>), [management](<https://devfeed.tech/tags/management.md>), [outage](<https://devfeed.tech/tags/outage.md>), [outages](<https://devfeed.tech/tags/outages.md>), [post-mortem](<https://devfeed.tech/tags/post-mortem.md>), [workers](<https://devfeed.tech/tags/workers.md>)

### AI overview

The article examines the November 2025 Cloudflare outage, explaining how an oversized Bot Management configuration caused HTTP 5XX errors and cascading failures across dependent services. It highlights configuration propagation, service dependencies, and chaos engineering as reliability considerations.

### Source excerpt

In November 2025, a misconfigured Cloudflare service led to a partial outage. Learn what happened, and what you can do to reduce the impact of similar outages.

[Next page](<https://devfeed.tech/tags/outages.md?cursor=WyIyMDI1LTExLTIwVDAwOjAwOjAwKzAwOjAwIiwgImM5ZTNjODFjLTg4MzMtNGMyMy05MjVhLTAzMTcxNDlmNWU5NSJd>)