# Chaos Engineering

Chaos engineering is the discipline of experimenting on a software system in production in order to build confidence in the system's capability to withstand turbulent and unexpected conditions. Chaos engineering is a disciplined approach to identifying failures before they become outages

This is one page of public article previews, not the complete archive. Follow Next page to continue. Summaries are not the original full articles.

## Testing application resilience with Amazon SQS and AWS Fault Injection Service

DevFeed: [Testing application resilience with Amazon SQS and AWS Fault Injection Service](<https://devfeed.tech/articles/testing-application-resilience-with-amazon-sqs-and-aws-fault-injection-service-4651.md>)

Original publisher: [Read original article](<https://aws.amazon.com/blogs/architecture/testing-application-resilience-with-amazon-sqs-and-aws-fault-injection-service/>)

Author: Richard Whitworth

Published: 2026-09-09T21:33:59Z

Content type: tutorial

Language: en

Sources: [AWS Architecture Blog](<https://devfeed.tech/sources/aws-architecture-blog.md>)

Topics: [Amazon Simple Queue Service (SQS)](<https://devfeed.tech/topics/amazon-simple-queue-service-sqs.md>), [AWS Fault Injection Service (FIS)](<https://devfeed.tech/topics/aws-fault-injection-service-fis.md>), [Chaos Engineering](<https://devfeed.tech/topics/chaos-engineering.md>), [AWS IAM](<https://devfeed.tech/topics/aws-iam.md>), [AWS Identity and Access Management (IAM)](<https://devfeed.tech/topics/aws-identity-and-access-management-iam.md>)

Tags: [advanced-300](<https://devfeed.tech/tags/advanced-300.md>), [amazon-cloudwatch](<https://devfeed.tech/tags/amazon-cloudwatch.md>), [amazon-simple-queue-service-sqs](<https://devfeed.tech/tags/amazon-simple-queue-service-sqs.md>), [amazon-sqs](<https://devfeed.tech/tags/amazon-sqs.md>), [automation](<https://devfeed.tech/tags/automation.md>), [aws](<https://devfeed.tech/tags/aws.md>), [aws-fault-injection-service-fis](<https://devfeed.tech/tags/aws-fault-injection-service-fis.md>), [experiments](<https://devfeed.tech/tags/experiments.md>), [how-to](<https://devfeed.tech/tags/how-to.md>), [iam](<https://devfeed.tech/tags/iam.md>), [observability](<https://devfeed.tech/tags/observability.md>), [resilience](<https://devfeed.tech/tags/resilience.md>), [technical-how-to](<https://devfeed.tech/tags/technical-how-to.md>), [testing](<https://devfeed.tech/tags/testing.md>)

### AI overview

A tutorial on testing application resilience when Amazon SQS data-plane operations fail. It uses AWS Fault Injection Service and Systems Manager Automation to progressively deny queue access, evaluate recovery and observability with CloudWatch metrics, and avoid IAM deny-policy lockouts.

### Source excerpt

Learn how to use AWS Fault Injection Service and AWS Systems Manager Automation to run progressive chaos experiments against Amazon SQS queues. Validate that your retry logic, circuit breakers, and dead-letter queues actually work under failure before a real outage hits production.

## Validating multi-Region DR for Terraform Enterprise with AWS FIS

DevFeed: [Validating multi-Region DR for Terraform Enterprise with AWS FIS](<https://devfeed.tech/articles/validating-multi-region-dr-for-terraform-enterprise-with-aws-fis-4653.md>)

Original publisher: [Read original article](<https://aws.amazon.com/blogs/architecture/validating-multi-region-dr-for-terraform-enterprise-with-aws-fis/>)

Author: Frenil Randeria

Published: 2026-09-09T21:05:02Z

Content type: article

Language: en

Sources: [AWS Architecture Blog](<https://devfeed.tech/sources/aws-architecture-blog.md>)

Topics: [Terraform](<https://devfeed.tech/topics/terraform.md>), [AWS Fault Injection Service (FIS)](<https://devfeed.tech/topics/aws-fault-injection-service-fis.md>), [Chaos Engineering](<https://devfeed.tech/topics/chaos-engineering.md>), [Amazon S3](<https://devfeed.tech/topics/amazon-s3.md>), [Amazon EC2](<https://devfeed.tech/topics/amazon-ec2.md>), [Databases](<https://devfeed.tech/topics/databases.md>)

Tags: [advanced-300](<https://devfeed.tech/tags/advanced-300.md>), [amazon-aurora](<https://devfeed.tech/tags/amazon-aurora.md>), [amazon-ec2](<https://devfeed.tech/tags/amazon-ec2.md>), [amazon-s3](<https://devfeed.tech/tags/amazon-s3.md>), [aws](<https://devfeed.tech/tags/aws.md>), [aws-fault-injection-service-fis](<https://devfeed.tech/tags/aws-fault-injection-service-fis.md>), [disaster-recovery](<https://devfeed.tech/tags/disaster-recovery.md>), [postgresql](<https://devfeed.tech/tags/postgresql.md>), [technical-how-to](<https://devfeed.tech/tags/technical-how-to.md>), [terraform](<https://devfeed.tech/tags/terraform.md>), [validation](<https://devfeed.tech/tags/validation.md>)

### AI overview

This article explains how to validate customer-operated multi-Region disaster recovery for Terraform Enterprise on AWS using three-phase AWS FIS experiments. It covers failover and failback testing, hidden dependency discovery, and reported recovery times of 12-14 minutes.

### Source excerpt

Learn how AWS, HashiCorp, and Athenahealth designed and chaos-tested a multi-Region disaster recovery strategy for Terraform Enterprise on AWS. This post walks through three-phase AWS Fault Injection Service experiments across Amazon EC2, Aurora, and Amazon S3, the 12-14 minute recovery times achieved, and the state file dependency pitfall to avoid.

## Chaos Hub & MCP Prompt Library: Harness Resilience Testing

DevFeed: [Chaos Hub & MCP Prompt Library: Harness Resilience Testing](<https://devfeed.tech/articles/chaos-hub-mcp-prompt-library-harness-resilience-testing-13378.md>)

Original publisher: [Read original article](<https://www.harness.io/blog/chaos-hub-in-docs-prompt-library-for-mcp-whats-new-in-resilience-testing>)

Author: Pritesh Kiri

Published: 2026-08-05T00:00:00Z

Content type: article

Language: en

Sources: [Harness Blog](<https://devfeed.tech/sources/harness-blog.md>)

Topics: [Resilience](<https://devfeed.tech/topics/resilience.md>), [Chaos Engineering](<https://devfeed.tech/topics/chaos-engineering.md>), [Model Context Protocol](<https://devfeed.tech/topics/model-context-protocol.md>), [AWS Fault Injection Service (FIS)](<https://devfeed.tech/topics/aws-fault-injection-service-fis.md>), [MCP Server](<https://devfeed.tech/topics/mcp-server.md>), [Documentation](<https://devfeed.tech/topics/documentation.md>), [cursor](<https://devfeed.tech/topics/cursor.md>), [Claude](<https://devfeed.tech/topics/claude.md>)

Tags: [chaos](<https://devfeed.tech/tags/chaos.md>), [chaos-engineering](<https://devfeed.tech/tags/chaos-engineering.md>), [claude](<https://devfeed.tech/tags/claude.md>), [cursor](<https://devfeed.tech/tags/cursor.md>), [documentation](<https://devfeed.tech/tags/documentation.md>), [mcp](<https://devfeed.tech/tags/mcp.md>), [mcp-server](<https://devfeed.tech/tags/mcp-server.md>), [resilience](<https://devfeed.tech/tags/resilience.md>)

### AI overview

Harness Resilience Testing has added two documentation updates: Chaos Hub is now available directly in the documentation, and a Prompt Library provides ready-to-use prompts for running resilience workflows through Harness MCP using natural language.

### Source excerpt

New in Harness Resilience Testing: Chaos Hub now lives in the docs, plus a Prompt Library for running chaos experiments via MCP using natural language. | Blog

## Eliminate Reliability Blind Spots in AWS, Azure, and GCP

DevFeed: [Eliminate Reliability Blind Spots in AWS, Azure, and GCP](<https://devfeed.tech/articles/eliminate-reliability-blind-spots-in-aws-azure-and-gcp-11565.md>)

Original publisher: [Read original article](<https://www.gremlin.com/blog/eliminate-reliability-blind-spots-detected-risks-aws-azure-gcp>)

Author: Andre Newman

Published: 2026-07-14T00:00:00Z

Content type: article

Language: en

Sources: [Gremlin Blog](<https://devfeed.tech/sources/gremlin-blog.md>)

Topics: [Cloud](<https://devfeed.tech/topics/cloud.md>), [Kubernetes](<https://devfeed.tech/topics/kubernetes.md>), [Resilience](<https://devfeed.tech/topics/resilience.md>), [Amazon Web Services](<https://devfeed.tech/topics/aws.md>), [Azure](<https://devfeed.tech/topics/azure.md>), [Chaos Engineering](<https://devfeed.tech/topics/chaos-engineering.md>)

Tags: [availability](<https://devfeed.tech/tags/availability.md>), [aws](<https://devfeed.tech/tags/aws.md>), [azure](<https://devfeed.tech/tags/azure.md>), [containers](<https://devfeed.tech/tags/containers.md>), [deployment](<https://devfeed.tech/tags/deployment.md>), [experimentation](<https://devfeed.tech/tags/experimentation.md>), [features](<https://devfeed.tech/tags/features.md>), [gcp](<https://devfeed.tech/tags/gcp.md>), [gremlin](<https://devfeed.tech/tags/gremlin.md>), [kubernetes](<https://devfeed.tech/tags/kubernetes.md>), [orchestration](<https://devfeed.tech/tags/orchestration.md>), [outages](<https://devfeed.tech/tags/outages.md>), [resilience](<https://devfeed.tech/tags/resilience.md>), [uptime](<https://devfeed.tech/tags/uptime.md>)

### AI overview

Gremlin's expanded Detected Risks feature automatically identifies high-priority reliability risks across AWS, Azure, GCP, and Kubernetes environments. It is designed to reveal issues such as misconfigured deployments, crash-looping containers, missing readiness probes, and poor pod distribution before they cause outages.

### Source excerpt

Discover reliability risks without running a single test. See how Gremlin identifies high-priority reliability risks across AWS, Azure, and GCP to prevent outages.

## Creating an agentic feedback loop with reliability guardrails

DevFeed: [Creating an agentic feedback loop with reliability guardrails](<https://devfeed.tech/articles/creating-an-agentic-feedback-loop-with-reliability-guardrails-11564.md>)

Original publisher: [Read original article](<https://www.gremlin.com/blog/creating-an-agentic-feedback-loop-with-reliability-guardrails>)

Author: Gavin Cahill

Published: 2026-06-25T00:00:00Z

Content type: article

Language: en

Sources: [Gremlin Blog](<https://devfeed.tech/sources/gremlin-blog.md>)

Topics: [AI Development](<https://devfeed.tech/topics/ai-development.md>), [AI-assisted coding](<https://devfeed.tech/topics/ai-assisted-coding.md>), [Chaos Engineering](<https://devfeed.tech/topics/chaos-engineering.md>), [Resilience](<https://devfeed.tech/topics/resilience.md>), [CI/CD](<https://devfeed.tech/topics/cicd.md>), [Code review](<https://devfeed.tech/topics/code-review.md>)

Tags: [agentic-ai](<https://devfeed.tech/tags/agentic-ai.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-for-code](<https://devfeed.tech/tags/ai-for-code.md>), [chaos-engineering](<https://devfeed.tech/tags/chaos-engineering.md>), [ci-cd](<https://devfeed.tech/tags/ci-cd.md>), [code-generation](<https://devfeed.tech/tags/code-generation.md>), [resilience](<https://devfeed.tech/tags/resilience.md>), [testing](<https://devfeed.tech/tags/testing.md>)

### AI overview

The article explains how reliability guardrails and resilience testing can create an agentic feedback loop for AI coding. It argues that agentic code review and QA alone may miss production failures, so fault injection and Chaos Engineering can validate system behavior under realistic failures and provide data that improves AI-generated code. It also discusses using resilience tests as an automated CI/CD gate.

### Source excerpt

Reliability guardrails are essential for ensuring resilience with AI development, but they can also be used to create a feedback loop for AI context.

## Announcing no-code application fault injection

DevFeed: [Announcing no-code application fault injection](<https://devfeed.tech/articles/announcing-no-code-application-fault-injection-11559.md>)

Original publisher: [Read original article](<https://www.gremlin.com/blog/announcing-failure-flags-no-code-application-fault-injection>)

Author: Andre Newman

Published: 2026-06-02T00:00:00Z

Content type: release

Language: en

Sources: [Gremlin Blog](<https://devfeed.tech/sources/gremlin-blog.md>)

Topics: [Failure Flags](<https://devfeed.tech/topics/failure-flags.md>), [Serverless](<https://devfeed.tech/topics/serverless.md>), [Chaos Engineering](<https://devfeed.tech/topics/chaos-engineering.md>), [Testing](<https://devfeed.tech/topics/testing.md>), [Availability](<https://devfeed.tech/topics/availability.md>), [Latency](<https://devfeed.tech/topics/latency.md>), [observability](<https://devfeed.tech/topics/observability.md>)

Tags: [availability](<https://devfeed.tech/tags/availability.md>), [failure-flags](<https://devfeed.tech/tags/failure-flags.md>), [gremlin](<https://devfeed.tech/tags/gremlin.md>), [health-checks](<https://devfeed.tech/tags/health-checks.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [latency](<https://devfeed.tech/tags/latency.md>), [no-code](<https://devfeed.tech/tags/no-code.md>), [observability](<https://devfeed.tech/tags/observability.md>), [performance](<https://devfeed.tech/tags/performance.md>), [serverless](<https://devfeed.tech/tags/serverless.md>), [testing](<https://devfeed.tech/tags/testing.md>), [uptime](<https://devfeed.tech/tags/uptime.md>)

### AI overview

Gremlin announces Failure Flags by proxy, a no-code application fault injection solution for serverless and managed applications. The sidecar proxy enables reliability tests such as simulating outages, adding latency, and generating exceptions without code changes. Intelligent Health Checks automatically monitor network throughput, latency, and error rate during tests.

### Source excerpt

Gremlin announces Failure Flags by proxy, a no-code application fault injection solution for serverless and managed applications. Learn more in our latest blog post.

## Failure Modes in Distributed Systems and Patterns for Resilient Design

DevFeed: [Failure Modes in Distributed Systems and Patterns for Resilient Design](<https://devfeed.tech/articles/all-the-distributed-systems-failures-in-1-email-18123.md>)

Original publisher: [Read original article](<https://hungrymindsdev.substack.com/p/all-the-distributed-systems-failures>)

Author: Alexandre Zajac

Published: 2026-06-01T15:30:24Z

Content type: article

Language: en

Sources: [Hungry Minds](<https://devfeed.tech/sources/hungry-minds.md>)

Topics: [distributed-systems](<https://devfeed.tech/topics/distributed-systems.md>), [systems](<https://devfeed.tech/topics/systems.md>), [Resilience](<https://devfeed.tech/topics/resilience.md>), [Architecture & Design](<https://devfeed.tech/topics/architecture-design.md>), [Chaos Engineering](<https://devfeed.tech/topics/chaos-engineering.md>)

Tags: [architecture](<https://devfeed.tech/tags/architecture.md>), [chaos-engineering](<https://devfeed.tech/tags/chaos-engineering.md>), [distributed-systems](<https://devfeed.tech/tags/distributed-systems.md>), [resilience](<https://devfeed.tech/tags/resilience.md>), [software-engineering](<https://devfeed.tech/tags/software-engineering.md>), [systems](<https://devfeed.tech/tags/systems.md>)

### AI overview

The article explains recurring failure modes in distributed systems, including Byzantine failures, split-brain scenarios, cascading timeouts, and partial failures. It recommends semantic health checks, defensive timeouts, circuit breakers, bulkheads, and explicit failure-mode testing to improve resilience.

### Source excerpt

PLUS: SWE job market 2026 👨💻, Visual debugging for ML ⚡, S-tier demo framework 👨💻

## Why agentic AI development needs reliability guardrails

DevFeed: [Why agentic AI development needs reliability guardrails](<https://devfeed.tech/articles/why-agentic-ai-development-needs-reliability-guardrails-11742.md>)

Original publisher: [Read original article](<https://www.gremlin.com/blog/why-agentic-ai-development-needs-reliability-guardrails>)

Author: Gavin Cahill

Published: 2026-05-15T00:00:00Z

Content type: article

Language: en

Sources: [Gremlin Blog](<https://devfeed.tech/sources/gremlin-blog.md>)

Topics: [agentic-coding](<https://devfeed.tech/topics/agentic-coding.md>), [AI-assisted coding](<https://devfeed.tech/topics/ai-assisted-coding.md>), [Resilience](<https://devfeed.tech/topics/resilience.md>), [Chaos Engineering](<https://devfeed.tech/topics/chaos-engineering.md>), [systems](<https://devfeed.tech/topics/systems.md>), [Latency](<https://devfeed.tech/topics/latency.md>), [Database](<https://devfeed.tech/topics/database.md>)

Tags: [agentic](<https://devfeed.tech/tags/agentic.md>), [agentic-ai](<https://devfeed.tech/tags/agentic-ai.md>), [agentic-coding](<https://devfeed.tech/tags/agentic-coding.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-for-code](<https://devfeed.tech/tags/ai-for-code.md>), [database](<https://devfeed.tech/tags/database.md>), [latency](<https://devfeed.tech/tags/latency.md>), [resilience](<https://devfeed.tech/tags/resilience.md>), [systems](<https://devfeed.tech/tags/systems.md>), [testing](<https://devfeed.tech/tags/testing.md>)

### AI overview

Agentic AI is accelerating code generation and deployment, but the resulting increase in code volume and reported issue rates raises reliability risks. The article argues for scalable reliability guardrails and fault-injection testing to verify resilience and reduce the risk of outage-causing failures.

### Source excerpt

Companies are moving faster than ever with agentic AI, but that means more risks. Without reliability guardrails, they risk costly outages.

## Reliability Resolutions: How to build effective reliability programs that won't fade away

DevFeed: [Reliability Resolutions: How to build effective reliability programs that won't fade away](<https://devfeed.tech/articles/reliability-resolutions-how-to-build-effective-reliability-programs-that-won-t-fade-away-11608.md>)

Original publisher: [Read original article](<https://www.gremlin.com/blog/how-to-build-effective-reliability-programs-that-wont-fade-away>)

Author: Gavin Cahill

Published: 2026-01-21T00:00:00Z

Content type: article

Language: en

Sources: [Gremlin Blog](<https://devfeed.tech/sources/gremlin-blog.md>)

Topics: [Chaos Engineering](<https://devfeed.tech/topics/chaos-engineering.md>), [Availability](<https://devfeed.tech/topics/availability.md>), [systems](<https://devfeed.tech/topics/systems.md>), [data](<https://devfeed.tech/topics/data.md>)

Tags: [availability](<https://devfeed.tech/tags/availability.md>), [chaos-engineering](<https://devfeed.tech/tags/chaos-engineering.md>), [data](<https://devfeed.tech/tags/data.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [reliability-management](<https://devfeed.tech/tags/reliability-management.md>), [systems](<https://devfeed.tech/tags/systems.md>), [testing](<https://devfeed.tech/tags/testing.md>), [uptime](<https://devfeed.tech/tags/uptime.md>)

### AI overview

This article explains how to build reliability and Chaos Engineering programs that produce lasting results. It recommends aligning reliability work with company goals, assigning ownership, creating repeatable processes, identifying data gaps, and testing specific failure modes on critical systems. Progress can be demonstrated through evidence such as validated failover and achievement of uptime targets.

### Source excerpt

We're already almost through January. How are your reliability resolutions faring? Check out these key questions to help you follow-through and build an effective reliability program.

## How to test application resiliency by simulating the Cloudflare December 2025 outage

DevFeed: [How to test application resiliency by simulating the Cloudflare December 2025 outage](<https://devfeed.tech/articles/how-to-test-application-resiliency-by-simulating-the-cloudflare-december-2025-outage-11636.md>)

Original publisher: [Read original article](<https://www.gremlin.com/blog/how-to-test-application-resiliency-by-simulating-the-cloudflare-december-2025-outage>)

Author: Gavin Cahill

Published: 2025-12-19T00:00:00Z

Content type: tutorial

Language: en

Sources: [Gremlin Blog](<https://devfeed.tech/sources/gremlin-blog.md>)

Topics: [Failure Flags](<https://devfeed.tech/topics/failure-flags.md>), [Resilience](<https://devfeed.tech/topics/resilience.md>), [incident](<https://devfeed.tech/topics/incident.md>), [Chaos Engineering](<https://devfeed.tech/topics/chaos-engineering.md>), [SDKs](<https://devfeed.tech/topics/sdks.md>), [JavaScript](<https://devfeed.tech/topics/javascript.md>), [Node.js](<https://devfeed.tech/topics/node-js.md>), [Python](<https://devfeed.tech/topics/python.md>), [Kubernetes](<https://devfeed.tech/topics/kubernetes.md>), [AWS Lambda](<https://devfeed.tech/topics/aws-lambda.md>), [Azure](<https://devfeed.tech/topics/azure.md>)

Tags: [2025](<https://devfeed.tech/tags/2025.md>), [after-the-retrospective](<https://devfeed.tech/tags/after-the-retrospective.md>), [applications](<https://devfeed.tech/tags/applications.md>), [aws](<https://devfeed.tech/tags/aws.md>), [aws-lambda](<https://devfeed.tech/tags/aws-lambda.md>), [azure](<https://devfeed.tech/tags/azure.md>), [c-sharp](<https://devfeed.tech/tags/c-sharp.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [cloudflare](<https://devfeed.tech/tags/cloudflare.md>), [code](<https://devfeed.tech/tags/code.md>), [failure-flags](<https://devfeed.tech/tags/failure-flags.md>), [gremlin](<https://devfeed.tech/tags/gremlin.md>), [how-to](<https://devfeed.tech/tags/how-to.md>), [http](<https://devfeed.tech/tags/http.md>), [incident](<https://devfeed.tech/tags/incident.md>), [internet-traffic](<https://devfeed.tech/tags/internet-traffic.md>), [javascript](<https://devfeed.tech/tags/javascript.md>), [kubernetes](<https://devfeed.tech/tags/kubernetes.md>), [node-js](<https://devfeed.tech/tags/node-js.md>), [outage](<https://devfeed.tech/tags/outage.md>), [python](<https://devfeed.tech/tags/python.md>)

### AI overview

This tutorial explains how to use Gremlin Failure Flags to simulate HTTP 500 errors associated with the December 2025 Cloudflare outage. It describes application-layer fault injection, the distinction from network-layer experiments, and supported deployment environments and SDK languages.

### Source excerpt

Use Gremlin Failure Flags to safely simulate 500 error codes and test resiliency to outages, like the December 5th Cloudflare incident.

## AWS re:Invent 2025: The top sessions SREs should attend

DevFeed: [AWS re:Invent 2025: The top sessions SREs should attend](<https://devfeed.tech/articles/aws-re-invent-2025-the-top-sessions-sres-should-attend-11614.md>)

Original publisher: [Read original article](<https://incident.io/blog/aws-re-invent-2025>)

Author: Kate Bernacchi-Sass

Published: 2025-11-20T18:29:00Z

Content type: article

Language: en

Sources: [The incident.io Blog](<https://devfeed.tech/sources/the-incident-io-blog.md>)

Topics: [Amazon Web Services](<https://devfeed.tech/topics/aws.md>), [site-reliability-engineering](<https://devfeed.tech/topics/site-reliability-engineering.md>), [Resilience](<https://devfeed.tech/topics/resilience.md>), [Chaos Engineering](<https://devfeed.tech/topics/chaos-engineering.md>), [Incident response](<https://devfeed.tech/topics/incident-response.md>), [Availability](<https://devfeed.tech/topics/availability.md>), [Serverless](<https://devfeed.tech/topics/serverless.md>)

Tags: [2025](<https://devfeed.tech/tags/2025.md>), [architecture](<https://devfeed.tech/tags/architecture.md>), [availability](<https://devfeed.tech/tags/availability.md>), [aws](<https://devfeed.tech/tags/aws.md>), [best-practices](<https://devfeed.tech/tags/best-practices.md>), [chaos-engineering](<https://devfeed.tech/tags/chaos-engineering.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [incident](<https://devfeed.tech/tags/incident.md>), [incident-channel](<https://devfeed.tech/tags/incident-channel.md>), [incident-management](<https://devfeed.tech/tags/incident-management.md>), [incident-response](<https://devfeed.tech/tags/incident-response.md>), [on-call](<https://devfeed.tech/tags/on-call.md>), [operational](<https://devfeed.tech/tags/operational.md>), [outage](<https://devfeed.tech/tags/outage.md>), [post-mortem](<https://devfeed.tech/tags/post-mortem.md>), [production](<https://devfeed.tech/tags/production.md>), [resilience](<https://devfeed.tech/tags/resilience.md>), [scale](<https://devfeed.tech/tags/scale.md>), [slack-incident](<https://devfeed.tech/tags/slack-incident.md>), [sre](<https://devfeed.tech/tags/sre.md>), [systems](<https://devfeed.tech/tags/systems.md>), [techniques](<https://devfeed.tech/tags/techniques.md>)

### AI overview

A curated guide to AWS re:Invent 2025 sessions for SREs and others focused on reliability, incident response, on-call operations, cloud resilience, and resilient systems. It highlights architecture lessons, resilience practices, and a session on testing AWS Lambda with chaos engineering and fault-injection techniques.

### Source excerpt

The top sessions every SRE should see at this year's AWS re:Invent.

## How we prepare Shopify for BFCM

DevFeed: [How we prepare Shopify for BFCM](<https://devfeed.tech/articles/how-we-prepare-shopify-for-bfcm-1307.md>)

Original publisher: [Read original article](<https://shopify.engineering/bfcm-readiness-2025>)

Author: Kyle Petroski; Matthew Frail

Published: 2025-11-20T14:40:48Z

Content type: article

Language: en

Sources: [Shopify Engineering](<https://devfeed.tech/sources/shopify-engineering.md>), [Shopify Engineering - Shopify Engineering](<https://devfeed.tech/sources/shopify-engineering-shopify-engineering.md>)

Topics: [Shopify](<https://devfeed.tech/topics/shopify.md>), [Chaos Engineering](<https://devfeed.tech/topics/chaos-engineering.md>), [Resilience](<https://devfeed.tech/topics/resilience.md>), [Cloud](<https://devfeed.tech/topics/cloud.md>), [Google Cloud Platform (GCP)](<https://devfeed.tech/topics/google-cloud.md>)

Tags: [chaos-engineering](<https://devfeed.tech/tags/chaos-engineering.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [google-cloud](<https://devfeed.tech/tags/google-cloud.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [performance](<https://devfeed.tech/tags/performance.md>), [production](<https://devfeed.tech/tags/production.md>), [resilience](<https://devfeed.tech/tags/resilience.md>), [shopify](<https://devfeed.tech/tags/shopify.md>)

### AI overview

Shopify describes its year-round preparation for Black Friday Cyber Monday by modeling traffic, expanding capacity across multiple Google Cloud regions, reviewing infrastructure changes, and conducting risk assessments. Large-scale Game Days and fire drills simulated extreme production load, exposing issues such as Kafka bottlenecks, memory pressure, and timeouts that were fixed and revalidated.

### Source excerpt

From March to October we simulated traffic tsunamis, injected chaos, and fixed every bottleneck before our merchants needed us most.

## Reliability lessons from the 2025 Cloudflare outage

DevFeed: [Reliability lessons from the 2025 Cloudflare outage](<https://devfeed.tech/articles/reliability-lessons-from-the-2025-cloudflare-outage-11697.md>)

Original publisher: [Read original article](<https://www.gremlin.com/blog/reliability-lessons-from-the-2025-cloudflare-outage>)

Author: Andre Newman

Published: 2025-11-20T00:00:00Z

Content type: article

Language: en

Sources: [Gremlin Blog](<https://devfeed.tech/sources/gremlin-blog.md>)

Topics: [Cloudflare](<https://devfeed.tech/topics/cloudflare.md>), [Bot Management](<https://devfeed.tech/topics/bot-management.md>), [configuration](<https://devfeed.tech/topics/configuration.md>), [Chaos Engineering](<https://devfeed.tech/topics/chaos-engineering.md>), [Post Mortem](<https://devfeed.tech/topics/post-mortem.md>), [Network](<https://devfeed.tech/topics/network.md>), [Workers](<https://devfeed.tech/topics/workers.md>)

Tags: [2025](<https://devfeed.tech/tags/2025.md>), [after-the-retrospective](<https://devfeed.tech/tags/after-the-retrospective.md>), [bot-management](<https://devfeed.tech/tags/bot-management.md>), [chaos-engineering](<https://devfeed.tech/tags/chaos-engineering.md>), [cloudflare](<https://devfeed.tech/tags/cloudflare.md>), [configuration](<https://devfeed.tech/tags/configuration.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [http](<https://devfeed.tech/tags/http.md>), [internet](<https://devfeed.tech/tags/internet.md>), [management](<https://devfeed.tech/tags/management.md>), [outage](<https://devfeed.tech/tags/outage.md>), [outages](<https://devfeed.tech/tags/outages.md>), [post-mortem](<https://devfeed.tech/tags/post-mortem.md>), [workers](<https://devfeed.tech/tags/workers.md>)

### AI overview

The article examines the November 2025 Cloudflare outage, explaining how an oversized Bot Management configuration caused HTTP 5XX errors and cascading failures across dependent services. It highlights configuration propagation, service dependencies, and chaos engineering as reliability considerations.

### Source excerpt

In November 2025, a misconfigured Cloudflare service led to a partial outage. Learn what happened, and what you can do to reduce the impact of similar outages.

## Improve Kubernetes reliability faster with Gremlin and Dynatrace

DevFeed: [Improve Kubernetes reliability faster with Gremlin and Dynatrace](<https://devfeed.tech/articles/improve-kubernetes-reliability-faster-with-gremlin-and-dynatrace-11649.md>)

Original publisher: [Read original article](<https://www.gremlin.com/blog/improve-kubernetes-reliability-faster-with-gremlin-and-dynatrace>)

Author: Gavin Cahill

Published: 2025-11-10T00:00:00Z

Content type: article

Language: en

Sources: [Gremlin Blog](<https://devfeed.tech/sources/gremlin-blog.md>)

Topics: [Kubernetes](<https://devfeed.tech/topics/kubernetes.md>), [dynatrace](<https://devfeed.tech/topics/dynatrace.md>), [observability](<https://devfeed.tech/topics/observability.md>), [Chaos Engineering](<https://devfeed.tech/topics/chaos-engineering.md>), [ai observability](<https://devfeed.tech/topics/ai-observability.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [announcements](<https://devfeed.tech/tags/announcements.md>), [cloud-native](<https://devfeed.tech/tags/cloud-native.md>), [dynatrace](<https://devfeed.tech/tags/dynatrace.md>), [gremlin](<https://devfeed.tech/tags/gremlin.md>), [integration](<https://devfeed.tech/tags/integration.md>), [kubernetes](<https://devfeed.tech/tags/kubernetes.md>), [observability](<https://devfeed.tech/tags/observability.md>), [testing](<https://devfeed.tech/tags/testing.md>)

### AI overview

The article presents a strategic integration between Gremlin and Dynatrace that automatically discovers Kubernetes services configured in Dynatrace, making fault-injection testing faster to set up. It explains how health checks use observability metrics such as error rates and latency to evaluate experiments, validate monitoring and alerting, and improve reliability across cloud-native architectures.

### Source excerpt

It's easier than ever to start testing Kubernetes with Dynatrace and Gremlin. The new strategic integration automatically discovers objects to make testing set up simple and fast.

## How to test the reliability of a Point of Sale (POS) system

DevFeed: [How to test the reliability of a Point of Sale (POS) system](<https://devfeed.tech/articles/how-to-test-the-reliability-of-a-point-of-sale-pos-system-11640.md>)

Original publisher: [Read original article](<https://www.gremlin.com/blog/how-to-test-the-reliability-of-a-point-of-sale-pos-system>)

Author: Gavin Cahill

Published: 2025-10-20T00:00:00Z

Content type: tutorial

Language: en

Sources: [Gremlin Blog](<https://devfeed.tech/sources/gremlin-blog.md>)

Topics: [Chaos Engineering](<https://devfeed.tech/topics/chaos-engineering.md>), [Resilience](<https://devfeed.tech/topics/resilience.md>), [autoscaling](<https://devfeed.tech/topics/autoscaling.md>), [Microservice](<https://devfeed.tech/topics/microservice.md>), [Complex Systems](<https://devfeed.tech/topics/complex-systems.md>), [systems](<https://devfeed.tech/topics/systems.md>)

Tags: [autoscaling](<https://devfeed.tech/tags/autoscaling.md>), [best-practices](<https://devfeed.tech/tags/best-practices.md>), [chaos-engineering](<https://devfeed.tech/tags/chaos-engineering.md>), [cpu](<https://devfeed.tech/tags/cpu.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [gremlin](<https://devfeed.tech/tags/gremlin.md>), [how-to](<https://devfeed.tech/tags/how-to.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [memory](<https://devfeed.tech/tags/memory.md>), [microservices](<https://devfeed.tech/tags/microservices.md>), [outage](<https://devfeed.tech/tags/outage.md>), [reliability-management](<https://devfeed.tech/tags/reliability-management.md>), [resilience](<https://devfeed.tech/tags/resilience.md>), [retail](<https://devfeed.tech/tags/retail.md>), [testing](<https://devfeed.tech/tags/testing.md>)

### AI overview

This tutorial explains how to test the reliability of retail Point of Sale systems using Gremlin and Chaos Engineering. It focuses on resilience testing for microservice-based checkout systems, including autoscaling, CPU, memory, and disk I/O capacity, to identify failure conditions and reduce outages.

### Source excerpt

Find out how to use Gremlin and Chaos Engineering to make sure your Point of Sale system is reliable.

## Chaos Engineering works, but it has to scale

DevFeed: [Chaos Engineering works, but it has to scale](<https://devfeed.tech/articles/chaos-engineering-works-but-it-has-to-scale-11563.md>)

Original publisher: [Read original article](<https://www.gremlin.com/blog/chaos-engineering-works-but-it-has-to-scale>)

Author: Gavin Cahill

Published: 2025-10-07T00:00:00Z

Content type: article

Language: en

Sources: [Gremlin Blog](<https://devfeed.tech/sources/gremlin-blog.md>)

Topics: [Chaos Engineering](<https://devfeed.tech/topics/chaos-engineering.md>), [Resilience](<https://devfeed.tech/topics/resilience.md>), [Testing](<https://devfeed.tech/topics/testing.md>), [distributed-systems](<https://devfeed.tech/topics/distributed-systems.md>), [Availability](<https://devfeed.tech/topics/availability.md>), [site-reliability-engineering](<https://devfeed.tech/topics/site-reliability-engineering.md>)

Tags: [availability](<https://devfeed.tech/tags/availability.md>), [chaos-engineering](<https://devfeed.tech/tags/chaos-engineering.md>), [outages](<https://devfeed.tech/tags/outages.md>), [resilience](<https://devfeed.tech/tags/resilience.md>), [sre](<https://devfeed.tech/tags/sre.md>), [testing](<https://devfeed.tech/tags/testing.md>), [tests](<https://devfeed.tech/tags/tests.md>), [uptime](<https://devfeed.tech/tags/uptime.md>)

### AI overview

Chaos Engineering can uncover failure modes and help prevent outages, but organization-wide adoption may stall when expertise is concentrated in a small number of teams. The article recommends scaling the practice through standards, validation testing, and reporting so that reliability improvements extend beyond critical services.

### Source excerpt

Chaos Engineering effectively improves the reliability of systems, but it can run into snags when you try to scale. Build on Chaos Engineering with these key actions.

## How to get fast, easy insights with the Gremlin MCP Server

DevFeed: [How to get fast, easy insights with the Gremlin MCP Server](<https://devfeed.tech/articles/how-to-get-fast-easy-insights-with-the-gremlin-mcp-server-11620.md>)

Original publisher: [Read original article](<https://www.gremlin.com/blog/how-to-get-fast-easy-insights-with-the-gremlin-mcp-server>)

Author: Gavin Cahill

Published: 2025-08-28T00:00:00Z

Content type: article

Language: en

Sources: [Gremlin Blog](<https://devfeed.tech/sources/gremlin-blog.md>)

Topics: [MCP Server](<https://devfeed.tech/topics/mcp-server.md>), [Model Context Protocol](<https://devfeed.tech/topics/model-context-protocol.md>), [Chaos Engineering](<https://devfeed.tech/topics/chaos-engineering.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [test-coverage](<https://devfeed.tech/topics/test-coverage.md>), [Authorization](<https://devfeed.tech/topics/authorization.md>), [Containers](<https://devfeed.tech/topics/containers.md>), [dashboards](<https://devfeed.tech/topics/dashboards.md>), [Security](<https://devfeed.tech/topics/security.md>), [API keys](<https://devfeed.tech/topics/api-keys.md>)

Tags: [access-control](<https://devfeed.tech/tags/access-control.md>), [ai](<https://devfeed.tech/tags/ai.md>), [api](<https://devfeed.tech/tags/api.md>), [api-keys](<https://devfeed.tech/tags/api-keys.md>), [applications](<https://devfeed.tech/tags/applications.md>), [chaos-engineering](<https://devfeed.tech/tags/chaos-engineering.md>), [chatgpt](<https://devfeed.tech/tags/chatgpt.md>), [claude](<https://devfeed.tech/tags/claude.md>), [components](<https://devfeed.tech/tags/components.md>), [core](<https://devfeed.tech/tags/core.md>), [data](<https://devfeed.tech/tags/data.md>), [gremlin](<https://devfeed.tech/tags/gremlin.md>), [hardening](<https://devfeed.tech/tags/hardening.md>), [how-to](<https://devfeed.tech/tags/how-to.md>), [insights](<https://devfeed.tech/tags/insights.md>), [installation](<https://devfeed.tech/tags/installation.md>), [llm](<https://devfeed.tech/tags/llm.md>), [management](<https://devfeed.tech/tags/management.md>), [mcp](<https://devfeed.tech/tags/mcp.md>), [mcp-server](<https://devfeed.tech/tags/mcp-server.md>)

### AI overview

The article explains how the Gremlin MCP Server connects an LLM client to the Gremlin API so teams can explore reliability-testing data and identify insights using natural-language prompts. It covers the client-server architecture, containerized deployment, security hardening, non-destructive API operations, dashboards, reporting, evaluation, and RBAC-based control of API keys.

### Source excerpt

Find out how to quickly and easily uncover new reliability insights by using the Gremlin MCP Server and your favorite LLM.

## Fix issues faster with Recommended Remediations

DevFeed: [Fix issues faster with Recommended Remediations](<https://devfeed.tech/articles/fix-issues-faster-with-recommended-remediations-11569.md>)

Original publisher: [Read original article](<https://www.gremlin.com/blog/fix-issues-faster-with-recommended-remediations>)

Author: Gavin Cahill

Published: 2025-08-22T00:00:00Z

Content type: release

Language: en

Sources: [Gremlin Blog](<https://devfeed.tech/sources/gremlin-blog.md>)

Topics: [Chaos Engineering](<https://devfeed.tech/topics/chaos-engineering.md>), [Testing](<https://devfeed.tech/topics/testing.md>)

Tags: [chaos-engineering](<https://devfeed.tech/tags/chaos-engineering.md>), [features](<https://devfeed.tech/tags/features.md>), [gremlin](<https://devfeed.tech/tags/gremlin.md>), [recommendations](<https://devfeed.tech/tags/recommendations.md>), [uptime](<https://devfeed.tech/tags/uptime.md>)

### AI overview

Gremlin's Recommended Remediation analyzes fault-injection test results to identify likely failure causes and suggest ways to address reliability issues. It builds on Experiment Analysis by combining test data, metrics, health checks, and key events, then applies reliability expertise to produce tailored recommendations.

### Source excerpt

Recommended Remediation speeds up teams with tailored suggestions to help you address reliability risks before they cause failures.

## How Experiment Analysis uncovers the cause behind failures

DevFeed: [How Experiment Analysis uncovers the cause behind failures](<https://devfeed.tech/articles/how-experiment-analysis-uncovers-the-cause-behind-failures-11589.md>)

Original publisher: [Read original article](<https://www.gremlin.com/blog/how-experiment-analysis-uncovers-the-cause-behind-failures>)

Author: Gavin Cahill

Published: 2025-08-15T00:00:00Z

Content type: article

Language: en

Sources: [Gremlin Blog](<https://devfeed.tech/sources/gremlin-blog.md>)

Topics: [experiments](<https://devfeed.tech/topics/experiments.md>), [Chaos Engineering](<https://devfeed.tech/topics/chaos-engineering.md>), [observability](<https://devfeed.tech/topics/observability.md>), [Machine learning](<https://devfeed.tech/topics/machine-learning.md>), [Availability](<https://devfeed.tech/topics/availability.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [analysis](<https://devfeed.tech/tags/analysis.md>), [availability](<https://devfeed.tech/tags/availability.md>), [chaos-engineering](<https://devfeed.tech/tags/chaos-engineering.md>), [experiment](<https://devfeed.tech/tags/experiment.md>), [failover](<https://devfeed.tech/tags/failover.md>), [gremlin](<https://devfeed.tech/tags/gremlin.md>), [observability](<https://devfeed.tech/tags/observability.md>), [testing](<https://devfeed.tech/tags/testing.md>)

### AI overview

The article explains how Gremlin's Experiment Analysis uses machine learning and observability context to identify why Chaos Engineering tests fail, reducing the manual effort needed to investigate system behavior. It also emphasizes defining expected behavior as the basis for pass/fail decisions, illustrated by a zone failover scenario.

### Source excerpt

Using knowledge of your data and your systems, Experiment Analysis connects the cause with the effect to help you pinpoint the root problem faster.

## Antifragile Systems and Teams

DevFeed: [Antifragile Systems and Teams](<https://devfeed.tech/articles/antifragile-systems-and-teams-27864.md>)

Original publisher: [Read original article](<https://gagor.pro/book/2025/antifragile-systems-and-teams/>)

Author: Tom

Published: 2025-08-14T00:00:00Z

Content type: opinion

Language: en

Sources: [Tomasz Gągor](<https://devfeed.tech/sources/tomasz-gagor.md>)

Topics: [systems](<https://devfeed.tech/topics/systems.md>), [Resilience](<https://devfeed.tech/topics/resilience.md>), [DevOps](<https://devfeed.tech/topics/devops.md>), [Chaos Engineering](<https://devfeed.tech/topics/chaos-engineering.md>), [Automation](<https://devfeed.tech/topics/automation.md>), [Tooling](<https://devfeed.tech/topics/tooling.md>), [Netflix](<https://devfeed.tech/topics/netflix.md>)

Tags: [automation](<https://devfeed.tech/tags/automation.md>), [chaos](<https://devfeed.tech/tags/chaos.md>), [chaos-engineering](<https://devfeed.tech/tags/chaos-engineering.md>), [culture](<https://devfeed.tech/tags/culture.md>), [devops](<https://devfeed.tech/tags/devops.md>), [netflix](<https://devfeed.tech/tags/netflix.md>), [resilience](<https://devfeed.tech/tags/resilience.md>), [systems](<https://devfeed.tech/tags/systems.md>), [technology](<https://devfeed.tech/tags/technology.md>)

### AI overview

A concise review of Dave Zwieback's book about antifragile systems and teams. It explains how organizations can use change, mistakes, and small failures to learn and become stronger, emphasizing DevOps practices such as culture, automation, measurement, sharing, chaos engineering, and frequent small deployments.

### Source excerpt

Antifragile Systems and Teams Author: Dave Zwieback This is a short read, but it does a solid job of capturing an important idea: that the healthiest systems and teams aren't just resistant to change, they actually get stronger through it. Zwieback contrasts fragile organizations - those that try to lock things down and prevent every possible failure - with antifragile ones, which use volatility, mistakes, and small shocks as opportunities to learn and improve.

## Reliability Intelligence: your reliability expert

DevFeed: [Reliability Intelligence: your reliability expert](<https://devfeed.tech/articles/reliability-intelligence-your-reliability-expert-11693.md>)

Original publisher: [Read original article](<https://www.gremlin.com/blog/reliability-intelligence-your-reliability-expert>)

Author: Gavin Cahill

Published: 2025-08-11T00:00:00Z

Content type: release

Language: en

Sources: [Gremlin Blog](<https://devfeed.tech/sources/gremlin-blog.md>)

Topics: [Chaos Engineering](<https://devfeed.tech/topics/chaos-engineering.md>), [experiments](<https://devfeed.tech/topics/experiments.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [telemetry](<https://devfeed.tech/topics/telemetry.md>), [site-reliability-engineering](<https://devfeed.tech/topics/site-reliability-engineering.md>), [MCP Server](<https://devfeed.tech/topics/mcp-server.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [announcements](<https://devfeed.tech/tags/announcements.md>), [chaos-engineering](<https://devfeed.tech/tags/chaos-engineering.md>), [experiment](<https://devfeed.tech/tags/experiment.md>), [mcp-server](<https://devfeed.tech/tags/mcp-server.md>), [release](<https://devfeed.tech/tags/release.md>), [sre](<https://devfeed.tech/tags/sre.md>), [telemetry](<https://devfeed.tech/tags/telemetry.md>)

### AI overview

Gremlin announces Reliability Intelligence, a release that combines Experiment Analysis, Recommended Remediation, and an MCP Server. It uses Gremlin's reliability expertise, telemetry, and trace data to help engineers run reliability tests, identify root causes, detect regressions, and remediate issues faster without slowing deployment velocity.

### Source excerpt

Gremlin's Reliability Intelligence combines Experiment Analysis, Recommended Remediation, and an MCP Server to help teams increase reliability faster than ever.

## 4 Chaos Engineering recommendations from Gartner

DevFeed: [4 Chaos Engineering recommendations from Gartner](<https://devfeed.tech/articles/4-chaos-engineering-recommendations-from-gartner-11557.md>)

Original publisher: [Read original article](<https://www.gremlin.com/blog/4-chaos-engineering-recommendations-from-gartner>)

Author: Gavin Cahill

Published: 2025-07-11T00:00:00Z

Content type: article

Language: en

Sources: [Gremlin Blog](<https://devfeed.tech/sources/gremlin-blog.md>)

Topics: [Chaos Engineering](<https://devfeed.tech/topics/chaos-engineering.md>), [Resilience](<https://devfeed.tech/topics/resilience.md>), [genai](<https://devfeed.tech/topics/genai.md>), [Testing](<https://devfeed.tech/topics/testing.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [analysts](<https://devfeed.tech/tags/analysts.md>), [chaos-engineering](<https://devfeed.tech/tags/chaos-engineering.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [generative-ai](<https://devfeed.tech/tags/generative-ai.md>), [resilience](<https://devfeed.tech/tags/resilience.md>), [test](<https://devfeed.tech/tags/test.md>), [testing](<https://devfeed.tech/tags/testing.md>)

### AI overview

The article presents four Gartner recommendations for adopting Chaos Engineering, including testing generative AI API fallback patterns and using scenario-based GameDays to evaluate outage response and identify reliability risks.

### Source excerpt

Gartner recently published their 2025 Hype Cycle for Infrastructure Platforms. Here are four key recommendations from them for adopting Chaos Engineering.

## Infographic: Resilience and reliability in the cloud

DevFeed: [Infographic: Resilience and reliability in the cloud](<https://devfeed.tech/articles/infographic-resilience-and-reliability-in-the-cloud-11651.md>)

Original publisher: [Read original article](<https://www.gremlin.com/blog/infographic-resilience-and-reliability-in-the-cloud>)

Author: Gavin Cahill

Published: 2025-02-25T00:00:00Z

Content type: article

Language: en

Sources: [Gremlin Blog](<https://devfeed.tech/sources/gremlin-blog.md>)

Topics: [Resilience](<https://devfeed.tech/topics/resilience.md>), [Cloud](<https://devfeed.tech/topics/cloud.md>), [Chaos Engineering](<https://devfeed.tech/topics/chaos-engineering.md>), [observability](<https://devfeed.tech/topics/observability.md>), [Amazon Web Services](<https://devfeed.tech/topics/aws.md>)

Tags: [aws](<https://devfeed.tech/tags/aws.md>), [chaos-engineering](<https://devfeed.tech/tags/chaos-engineering.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [cost-of-downtime](<https://devfeed.tech/tags/cost-of-downtime.md>), [observability](<https://devfeed.tech/tags/observability.md>), [outages](<https://devfeed.tech/tags/outages.md>), [resilience](<https://devfeed.tech/tags/resilience.md>), [testing](<https://devfeed.tech/tags/testing.md>), [uptime](<https://devfeed.tech/tags/uptime.md>)

### AI overview

This infographic explains why resilience and reliability matter for cloud-based software systems. It presents outage impacts and causes, discusses resilience testing and Chaos Engineering with tools such as Gremlin, and describes how organizations can evaluate the return on investment of reliability efforts.

### Source excerpt

Created in partnership with AWS, this infographic shows the impact of outages, the most common causes of outages, and the results companies get from investing in resilience.

## Announcing Gremlin Private Edition

DevFeed: [Announcing Gremlin Private Edition](<https://devfeed.tech/articles/announcing-gremlin-private-edition-11560.md>)

Original publisher: [Read original article](<https://www.gremlin.com/blog/announcing-gremlin-private-edition>)

Author: Andre Newman

Published: 2025-02-11T00:00:00Z

Content type: article

Language: en

Sources: [Gremlin Blog](<https://devfeed.tech/sources/gremlin-blog.md>)

Topics: [Chaos Engineering](<https://devfeed.tech/topics/chaos-engineering.md>), [Kubernetes](<https://devfeed.tech/topics/kubernetes.md>), [Security](<https://devfeed.tech/topics/security.md>), [Software as a service](<https://devfeed.tech/topics/saas.md>), [datacenter](<https://devfeed.tech/topics/datacenter.md>)

Tags: [announcement](<https://devfeed.tech/tags/announcement.md>), [announcements](<https://devfeed.tech/tags/announcements.md>), [blog](<https://devfeed.tech/tags/blog.md>), [gremlin](<https://devfeed.tech/tags/gremlin.md>), [kubernetes](<https://devfeed.tech/tags/kubernetes.md>), [monitoring](<https://devfeed.tech/tags/monitoring.md>), [security](<https://devfeed.tech/tags/security.md>), [self-hosted](<https://devfeed.tech/tags/self-hosted.md>), [testing](<https://devfeed.tech/tags/testing.md>)

### AI overview

Gremlin Private Edition is a self-hosted version of Gremlin that runs entirely within an organization's private network. It provides the SaaS product's reliability-testing features while allowing organizations to control deployment, access, and data storage. The platform runs on Kubernetes and supports private, cloud, datacenter, and air-gapped environments.

### Source excerpt

Gremlin Private Edition is a private, secure, self-hosted Gremlin instance that runs entirely within your network. Learn more in our announcement blog.

[Next page](<https://devfeed.tech/topics/chaos-engineering.md?cursor=WyIyMDI1LTAyLTExVDAwOjAwOjAwKzAwOjAwIiwgIjMxZDNkNmJlLWU3YmItNDQyNC1iNmQyLWQ5YmU2ODcxMTMwZCJd>)