# Gremlin

Published articles for Gremlin.

This is one page of public article previews, not the complete archive. Follow Next page to continue. Summaries are not the original full articles.

## The Gremlin app for Dynatrace: resilience testing and reliability scoring, built on the observability you already trust

DevFeed: [The Gremlin app for Dynatrace: resilience testing and reliability scoring, built on the observability you already trust](<https://devfeed.tech/articles/the-gremlin-app-for-dynatrace-resilience-testing-and-reliability-scoring-built-on-the-observability-you-already-trust-11572.md>)

Original publisher: [Read original article](<https://www.gremlin.com/blog/gremlin-app-for-dynatrace>)

Author: Ryan Detwiller

Published: 2026-07-28T00:00:00Z

Content type: article

Language: en

Sources: [Gremlin Blog](<https://devfeed.tech/sources/gremlin-blog.md>)

Topics: [Resilience](<https://devfeed.tech/topics/resilience.md>), [dynatrace](<https://devfeed.tech/topics/dynatrace.md>), [observability](<https://devfeed.tech/topics/observability.md>), [Instrumentation](<https://devfeed.tech/topics/instrumentation.md>), [distributed-systems](<https://devfeed.tech/topics/distributed-systems.md>)

Tags: [announcements](<https://devfeed.tech/tags/announcements.md>), [distributed-systems](<https://devfeed.tech/tags/distributed-systems.md>), [dynatrace](<https://devfeed.tech/tags/dynatrace.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [gremlin](<https://devfeed.tech/tags/gremlin.md>), [incident](<https://devfeed.tech/tags/incident.md>), [instrumentation](<https://devfeed.tech/tags/instrumentation.md>), [metrics](<https://devfeed.tech/tags/metrics.md>), [observability](<https://devfeed.tech/tags/observability.md>), [real-time](<https://devfeed.tech/tags/real-time.md>), [resilience](<https://devfeed.tech/tags/resilience.md>), [safety](<https://devfeed.tech/tags/safety.md>), [systems](<https://devfeed.tech/tags/systems.md>), [testing](<https://devfeed.tech/tags/testing.md>), [uptime](<https://devfeed.tech/tags/uptime.md>)

### AI overview

The Gremlin app for Dynatrace adds resilience testing and reliability scoring to Dynatrace workflows. Teams can run reliability tests, observe their impact in real time, and track service-level reliability scores using existing Dynatrace metrics, alerts, instrumentation, and health checks.

### Source excerpt

With the Gremlin app for Dyantrace, you get resilience testing and reliability scoring built on the observability you already trust.

## Eliminate Reliability Blind Spots in AWS, Azure, and GCP

DevFeed: [Eliminate Reliability Blind Spots in AWS, Azure, and GCP](<https://devfeed.tech/articles/eliminate-reliability-blind-spots-in-aws-azure-and-gcp-11565.md>)

Original publisher: [Read original article](<https://www.gremlin.com/blog/eliminate-reliability-blind-spots-detected-risks-aws-azure-gcp>)

Author: Andre Newman

Published: 2026-07-14T00:00:00Z

Content type: article

Language: en

Sources: [Gremlin Blog](<https://devfeed.tech/sources/gremlin-blog.md>)

Topics: [Cloud](<https://devfeed.tech/topics/cloud.md>), [Kubernetes](<https://devfeed.tech/topics/kubernetes.md>), [Resilience](<https://devfeed.tech/topics/resilience.md>), [Amazon Web Services](<https://devfeed.tech/topics/aws.md>), [Azure](<https://devfeed.tech/topics/azure.md>), [Chaos Engineering](<https://devfeed.tech/topics/chaos-engineering.md>)

Tags: [availability](<https://devfeed.tech/tags/availability.md>), [aws](<https://devfeed.tech/tags/aws.md>), [azure](<https://devfeed.tech/tags/azure.md>), [containers](<https://devfeed.tech/tags/containers.md>), [deployment](<https://devfeed.tech/tags/deployment.md>), [experimentation](<https://devfeed.tech/tags/experimentation.md>), [features](<https://devfeed.tech/tags/features.md>), [gcp](<https://devfeed.tech/tags/gcp.md>), [gremlin](<https://devfeed.tech/tags/gremlin.md>), [kubernetes](<https://devfeed.tech/tags/kubernetes.md>), [orchestration](<https://devfeed.tech/tags/orchestration.md>), [outages](<https://devfeed.tech/tags/outages.md>), [resilience](<https://devfeed.tech/tags/resilience.md>), [uptime](<https://devfeed.tech/tags/uptime.md>)

### AI overview

Gremlin's expanded Detected Risks feature automatically identifies high-priority reliability risks across AWS, Azure, GCP, and Kubernetes environments. It is designed to reveal issues such as misconfigured deployments, crash-looping containers, missing readiness probes, and poor pod distribution before they cause outages.

### Source excerpt

Discover reliability risks without running a single test. See how Gremlin identifies high-priority reliability risks across AWS, Azure, and GCP to prevent outages.

## Announcing no-code application fault injection

DevFeed: [Announcing no-code application fault injection](<https://devfeed.tech/articles/announcing-no-code-application-fault-injection-11559.md>)

Original publisher: [Read original article](<https://www.gremlin.com/blog/announcing-failure-flags-no-code-application-fault-injection>)

Author: Andre Newman

Published: 2026-06-02T00:00:00Z

Content type: release

Language: en

Sources: [Gremlin Blog](<https://devfeed.tech/sources/gremlin-blog.md>)

Topics: [Failure Flags](<https://devfeed.tech/topics/failure-flags.md>), [Serverless](<https://devfeed.tech/topics/serverless.md>), [Chaos Engineering](<https://devfeed.tech/topics/chaos-engineering.md>), [Testing](<https://devfeed.tech/topics/testing.md>), [Availability](<https://devfeed.tech/topics/availability.md>), [Latency](<https://devfeed.tech/topics/latency.md>), [observability](<https://devfeed.tech/topics/observability.md>)

Tags: [availability](<https://devfeed.tech/tags/availability.md>), [failure-flags](<https://devfeed.tech/tags/failure-flags.md>), [gremlin](<https://devfeed.tech/tags/gremlin.md>), [health-checks](<https://devfeed.tech/tags/health-checks.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [latency](<https://devfeed.tech/tags/latency.md>), [no-code](<https://devfeed.tech/tags/no-code.md>), [observability](<https://devfeed.tech/tags/observability.md>), [performance](<https://devfeed.tech/tags/performance.md>), [serverless](<https://devfeed.tech/tags/serverless.md>), [testing](<https://devfeed.tech/tags/testing.md>), [uptime](<https://devfeed.tech/tags/uptime.md>)

### AI overview

Gremlin announces Failure Flags by proxy, a no-code application fault injection solution for serverless and managed applications. The sidecar proxy enables reliability tests such as simulating outages, adding latency, and generating exceptions without code changes. Intelligent Health Checks automatically monitor network throughput, latency, and error rate during tests.

### Source excerpt

Gremlin announces Failure Flags by proxy, a no-code application fault injection solution for serverless and managed applications. Learn more in our latest blog post.

## How Gremlin makes disaster recovery testing easier and faster

DevFeed: [How Gremlin makes disaster recovery testing easier and faster](<https://devfeed.tech/articles/how-gremlin-makes-disaster-recovery-testing-easier-and-faster-11593.md>)

Original publisher: [Read original article](<https://www.gremlin.com/blog/how-gremlin-makes-disaster-recovery-testing-easier-and-faster>)

Author: Gavin Cahill

Published: 2026-03-04T00:00:00Z

Content type: tutorial

Language: en

Sources: [Gremlin Blog](<https://devfeed.tech/sources/gremlin-blog.md>)

Topics: [Disaster Recovery](<https://devfeed.tech/topics/disaster-recovery.md>), [Resilience](<https://devfeed.tech/topics/resilience.md>), [Testing](<https://devfeed.tech/topics/testing.md>), [Cloud](<https://devfeed.tech/topics/cloud.md>)

Tags: [backup](<https://devfeed.tech/tags/backup.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [disaster-recovery](<https://devfeed.tech/tags/disaster-recovery.md>), [failover](<https://devfeed.tech/tags/failover.md>), [gremlin](<https://devfeed.tech/tags/gremlin.md>), [resilience](<https://devfeed.tech/tags/resilience.md>), [testing](<https://devfeed.tech/tags/testing.md>), [use-cases](<https://devfeed.tech/tags/use-cases.md>)

### AI overview

The article explains how Gremlin's Disaster Recovery Testing helps teams test disaster recovery plans by simulating failures such as zone evacuations, region failovers, and cloud-provider outages. It recommends establishing service baselines with test suites, running tests regularly, and repeating them to verify fixes.

### Source excerpt

Gremlin's Disaster Recovery Testing makes it easy to run zone evacuations, region failovers, and more for a fraction of the lift of traditional disaster recovery testing.

## Announcing Disaster Recovery Testing

DevFeed: [Announcing Disaster Recovery Testing](<https://devfeed.tech/articles/announcing-disaster-recovery-testing-11558.md>)

Original publisher: [Read original article](<https://www.gremlin.com/blog/announcing-disaster-recovery-testing>)

Author: Andre Newman

Published: 2026-02-03T00:00:00Z

Content type: release

Language: en

Sources: [Gremlin Blog](<https://devfeed.tech/sources/gremlin-blog.md>)

Topics: [Disaster Recovery](<https://devfeed.tech/topics/disaster-recovery.md>), [Testing](<https://devfeed.tech/topics/testing.md>), [Incident response](<https://devfeed.tech/topics/incident-response.md>), [Resilience](<https://devfeed.tech/topics/resilience.md>), [Availability](<https://devfeed.tech/topics/availability.md>), [Cloud](<https://devfeed.tech/topics/cloud.md>), [datacenter](<https://devfeed.tech/topics/datacenter.md>), [Amazon Web Services](<https://devfeed.tech/topics/aws.md>), [Azure](<https://devfeed.tech/topics/azure.md>)

Tags: [announcements](<https://devfeed.tech/tags/announcements.md>), [availability](<https://devfeed.tech/tags/availability.md>), [aws](<https://devfeed.tech/tags/aws.md>), [azure](<https://devfeed.tech/tags/azure.md>), [disaster-recovery](<https://devfeed.tech/tags/disaster-recovery.md>), [failover](<https://devfeed.tech/tags/failover.md>), [gremlin](<https://devfeed.tech/tags/gremlin.md>), [health-checks](<https://devfeed.tech/tags/health-checks.md>), [incident-response](<https://devfeed.tech/tags/incident-response.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [metrics](<https://devfeed.tech/tags/metrics.md>), [resilience](<https://devfeed.tech/tags/resilience.md>), [systems](<https://devfeed.tech/tags/systems.md>), [testing](<https://devfeed.tech/tags/testing.md>)

### AI overview

Gremlin announces Disaster Recovery Testing, a feature for running organization-wide zone, region, and datacenter-scale experiments. It helps teams validate failover, disaster recovery, and incident response processes, with health checks that can automatically halt tests when key metrics exceed defined SLA limits.

### Source excerpt

Gremlin announces Disaster Recovery Testing for validating region failover processes, disaster recovery plans, incident response procedures, and more.

## How to test application resiliency by simulating the Cloudflare December 2025 outage

DevFeed: [How to test application resiliency by simulating the Cloudflare December 2025 outage](<https://devfeed.tech/articles/how-to-test-application-resiliency-by-simulating-the-cloudflare-december-2025-outage-11636.md>)

Original publisher: [Read original article](<https://www.gremlin.com/blog/how-to-test-application-resiliency-by-simulating-the-cloudflare-december-2025-outage>)

Author: Gavin Cahill

Published: 2025-12-19T00:00:00Z

Content type: tutorial

Language: en

Sources: [Gremlin Blog](<https://devfeed.tech/sources/gremlin-blog.md>)

Topics: [Failure Flags](<https://devfeed.tech/topics/failure-flags.md>), [Resilience](<https://devfeed.tech/topics/resilience.md>), [incident](<https://devfeed.tech/topics/incident.md>), [Chaos Engineering](<https://devfeed.tech/topics/chaos-engineering.md>), [SDKs](<https://devfeed.tech/topics/sdks.md>), [JavaScript](<https://devfeed.tech/topics/javascript.md>), [Node.js](<https://devfeed.tech/topics/node-js.md>), [Python](<https://devfeed.tech/topics/python.md>), [Kubernetes](<https://devfeed.tech/topics/kubernetes.md>), [AWS Lambda](<https://devfeed.tech/topics/aws-lambda.md>), [Azure](<https://devfeed.tech/topics/azure.md>)

Tags: [2025](<https://devfeed.tech/tags/2025.md>), [after-the-retrospective](<https://devfeed.tech/tags/after-the-retrospective.md>), [applications](<https://devfeed.tech/tags/applications.md>), [aws](<https://devfeed.tech/tags/aws.md>), [aws-lambda](<https://devfeed.tech/tags/aws-lambda.md>), [azure](<https://devfeed.tech/tags/azure.md>), [c-sharp](<https://devfeed.tech/tags/c-sharp.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [cloudflare](<https://devfeed.tech/tags/cloudflare.md>), [code](<https://devfeed.tech/tags/code.md>), [failure-flags](<https://devfeed.tech/tags/failure-flags.md>), [gremlin](<https://devfeed.tech/tags/gremlin.md>), [how-to](<https://devfeed.tech/tags/how-to.md>), [http](<https://devfeed.tech/tags/http.md>), [incident](<https://devfeed.tech/tags/incident.md>), [internet-traffic](<https://devfeed.tech/tags/internet-traffic.md>), [javascript](<https://devfeed.tech/tags/javascript.md>), [kubernetes](<https://devfeed.tech/tags/kubernetes.md>), [node-js](<https://devfeed.tech/tags/node-js.md>), [outage](<https://devfeed.tech/tags/outage.md>), [python](<https://devfeed.tech/tags/python.md>)

### AI overview

This tutorial explains how to use Gremlin Failure Flags to simulate HTTP 500 errors associated with the December 2025 Cloudflare outage. It describes application-layer fault injection, the distinction from network-layer experiments, and supported deployment environments and SDK languages.

### Source excerpt

Use Gremlin Failure Flags to safely simulate 500 error codes and test resiliency to outages, like the December 5th Cloudflare incident.

## Gremlin Release Roundup 2025: Reliability across AI, on-prem, and applications

DevFeed: [Gremlin Release Roundup 2025: Reliability across AI, on-prem, and applications](<https://devfeed.tech/articles/gremlin-release-roundup-2025-reliability-across-ai-on-prem-and-applications-11684.md>)

Original publisher: [Read original article](<https://www.gremlin.com/blog/release-roundup-2025>)

Author: Andre Newman

Published: 2025-12-15T00:00:00Z

Content type: release

Language: en

Sources: [Gremlin Blog](<https://devfeed.tech/sources/gremlin-blog.md>)

Topics: [on-prem](<https://devfeed.tech/topics/on-prem.md>), [Failure Flags](<https://devfeed.tech/topics/failure-flags.md>), [experiments](<https://devfeed.tech/topics/experiments.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [MCP Server](<https://devfeed.tech/topics/mcp-server.md>), [telemetry](<https://devfeed.tech/topics/telemetry.md>), [SRE](<https://devfeed.tech/topics/sre.md>)

Tags: [2025](<https://devfeed.tech/tags/2025.md>), [ai](<https://devfeed.tech/tags/ai.md>), [analysis](<https://devfeed.tech/tags/analysis.md>), [experiments](<https://devfeed.tech/tags/experiments.md>), [failure-flags](<https://devfeed.tech/tags/failure-flags.md>), [features](<https://devfeed.tech/tags/features.md>), [gremlin](<https://devfeed.tech/tags/gremlin.md>), [mcp](<https://devfeed.tech/tags/mcp.md>), [mcp-server](<https://devfeed.tech/tags/mcp-server.md>), [no-code](<https://devfeed.tech/tags/no-code.md>), [on-prem](<https://devfeed.tech/tags/on-prem.md>), [outages](<https://devfeed.tech/tags/outages.md>), [release](<https://devfeed.tech/tags/release.md>), [releases](<https://devfeed.tech/tags/releases.md>), [sre](<https://devfeed.tech/tags/sre.md>), [telemetry](<https://devfeed.tech/tags/telemetry.md>)

### AI overview

Gremlin's 2025 release roundup describes improvements aimed at preventing outages and making reliability testing easier. Highlights include Reliability Intelligence for analyzing failures and recommending remediations, a Gremlin MCP server for querying environments through an LLM, new experiments and Failure Flags capabilities, expanded platform support, streamlined onboarding, and web UI refinements.

### Source excerpt

This year's release roundup covers our new on-prem offering, new Failure Flags features, intelligent analysis of failed experiments, and much more.

## How to use Gremlin's Reliability Report

DevFeed: [How to use Gremlin's Reliability Report](<https://devfeed.tech/articles/how-to-use-gremlin-s-reliability-report-11642.md>)

Original publisher: [Read original article](<https://www.gremlin.com/blog/how-to-use-gremlins-reliability-report>)

Author: Gavin Cahill

Published: 2025-12-12T00:00:00Z

Content type: article

Language: en

Sources: [Gremlin Blog](<https://devfeed.tech/sources/gremlin-blog.md>)

Topics: [dashboards](<https://devfeed.tech/topics/dashboards.md>), [monitor](<https://devfeed.tech/topics/monitor.md>), [systems](<https://devfeed.tech/topics/systems.md>), [Cloud](<https://devfeed.tech/topics/cloud.md>), [Kubernetes](<https://devfeed.tech/topics/kubernetes.md>), [configuration](<https://devfeed.tech/topics/configuration.md>)

Tags: [blog](<https://devfeed.tech/tags/blog.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [configuration](<https://devfeed.tech/tags/configuration.md>), [gremlin](<https://devfeed.tech/tags/gremlin.md>), [how-to](<https://devfeed.tech/tags/how-to.md>), [kubernetes](<https://devfeed.tech/tags/kubernetes.md>), [leadership](<https://devfeed.tech/tags/leadership.md>), [monitor](<https://devfeed.tech/tags/monitor.md>), [organizational](<https://devfeed.tech/tags/organizational.md>), [reliability-management](<https://devfeed.tech/tags/reliability-management.md>)

### AI overview

Gremlin's Reliability Report provides organization-wide visibility into system reliability through reliability scores, detected risks, test-run counts, and service-level impacts. The article explains the report's dashboard sections, including six-month reliability trends and automatically detected Kubernetes and cloud risks, and describes how leadership can use the information to monitor and improve reliability.

### Source excerpt

Find out how our Reliability Report gives you visibility into your system's reliability--and how Gremlin uses it to improve reliability.

## Gremlin's unofficial Microsoft Ignite 2025 reliability track

DevFeed: [Gremlin's unofficial Microsoft Ignite 2025 reliability track](<https://devfeed.tech/articles/gremlin-s-unofficial-microsoft-ignite-2025-reliability-track-11580.md>)

Original publisher: [Read original article](<https://www.gremlin.com/blog/gremlins-unofficial-microsoft-ignite-2025-reliability-track>)

Author: Gavin Cahill

Published: 2025-11-12T00:00:00Z

Content type: article

Language: en

Sources: [Gremlin Blog](<https://devfeed.tech/sources/gremlin-blog.md>)

Topics: [Azure](<https://devfeed.tech/topics/azure.md>), [Microsoft](<https://devfeed.tech/topics/microsoft.md>), [Resilience](<https://devfeed.tech/topics/resilience.md>), [Cloud](<https://devfeed.tech/topics/cloud.md>), [configuration](<https://devfeed.tech/topics/configuration.md>), [autoscaling](<https://devfeed.tech/topics/autoscaling.md>), [Replication](<https://devfeed.tech/topics/replication.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>)

Tags: [2025](<https://devfeed.tech/tags/2025.md>), [ai](<https://devfeed.tech/tags/ai.md>), [autoscaling](<https://devfeed.tech/tags/autoscaling.md>), [azure](<https://devfeed.tech/tags/azure.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [cloud-native](<https://devfeed.tech/tags/cloud-native.md>), [conference](<https://devfeed.tech/tags/conference.md>), [configuration](<https://devfeed.tech/tags/configuration.md>), [gremlin](<https://devfeed.tech/tags/gremlin.md>), [industry](<https://devfeed.tech/tags/industry.md>), [microsoft](<https://devfeed.tech/tags/microsoft.md>), [recovery](<https://devfeed.tech/tags/recovery.md>), [replication](<https://devfeed.tech/tags/replication.md>), [resilience](<https://devfeed.tech/tags/resilience.md>), [secure-by-default](<https://devfeed.tech/tags/secure-by-default.md>), [shared-responsibility](<https://devfeed.tech/tags/shared-responsibility.md>), [storage](<https://devfeed.tech/tags/storage.md>)

### AI overview

Gremlin's unofficial Microsoft Ignite 2025 reliability track curates conference sessions about building resilient systems with Microsoft and Azure tools. The sessions cover reliability assessment and configuration, autoscaling and replication, failure simulation, recovery orchestration, shared responsibility, and secure backup practices for cloud and AI-ready applications.

### Source excerpt

Check out the Gremlin-curated unofficial track of reliability talks at Microsoft Ignite 2025.

## Improve Kubernetes reliability faster with Gremlin and Dynatrace

DevFeed: [Improve Kubernetes reliability faster with Gremlin and Dynatrace](<https://devfeed.tech/articles/improve-kubernetes-reliability-faster-with-gremlin-and-dynatrace-11649.md>)

Original publisher: [Read original article](<https://www.gremlin.com/blog/improve-kubernetes-reliability-faster-with-gremlin-and-dynatrace>)

Author: Gavin Cahill

Published: 2025-11-10T00:00:00Z

Content type: article

Language: en

Sources: [Gremlin Blog](<https://devfeed.tech/sources/gremlin-blog.md>)

Topics: [Kubernetes](<https://devfeed.tech/topics/kubernetes.md>), [dynatrace](<https://devfeed.tech/topics/dynatrace.md>), [observability](<https://devfeed.tech/topics/observability.md>), [Chaos Engineering](<https://devfeed.tech/topics/chaos-engineering.md>), [ai observability](<https://devfeed.tech/topics/ai-observability.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [announcements](<https://devfeed.tech/tags/announcements.md>), [cloud-native](<https://devfeed.tech/tags/cloud-native.md>), [dynatrace](<https://devfeed.tech/tags/dynatrace.md>), [gremlin](<https://devfeed.tech/tags/gremlin.md>), [integration](<https://devfeed.tech/tags/integration.md>), [kubernetes](<https://devfeed.tech/tags/kubernetes.md>), [observability](<https://devfeed.tech/tags/observability.md>), [testing](<https://devfeed.tech/tags/testing.md>)

### AI overview

The article presents a strategic integration between Gremlin and Dynatrace that automatically discovers Kubernetes services configured in Dynatrace, making fault-injection testing faster to set up. It explains how health checks use observability metrics such as error rates and latency to evaluate experiments, validate monitoring and alerting, and improve reliability across cloud-native architectures.

### Source excerpt

It's easier than ever to start testing Kubernetes with Dynatrace and Gremlin. The new strategic integration automatically discovers objects to make testing set up simple and fast.

## How to test the reliability of a Point of Sale (POS) system

DevFeed: [How to test the reliability of a Point of Sale (POS) system](<https://devfeed.tech/articles/how-to-test-the-reliability-of-a-point-of-sale-pos-system-11640.md>)

Original publisher: [Read original article](<https://www.gremlin.com/blog/how-to-test-the-reliability-of-a-point-of-sale-pos-system>)

Author: Gavin Cahill

Published: 2025-10-20T00:00:00Z

Content type: tutorial

Language: en

Sources: [Gremlin Blog](<https://devfeed.tech/sources/gremlin-blog.md>)

Topics: [Chaos Engineering](<https://devfeed.tech/topics/chaos-engineering.md>), [Resilience](<https://devfeed.tech/topics/resilience.md>), [autoscaling](<https://devfeed.tech/topics/autoscaling.md>), [Microservice](<https://devfeed.tech/topics/microservice.md>), [Complex Systems](<https://devfeed.tech/topics/complex-systems.md>), [systems](<https://devfeed.tech/topics/systems.md>)

Tags: [autoscaling](<https://devfeed.tech/tags/autoscaling.md>), [best-practices](<https://devfeed.tech/tags/best-practices.md>), [chaos-engineering](<https://devfeed.tech/tags/chaos-engineering.md>), [cpu](<https://devfeed.tech/tags/cpu.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [gremlin](<https://devfeed.tech/tags/gremlin.md>), [how-to](<https://devfeed.tech/tags/how-to.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [memory](<https://devfeed.tech/tags/memory.md>), [microservices](<https://devfeed.tech/tags/microservices.md>), [outage](<https://devfeed.tech/tags/outage.md>), [reliability-management](<https://devfeed.tech/tags/reliability-management.md>), [resilience](<https://devfeed.tech/tags/resilience.md>), [retail](<https://devfeed.tech/tags/retail.md>), [testing](<https://devfeed.tech/tags/testing.md>)

### AI overview

This tutorial explains how to test the reliability of retail Point of Sale systems using Gremlin and Chaos Engineering. It focuses on resilience testing for microservice-based checkout systems, including autoscaling, CPU, memory, and disk I/O capacity, to identify failure conditions and reduce outages.

### Source excerpt

Find out how to use Gremlin and Chaos Engineering to make sure your Point of Sale system is reliable.

## Three practices for improving service availability toward 99.999% uptime

DevFeed: [Three practices for improving service availability toward 99.999% uptime](<https://devfeed.tech/articles/3-things-you-can-do-to-get-closer-to-five-nines-11556.md>)

Original publisher: [Read original article](<https://www.gremlin.com/blog/3-things-you-can-do-to-get-closer-to-five-nines>)

Author: Andre Newman

Published: 2025-10-02T00:00:00Z

Content type: tutorial

Language: en

Sources: [Gremlin Blog](<https://devfeed.tech/sources/gremlin-blog.md>)

Topics: [Testing](<https://devfeed.tech/topics/testing.md>), [Availability](<https://devfeed.tech/topics/availability.md>), [Resilience](<https://devfeed.tech/topics/resilience.md>), [observability](<https://devfeed.tech/topics/observability.md>), [Software as a service](<https://devfeed.tech/topics/saas.md>)

Tags: [automated](<https://devfeed.tech/tags/automated.md>), [availability](<https://devfeed.tech/tags/availability.md>), [gremlin](<https://devfeed.tech/tags/gremlin.md>), [health-checks](<https://devfeed.tech/tags/health-checks.md>), [observability](<https://devfeed.tech/tags/observability.md>), [resilience](<https://devfeed.tech/tags/resilience.md>), [sre](<https://devfeed.tech/tags/sre.md>), [testing](<https://devfeed.tech/tags/testing.md>)

### AI overview

The article presents three practices for improving service availability toward 99.999% uptime: regular reliability testing, on-call rotations, and better visibility. It describes Gremlin's approach, including weekly tests, Health Checks, and observability metrics.

### Source excerpt

3 proven practices to achieve 99.999% uptime. Learn how teams reach five nines through weekly testing, on-call rotations, and visibility.

## How to get fast, easy insights with the Gremlin MCP Server

DevFeed: [How to get fast, easy insights with the Gremlin MCP Server](<https://devfeed.tech/articles/how-to-get-fast-easy-insights-with-the-gremlin-mcp-server-11620.md>)

Original publisher: [Read original article](<https://www.gremlin.com/blog/how-to-get-fast-easy-insights-with-the-gremlin-mcp-server>)

Author: Gavin Cahill

Published: 2025-08-28T00:00:00Z

Content type: article

Language: en

Sources: [Gremlin Blog](<https://devfeed.tech/sources/gremlin-blog.md>)

Topics: [MCP Server](<https://devfeed.tech/topics/mcp-server.md>), [Model Context Protocol](<https://devfeed.tech/topics/model-context-protocol.md>), [Chaos Engineering](<https://devfeed.tech/topics/chaos-engineering.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [test-coverage](<https://devfeed.tech/topics/test-coverage.md>), [Authorization](<https://devfeed.tech/topics/authorization.md>), [Containers](<https://devfeed.tech/topics/containers.md>), [dashboards](<https://devfeed.tech/topics/dashboards.md>), [Security](<https://devfeed.tech/topics/security.md>), [API keys](<https://devfeed.tech/topics/api-keys.md>)

Tags: [access-control](<https://devfeed.tech/tags/access-control.md>), [ai](<https://devfeed.tech/tags/ai.md>), [api](<https://devfeed.tech/tags/api.md>), [api-keys](<https://devfeed.tech/tags/api-keys.md>), [applications](<https://devfeed.tech/tags/applications.md>), [chaos-engineering](<https://devfeed.tech/tags/chaos-engineering.md>), [chatgpt](<https://devfeed.tech/tags/chatgpt.md>), [claude](<https://devfeed.tech/tags/claude.md>), [components](<https://devfeed.tech/tags/components.md>), [core](<https://devfeed.tech/tags/core.md>), [data](<https://devfeed.tech/tags/data.md>), [gremlin](<https://devfeed.tech/tags/gremlin.md>), [hardening](<https://devfeed.tech/tags/hardening.md>), [how-to](<https://devfeed.tech/tags/how-to.md>), [insights](<https://devfeed.tech/tags/insights.md>), [installation](<https://devfeed.tech/tags/installation.md>), [llm](<https://devfeed.tech/tags/llm.md>), [management](<https://devfeed.tech/tags/management.md>), [mcp](<https://devfeed.tech/tags/mcp.md>), [mcp-server](<https://devfeed.tech/tags/mcp-server.md>)

### AI overview

The article explains how the Gremlin MCP Server connects an LLM client to the Gremlin API so teams can explore reliability-testing data and identify insights using natural-language prompts. It covers the client-server architecture, containerized deployment, security hardening, non-destructive API operations, dashboards, reporting, evaluation, and RBAC-based control of API keys.

### Source excerpt

Find out how to quickly and easily uncover new reliability insights by using the Gremlin MCP Server and your favorite LLM.

## Fix issues faster with Recommended Remediations

DevFeed: [Fix issues faster with Recommended Remediations](<https://devfeed.tech/articles/fix-issues-faster-with-recommended-remediations-11569.md>)

Original publisher: [Read original article](<https://www.gremlin.com/blog/fix-issues-faster-with-recommended-remediations>)

Author: Gavin Cahill

Published: 2025-08-22T00:00:00Z

Content type: release

Language: en

Sources: [Gremlin Blog](<https://devfeed.tech/sources/gremlin-blog.md>)

Topics: [Chaos Engineering](<https://devfeed.tech/topics/chaos-engineering.md>), [Testing](<https://devfeed.tech/topics/testing.md>)

Tags: [chaos-engineering](<https://devfeed.tech/tags/chaos-engineering.md>), [features](<https://devfeed.tech/tags/features.md>), [gremlin](<https://devfeed.tech/tags/gremlin.md>), [recommendations](<https://devfeed.tech/tags/recommendations.md>), [uptime](<https://devfeed.tech/tags/uptime.md>)

### AI overview

Gremlin's Recommended Remediation analyzes fault-injection test results to identify likely failure causes and suggest ways to address reliability issues. It builds on Experiment Analysis by combining test data, metrics, health checks, and key events, then applies reliability expertise to produce tailored recommendations.

### Source excerpt

Recommended Remediation speeds up teams with tailored suggestions to help you address reliability risks before they cause failures.

## How Experiment Analysis uncovers the cause behind failures

DevFeed: [How Experiment Analysis uncovers the cause behind failures](<https://devfeed.tech/articles/how-experiment-analysis-uncovers-the-cause-behind-failures-11589.md>)

Original publisher: [Read original article](<https://www.gremlin.com/blog/how-experiment-analysis-uncovers-the-cause-behind-failures>)

Author: Gavin Cahill

Published: 2025-08-15T00:00:00Z

Content type: article

Language: en

Sources: [Gremlin Blog](<https://devfeed.tech/sources/gremlin-blog.md>)

Topics: [experiments](<https://devfeed.tech/topics/experiments.md>), [Chaos Engineering](<https://devfeed.tech/topics/chaos-engineering.md>), [observability](<https://devfeed.tech/topics/observability.md>), [Machine learning](<https://devfeed.tech/topics/machine-learning.md>), [Availability](<https://devfeed.tech/topics/availability.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [analysis](<https://devfeed.tech/tags/analysis.md>), [availability](<https://devfeed.tech/tags/availability.md>), [chaos-engineering](<https://devfeed.tech/tags/chaos-engineering.md>), [experiment](<https://devfeed.tech/tags/experiment.md>), [failover](<https://devfeed.tech/tags/failover.md>), [gremlin](<https://devfeed.tech/tags/gremlin.md>), [observability](<https://devfeed.tech/tags/observability.md>), [testing](<https://devfeed.tech/tags/testing.md>)

### AI overview

The article explains how Gremlin's Experiment Analysis uses machine learning and observability context to identify why Chaos Engineering tests fail, reducing the manual effort needed to investigate system behavior. It also emphasizes defining expected behavior as the basis for pass/fail decisions, illustrated by a zone failover scenario.

### Source excerpt

Using knowledge of your data and your systems, Experiment Analysis connects the cause with the effect to help you pinpoint the root problem faster.

## Measure your reliability risk, not your engineers

DevFeed: [Measure your reliability risk, not your engineers](<https://devfeed.tech/articles/measure-your-reliability-risk-not-your-engineers-11669.md>)

Original publisher: [Read original article](<https://www.gremlin.com/blog/measure-your-reliability-risk-not-your-engineers>)

Author: Gavin Cahill

Published: 2025-07-23T00:00:00Z

Content type: article

Language: en

Sources: [Gremlin Blog](<https://devfeed.tech/sources/gremlin-blog.md>)

Topics: [Resilience](<https://devfeed.tech/topics/resilience.md>), [site-reliability-engineering](<https://devfeed.tech/topics/site-reliability-engineering.md>), [systems](<https://devfeed.tech/topics/systems.md>)

Tags: [gremlin](<https://devfeed.tech/tags/gremlin.md>), [metrics](<https://devfeed.tech/tags/metrics.md>), [reliability-management](<https://devfeed.tech/tags/reliability-management.md>), [resilience](<https://devfeed.tech/tags/resilience.md>), [reviews](<https://devfeed.tech/tags/reviews.md>), [testing](<https://devfeed.tech/tags/testing.md>)

### AI overview

The article argues that organizations should measure system reliability risk rather than rely on engineer skill or QA testing. It presents resilience testing and Reliability Scores as ways to track current risk, identify failures, and support regular improvement.

### Source excerpt

Reliability metrics should uncover risks and enable your teams to improve reliability, not create defensiveness and blame games.

## Insights to keep AI applications reliable

DevFeed: [Insights to keep AI applications reliable](<https://devfeed.tech/articles/insights-to-keep-ai-applications-reliable-11653.md>)

Original publisher: [Read original article](<https://www.gremlin.com/blog/insights-to-keep-ai-applications-reliable>)

Author: Gavin Cahill

Published: 2025-06-23T00:00:00Z

Content type: article

Language: en

Sources: [Gremlin Blog](<https://devfeed.tech/sources/gremlin-blog.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [App](<https://devfeed.tech/topics/app.md>), [observability](<https://devfeed.tech/topics/observability.md>), [Resilience](<https://devfeed.tech/topics/resilience.md>), [Incident response](<https://devfeed.tech/topics/incident-response.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [GPU](<https://devfeed.tech/topics/gpu.md>), [Training AI Models](<https://devfeed.tech/topics/training-ai-models.md>), [Software as a service](<https://devfeed.tech/topics/saas.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [applications](<https://devfeed.tech/tags/applications.md>), [best-practices](<https://devfeed.tech/tags/best-practices.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [gremlin](<https://devfeed.tech/tags/gremlin.md>), [incident-response](<https://devfeed.tech/tags/incident-response.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [insights](<https://devfeed.tech/tags/insights.md>), [llm](<https://devfeed.tech/tags/llm.md>), [metrics](<https://devfeed.tech/tags/metrics.md>), [observability](<https://devfeed.tech/tags/observability.md>), [operations](<https://devfeed.tech/tags/operations.md>), [resilience](<https://devfeed.tech/tags/resilience.md>), [saas](<https://devfeed.tech/tags/saas.md>), [testing](<https://devfeed.tech/tags/testing.md>), [training](<https://devfeed.tech/tags/training.md>)

### AI overview

The article explains how engineering teams can keep AI applications reliable. It recommends applying established practices such as observability, resilience testing, service-level objectives, metrics, and incident response while accounting for AI-specific traffic patterns, infrastructure complexity, dependencies, SaaS boundaries, model training, batch processing, and GPU resource surges.

### Source excerpt

AI has become a massive investment for companies, but how do you keep AI applications reliable? Check out these insights from Gremlin, Nobl9, and Pagerduty to find out!

## Test serverless and application-level reliability with Failure Flags

DevFeed: [Test serverless and application-level reliability with Failure Flags](<https://devfeed.tech/articles/test-serverless-and-application-level-reliability-with-failure-flags-11715.md>)

Original publisher: [Read original article](<https://www.gremlin.com/blog/test-serverless-and-application-level-reliability-with-failure-flags>)

Author: Gavin Cahill

Published: 2025-03-13T00:00:00Z

Content type: article

Language: en

Sources: [Gremlin Blog](<https://devfeed.tech/sources/gremlin-blog.md>)

Topics: [Failure Flags](<https://devfeed.tech/topics/failure-flags.md>), [Resilience](<https://devfeed.tech/topics/resilience.md>), [Serverless](<https://devfeed.tech/topics/serverless.md>), [Containers](<https://devfeed.tech/topics/containers.md>), [Kubernetes](<https://devfeed.tech/topics/kubernetes.md>), [SDKs](<https://devfeed.tech/topics/sdks.md>), [AWS Lambda](<https://devfeed.tech/topics/aws-lambda.md>), [Node.js](<https://devfeed.tech/topics/node-js.md>), [API](<https://devfeed.tech/topics/api.md>), [Python](<https://devfeed.tech/topics/python.md>), [airflow](<https://devfeed.tech/topics/airflow.md>)

Tags: [api](<https://devfeed.tech/tags/api.md>), [aws-lambda](<https://devfeed.tech/tags/aws-lambda.md>), [configuration](<https://devfeed.tech/tags/configuration.md>), [failure-flags](<https://devfeed.tech/tags/failure-flags.md>), [go](<https://devfeed.tech/tags/go.md>), [gremlin](<https://devfeed.tech/tags/gremlin.md>), [java](<https://devfeed.tech/tags/java.md>), [kubernetes](<https://devfeed.tech/tags/kubernetes.md>), [latency](<https://devfeed.tech/tags/latency.md>), [mesh](<https://devfeed.tech/tags/mesh.md>), [net](<https://devfeed.tech/tags/net.md>), [node-js](<https://devfeed.tech/tags/node-js.md>), [python](<https://devfeed.tech/tags/python.md>), [resilience](<https://devfeed.tech/tags/resilience.md>), [sdk](<https://devfeed.tech/tags/sdk.md>), [serverless](<https://devfeed.tech/tags/serverless.md>)

### AI overview

Gremlin Failure Flags is presented as a generally available tool for testing application-level resilience in serverless, container, Kubernetes, and service mesh environments. It combines the Gremlin SaaS API with a sidecar or Lambda Extension and an SDK to safely introduce latency, errors, and selected data issues, while also validating observability, alerting, and automated recovery.

### Source excerpt

Run resilience tests at the application level in serverless, container, Kubernetes, and service mesh environments with Gremlin Failure Flags.

## How a major retailer tested critical serverless systems with Failure Flags

DevFeed: [How a major retailer tested critical serverless systems with Failure Flags](<https://devfeed.tech/articles/how-a-major-retailer-tested-critical-serverless-systems-with-failure-flags-11586.md>)

Original publisher: [Read original article](<https://www.gremlin.com/blog/how-a-major-retailer-tested-critical-serverless-systems-with-failure-flags>)

Author: Gavin Cahill

Published: 2025-03-12T00:00:00Z

Content type: article

Language: en

Sources: [Gremlin Blog](<https://devfeed.tech/sources/gremlin-blog.md>)

Topics: [Failure Flags](<https://devfeed.tech/topics/failure-flags.md>), [AWS Lambda](<https://devfeed.tech/topics/aws-lambda.md>), [Serverless](<https://devfeed.tech/topics/serverless.md>), [systems](<https://devfeed.tech/topics/systems.md>), [Amazon EC2](<https://devfeed.tech/topics/amazon-ec2.md>), [Amazon Web Services](<https://devfeed.tech/topics/aws.md>)

Tags: [amazon-ec2](<https://devfeed.tech/tags/amazon-ec2.md>), [aws-lambda](<https://devfeed.tech/tags/aws-lambda.md>), [failover](<https://devfeed.tech/tags/failover.md>), [failure-flags](<https://devfeed.tech/tags/failure-flags.md>), [gremlin](<https://devfeed.tech/tags/gremlin.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [outage](<https://devfeed.tech/tags/outage.md>), [serverless](<https://devfeed.tech/tags/serverless.md>), [testing](<https://devfeed.tech/tags/testing.md>)

### AI overview

A case study describes how a major apparel retailer used Gremlin Failure Flags to test regional failover for a critical payment service built on AWS Lambda. Because Lambda abstracts the underlying infrastructure, conventional infrastructure-level fault injection could not reproduce a regional failure. Failure Flags enabled the retailer to run the failover test in less than 30 minutes.

### Source excerpt

Find out how Gremlin helped a major retailer test region failover for a critical service built on AWS Lambda using Failure Flags.

## Simulating artificial intelligence (AI) service outages with Gremlin

DevFeed: [Simulating artificial intelligence (AI) service outages with Gremlin](<https://devfeed.tech/articles/simulating-artificial-intelligence-ai-service-outages-with-gremlin-11709.md>)

Original publisher: [Read original article](<https://www.gremlin.com/blog/simulating-artificial-intelligence-service-outages-with-gremlin>)

Author: Andre Newman

Published: 2025-03-06T00:00:00Z

Content type: tutorial

Language: en

Sources: [Gremlin Blog](<https://devfeed.tech/sources/gremlin-blog.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [AI Development](<https://devfeed.tech/topics/ai-development.md>), [Network](<https://devfeed.tech/topics/network.md>), [API](<https://devfeed.tech/topics/api.md>), [Latency](<https://devfeed.tech/topics/latency.md>), [SDKs](<https://devfeed.tech/topics/sdks.md>), [Software](<https://devfeed.tech/topics/software.md>), [Amazon Web Services](<https://devfeed.tech/topics/aws.md>), [TLS (Transport Layer Security)](<https://devfeed.tech/topics/tls.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [api](<https://devfeed.tech/tags/api.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [gremlin](<https://devfeed.tech/tags/gremlin.md>), [how-to](<https://devfeed.tech/tags/how-to.md>), [latency](<https://devfeed.tech/tags/latency.md>), [network](<https://devfeed.tech/tags/network.md>), [outages](<https://devfeed.tech/tags/outages.md>), [sdks](<https://devfeed.tech/tags/sdks.md>), [tls](<https://devfeed.tech/tags/tls.md>)

### AI overview

This article explains how AI-as-a-Service providers host AI models for applications and expose them through web interfaces, APIs, and SDKs. It describes provider outages, dependency failures, network latency, and certificate expiration as risks that can affect application reliability, and discusses using Gremlin to simulate these failures and verify resilience.

### Source excerpt

Learn how to leverage artificial intelligence (AI) services while avoiding downtime caused by outages.

## Announcing Gremlin Private Edition

DevFeed: [Announcing Gremlin Private Edition](<https://devfeed.tech/articles/announcing-gremlin-private-edition-11560.md>)

Original publisher: [Read original article](<https://www.gremlin.com/blog/announcing-gremlin-private-edition>)

Author: Andre Newman

Published: 2025-02-11T00:00:00Z

Content type: article

Language: en

Sources: [Gremlin Blog](<https://devfeed.tech/sources/gremlin-blog.md>)

Topics: [Chaos Engineering](<https://devfeed.tech/topics/chaos-engineering.md>), [Kubernetes](<https://devfeed.tech/topics/kubernetes.md>), [Security](<https://devfeed.tech/topics/security.md>), [Software as a service](<https://devfeed.tech/topics/saas.md>), [datacenter](<https://devfeed.tech/topics/datacenter.md>)

Tags: [announcement](<https://devfeed.tech/tags/announcement.md>), [announcements](<https://devfeed.tech/tags/announcements.md>), [blog](<https://devfeed.tech/tags/blog.md>), [gremlin](<https://devfeed.tech/tags/gremlin.md>), [kubernetes](<https://devfeed.tech/tags/kubernetes.md>), [monitoring](<https://devfeed.tech/tags/monitoring.md>), [security](<https://devfeed.tech/tags/security.md>), [self-hosted](<https://devfeed.tech/tags/self-hosted.md>), [testing](<https://devfeed.tech/tags/testing.md>)

### AI overview

Gremlin Private Edition is a self-hosted version of Gremlin that runs entirely within an organization's private network. It provides the SaaS product's reliability-testing features while allowing organizations to control deployment, access, and data storage. The platform runs on Kubernetes and supports private, cloud, datacenter, and air-gapped environments.

### Source excerpt

Gremlin Private Edition is a private, secure, self-hosted Gremlin instance that runs entirely within your network. Learn more in our announcement blog.

## How the Gremlin agent fails safely

DevFeed: [How the Gremlin agent fails safely](<https://devfeed.tech/articles/how-the-gremlin-agent-fails-safely-11603.md>)

Original publisher: [Read original article](<https://www.gremlin.com/blog/how-the-gremlin-agent-fails-safely>)

Author: Andre Newman

Published: 2025-01-30T00:00:00Z

Content type: article

Language: en

Sources: [Gremlin Blog](<https://devfeed.tech/sources/gremlin-blog.md>)

Topics: [Chaos Engineering](<https://devfeed.tech/topics/chaos-engineering.md>), [experiments](<https://devfeed.tech/topics/experiments.md>), [Integration testing](<https://devfeed.tech/topics/integration-testing.md>), [Linux](<https://devfeed.tech/topics/linux.md>), [Network](<https://devfeed.tech/topics/network.md>)

Tags: [architecture](<https://devfeed.tech/tags/architecture.md>), [chaos-engineering](<https://devfeed.tech/tags/chaos-engineering.md>), [data-center](<https://devfeed.tech/tags/data-center.md>), [gremlin](<https://devfeed.tech/tags/gremlin.md>), [integration](<https://devfeed.tech/tags/integration.md>), [linux](<https://devfeed.tech/tags/linux.md>), [network](<https://devfeed.tech/tags/network.md>), [server](<https://devfeed.tech/tags/server.md>), [testing](<https://devfeed.tech/tags/testing.md>)

### AI overview

The article explains how the Gremlin agent makes Chaos Engineering and reliability testing safer. Agents periodically check in with the Gremlin Control Plane and automatically stop active experiments and restore systems if communication is lost.

### Source excerpt

Reliability testing shouldn't feel risky. Learn how Gremlin makes testing safer with fail-safe agents and automatic rollbacks.

## How to fix the root cause of a failed reliability test

DevFeed: [How to fix the root cause of a failed reliability test](<https://devfeed.tech/articles/how-to-fix-the-root-cause-of-a-failed-reliability-test-11617.md>)

Original publisher: [Read original article](<https://www.gremlin.com/blog/how-to-fix-the-root-cause-of-a-failed-reliability-test>)

Author: Andre Newman

Published: 2025-01-21T00:00:00Z

Content type: tutorial

Language: en

Sources: [Gremlin Blog](<https://devfeed.tech/sources/gremlin-blog.md>)

Topics: [Resilience](<https://devfeed.tech/topics/resilience.md>), [Scalability](<https://devfeed.tech/topics/scalability.md>), [Availability](<https://devfeed.tech/topics/availability.md>), [cpu](<https://devfeed.tech/topics/cpu.md>), [Latency](<https://devfeed.tech/topics/latency.md>), [Network](<https://devfeed.tech/topics/network.md>), [observability](<https://devfeed.tech/topics/observability.md>), [TLS (Transport Layer Security)](<https://devfeed.tech/topics/tls.md>), [systems](<https://devfeed.tech/topics/systems.md>)

Tags: [availability](<https://devfeed.tech/tags/availability.md>), [blog](<https://devfeed.tech/tags/blog.md>), [chaos-engineering](<https://devfeed.tech/tags/chaos-engineering.md>), [cpu](<https://devfeed.tech/tags/cpu.md>), [gremlin](<https://devfeed.tech/tags/gremlin.md>), [how-to](<https://devfeed.tech/tags/how-to.md>), [latency](<https://devfeed.tech/tags/latency.md>), [metrics](<https://devfeed.tech/tags/metrics.md>), [network](<https://devfeed.tech/tags/network.md>), [resilience](<https://devfeed.tech/tags/resilience.md>), [scalability](<https://devfeed.tech/tags/scalability.md>), [systems](<https://devfeed.tech/tags/systems.md>), [tls](<https://devfeed.tech/tags/tls.md>)

### AI overview

This tutorial explains how to investigate failed Gremlin reliability tests and use their results to improve system resilience. It describes the Well-Architected Cloud Test Suite, covering scalability, redundancy, dependency failures, latency, and TLS certificate expiry, along with the role of Health Checks and observability metrics and alerts.

### Source excerpt

You've run your reliability tests, and unfortunately, some of them failed. No need to panic: we'll tell you how to turn that F into an A.

## Manage your reliability work more easily with Gremlin's newest features

DevFeed: [Manage your reliability work more easily with Gremlin's newest features](<https://devfeed.tech/articles/manage-your-reliability-work-more-easily-with-gremlin-s-newest-features-11661.md>)

Original publisher: [Read original article](<https://www.gremlin.com/blog/manage-your-reliability-work-more-easily-with-gremlins-newest-features>)

Author: Andre Newman

Published: 2025-01-06T00:00:00Z

Content type: release

Language: en

Sources: [Gremlin Blog](<https://devfeed.tech/sources/gremlin-blog.md>)

Topics: [monitor](<https://devfeed.tech/topics/monitor.md>), [Testing](<https://devfeed.tech/topics/testing.md>), [experiments](<https://devfeed.tech/topics/experiments.md>), [Failure Flags](<https://devfeed.tech/topics/failure-flags.md>), [CI/CD](<https://devfeed.tech/topics/cicd.md>)

Tags: [ci-cd](<https://devfeed.tech/tags/ci-cd.md>), [features](<https://devfeed.tech/tags/features.md>), [gremlin](<https://devfeed.tech/tags/gremlin.md>), [latest-release](<https://devfeed.tech/tags/latest-release.md>), [metrics](<https://devfeed.tech/tags/metrics.md>), [release](<https://devfeed.tech/tags/release.md>)

### AI overview

Gremlin announces three new screens--Now Running, What's Scheduled, and What Ran--to help teams monitor and manage ongoing, scheduled, and past reliability activities. The screens cover experiments, Scenarios, reliability tests, and Failure Flags.

### Source excerpt

Gremlin has three new ways to help you monitor and manage your reliability efforts. See how to track current, future, and past activity in Gremlin.

[Next page](<https://devfeed.tech/tags/gremlin.md?cursor=WyIyMDI1LTAxLTA2VDAwOjAwOjAwKzAwOjAwIiwgIjE2YTU1ZjhmLWYyYzMtNDg5OC05MDA5LWViYjNjMzdmNDcxMSJd>)