# chaos

Published articles for chaos.

This is one page of public article previews, not the complete archive. Follow Next page to continue. Summaries are not the original full articles.

## Run contained pen testing as chaos experiments securely

DevFeed: [Run contained pen testing as chaos experiments securely](<https://devfeed.tech/articles/run-contained-pen-testing-as-chaos-experiments-securely-13490.md>)

Original publisher: [Read original article](<https://www.harness.io/blog/the-contained-pen-test-proving-resilience-without-widening-the-blast-radius>)

Author: Uma Mukkara

Published: 2026-09-09T22:42:00Z

Content type: article

Language: en

Sources: [Harness Blog](<https://devfeed.tech/sources/harness-blog.md>)

Topics: [Resilience](<https://devfeed.tech/topics/resilience.md>), [Security](<https://devfeed.tech/topics/security.md>), [Testing](<https://devfeed.tech/topics/testing.md>)

Tags: [chaos](<https://devfeed.tech/tags/chaos.md>), [permission](<https://devfeed.tech/tags/permission.md>), [pipeline](<https://devfeed.tech/tags/pipeline.md>), [production](<https://devfeed.tech/tags/production.md>), [resilience](<https://devfeed.tech/tags/resilience.md>), [security](<https://devfeed.tech/tags/security.md>), [testing](<https://devfeed.tech/tags/testing.md>), [tool](<https://devfeed.tech/tags/tool.md>)

### AI overview

The article describes contained pen-testing experiments that run inside a release pipeline. Each experiment targets one service, one failure or intrusion condition, and one pipeline stage, producing an immutable audit record while avoiding broad production access.

### Source excerpt

| Blog

## Harness RT Agents Detect Resilience Risks and Generate Tests for CD Pipelines and Kubernetes Workloads

DevFeed: [Harness RT Agents Detect Resilience Risks and Generate Tests for CD Pipelines and Kubernetes Workloads](<https://devfeed.tech/articles/automate-resilience-testing-with-agents-13398.md>)

Original publisher: [Read original article](<https://www.harness.io/blog/find-resilience-risks-automatically-then-confirm-them>)

Author: Uma Mukkara

Published: 2026-08-24T00:00:00Z

Content type: release

Language: en

Sources: [Harness Blog](<https://devfeed.tech/sources/harness-blog.md>)

Topics: [Resilience](<https://devfeed.tech/topics/resilience.md>), [Testing](<https://devfeed.tech/topics/testing.md>), [Continuous Delivery (CD)](<https://devfeed.tech/topics/continuous-delivery.md>), [Kubernetes](<https://devfeed.tech/topics/kubernetes.md>), [dashboards](<https://devfeed.tech/topics/dashboards.md>)

Tags: [agentic](<https://devfeed.tech/tags/agentic.md>), [agents](<https://devfeed.tech/tags/agents.md>), [blog](<https://devfeed.tech/tags/blog.md>), [chaos](<https://devfeed.tech/tags/chaos.md>), [configuration](<https://devfeed.tech/tags/configuration.md>), [continuous-delivery](<https://devfeed.tech/tags/continuous-delivery.md>), [dashboards](<https://devfeed.tech/tags/dashboards.md>), [deployment](<https://devfeed.tech/tags/deployment.md>), [experiments](<https://devfeed.tech/tags/experiments.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [insights](<https://devfeed.tech/tags/insights.md>), [kubernetes](<https://devfeed.tech/tags/kubernetes.md>), [load](<https://devfeed.tech/tags/load.md>), [pipelines](<https://devfeed.tech/tags/pipelines.md>), [product](<https://devfeed.tech/tags/product.md>), [resilience](<https://devfeed.tech/tags/resilience.md>), [services](<https://devfeed.tech/tags/services.md>), [teams](<https://devfeed.tech/tags/teams.md>), [testing](<https://devfeed.tech/tags/testing.md>), [tests](<https://devfeed.tech/tags/tests.md>)

### AI overview

Harness announces an update to Resilience Testing called RT Agents. The agents analyze CD pipelines and Kubernetes workloads for resilience risks, recommend the testing needed to confirm those risks, and can generate and run chaos experiments or load tests and interpret the results.

### Source excerpt

RT Agents detect resilience risk in your CD pipelines and Kubernetes workloads, then generate and run chaos experiments or load tests to confirm it. | Blog

## Data Engineering Weekly #282

DevFeed: [Data Engineering Weekly #282](<https://devfeed.tech/articles/data-engineering-weekly-282-18262.md>)

Original publisher: [Read original article](<https://www.dataengineeringweekly.com/p/data-engineering-weekly-282>)

Author: Ananth Packkildurai

Published: 2026-08-10T01:21:26Z

Content type: article

Language: en

Sources: [Data Engineering Weekly](<https://devfeed.tech/sources/data-engineering-weekly.md>)

Topics: [data-engineering](<https://devfeed.tech/topics/data-engineering.md>), [data-platforms](<https://devfeed.tech/topics/data-platforms.md>), [data observability](<https://devfeed.tech/topics/data-observability.md>), [Graphs](<https://devfeed.tech/topics/graphs.md>), [semantic-layer](<https://devfeed.tech/topics/semantic-layer.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [Apache Iceberg](<https://devfeed.tech/topics/apache-iceberg.md>), [gRPC](<https://devfeed.tech/topics/grpc.md>), [Concurrency](<https://devfeed.tech/topics/concurrency.md>), [Caching](<https://devfeed.tech/topics/caching.md>), [Optimization](<https://devfeed.tech/topics/optimization.md>), [Streaming](<https://devfeed.tech/topics/streaming.md>)

Tags: [caching](<https://devfeed.tech/tags/caching.md>), [chaos](<https://devfeed.tech/tags/chaos.md>), [concurrency](<https://devfeed.tech/tags/concurrency.md>), [data](<https://devfeed.tech/tags/data.md>), [data-engineering](<https://devfeed.tech/tags/data-engineering.md>), [data-platforms](<https://devfeed.tech/tags/data-platforms.md>), [data-quality](<https://devfeed.tech/tags/data-quality.md>), [knowledge-graphs](<https://devfeed.tech/tags/knowledge-graphs.md>), [llms](<https://devfeed.tech/tags/llms.md>), [newsletter](<https://devfeed.tech/tags/newsletter.md>), [observability](<https://devfeed.tech/tags/observability.md>), [pipelines](<https://devfeed.tech/tags/pipelines.md>), [semantic-layer](<https://devfeed.tech/tags/semantic-layer.md>), [streaming](<https://devfeed.tech/tags/streaming.md>)

### AI overview

Data Engineering Weekly #282 is a newsletter covering data platform fundamentals, semantic layers, ontology-backed knowledge graphs, converged databases, AI modernization, and Netflix's real-time distributed graph query architecture. It highlights composable architectures, data quality and observability, evolving schemas supported by LLM-assisted extraction, Iceberg full-text search, and optimization techniques including concurrency control, streaming filters, and caching.

### Source excerpt

The Weekly Data Engineering Newsletter

## Chaos Hub & MCP Prompt Library: Harness Resilience Testing

DevFeed: [Chaos Hub & MCP Prompt Library: Harness Resilience Testing](<https://devfeed.tech/articles/chaos-hub-mcp-prompt-library-harness-resilience-testing-13378.md>)

Original publisher: [Read original article](<https://www.harness.io/blog/chaos-hub-in-docs-prompt-library-for-mcp-whats-new-in-resilience-testing>)

Author: Pritesh Kiri

Published: 2026-08-05T00:00:00Z

Content type: article

Language: en

Sources: [Harness Blog](<https://devfeed.tech/sources/harness-blog.md>)

Topics: [Resilience](<https://devfeed.tech/topics/resilience.md>), [Chaos Engineering](<https://devfeed.tech/topics/chaos-engineering.md>), [Model Context Protocol](<https://devfeed.tech/topics/model-context-protocol.md>), [AWS Fault Injection Service (FIS)](<https://devfeed.tech/topics/aws-fault-injection-service-fis.md>), [MCP Server](<https://devfeed.tech/topics/mcp-server.md>), [Documentation](<https://devfeed.tech/topics/documentation.md>), [cursor](<https://devfeed.tech/topics/cursor.md>), [Claude](<https://devfeed.tech/topics/claude.md>)

Tags: [chaos](<https://devfeed.tech/tags/chaos.md>), [chaos-engineering](<https://devfeed.tech/tags/chaos-engineering.md>), [claude](<https://devfeed.tech/tags/claude.md>), [cursor](<https://devfeed.tech/tags/cursor.md>), [documentation](<https://devfeed.tech/tags/documentation.md>), [mcp](<https://devfeed.tech/tags/mcp.md>), [mcp-server](<https://devfeed.tech/tags/mcp-server.md>), [resilience](<https://devfeed.tech/tags/resilience.md>)

### AI overview

Harness Resilience Testing has added two documentation updates: Chaos Hub is now available directly in the documentation, and a Prompt Library provides ready-to-use prompts for running resilience workflows through Harness MCP using natural language.

### Source excerpt

New in Harness Resilience Testing: Chaos Hub now lives in the docs, plus a Prompt Library for running chaos experiments via MCP using natural language. | Blog

## Taming chaos is a learnable skill

DevFeed: [Taming chaos is a learnable skill](<https://devfeed.tech/articles/taming-chaos-is-a-learnable-skill-37643.md>)

Original publisher: [Read original article](<https://swizec.com/blog/taming-chaos-is-a-learnable-skill>)

Author: hi@swizec.com (Swizec Teller)

Published: 2026-03-11T00:00:00Z

Content type: opinion

Language: en

Sources: [Swizec Teller](<https://devfeed.tech/sources/swizec-teller.md>)

Topics: [Software Engineering](<https://devfeed.tech/topics/software-engineering.md>), [engineering-culture](<https://devfeed.tech/topics/engineering-culture.md>), [Low-Code / Internal Tools](<https://devfeed.tech/topics/internal-tools.md>)

Tags: [approach](<https://devfeed.tech/tags/approach.md>), [chaos](<https://devfeed.tech/tags/chaos.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [exceptions](<https://devfeed.tech/tags/exceptions.md>), [internal-tools](<https://devfeed.tech/tags/internal-tools.md>), [operations](<https://devfeed.tech/tags/operations.md>), [software-engineering](<https://devfeed.tech/tags/software-engineering.md>), [startup](<https://devfeed.tech/tags/startup.md>)

### AI overview

The article argues that handling interruptions, operational surprises, exceptions, and changing priorities is a learnable software-engineering skill. It recommends owning system problems, managing scope and tradeoffs, validating quality requirements, and maintaining a long-term vision despite short-term disruption.

### Source excerpt

How you approach software engineering makes it harder or easier to handle interruptions and other chaos. Writing a behavioral interview made me realize this is a learnable skill!

## Antifragile Systems and Teams

DevFeed: [Antifragile Systems and Teams](<https://devfeed.tech/articles/antifragile-systems-and-teams-27864.md>)

Original publisher: [Read original article](<https://gagor.pro/book/2025/antifragile-systems-and-teams/>)

Author: Tom

Published: 2025-08-14T00:00:00Z

Content type: opinion

Language: en

Sources: [Tomasz Gągor](<https://devfeed.tech/sources/tomasz-gagor.md>)

Topics: [systems](<https://devfeed.tech/topics/systems.md>), [Resilience](<https://devfeed.tech/topics/resilience.md>), [DevOps](<https://devfeed.tech/topics/devops.md>), [Chaos Engineering](<https://devfeed.tech/topics/chaos-engineering.md>), [Automation](<https://devfeed.tech/topics/automation.md>), [Tooling](<https://devfeed.tech/topics/tooling.md>), [Netflix](<https://devfeed.tech/topics/netflix.md>)

Tags: [automation](<https://devfeed.tech/tags/automation.md>), [chaos](<https://devfeed.tech/tags/chaos.md>), [chaos-engineering](<https://devfeed.tech/tags/chaos-engineering.md>), [culture](<https://devfeed.tech/tags/culture.md>), [devops](<https://devfeed.tech/tags/devops.md>), [netflix](<https://devfeed.tech/tags/netflix.md>), [resilience](<https://devfeed.tech/tags/resilience.md>), [systems](<https://devfeed.tech/tags/systems.md>), [technology](<https://devfeed.tech/tags/technology.md>)

### AI overview

A concise review of Dave Zwieback's book about antifragile systems and teams. It explains how organizations can use change, mistakes, and small failures to learn and become stronger, emphasizing DevOps practices such as culture, automation, measurement, sharing, chaos engineering, and frequent small deployments.

### Source excerpt

Antifragile Systems and Teams Author: Dave Zwieback This is a short read, but it does a solid job of capturing an important idea: that the healthiest systems and teams aren't just resistant to change, they actually get stronger through it. Zwieback contrasts fragile organizations - those that try to lock things down and prevent every possible failure - with antifragile ones, which use volatility, mistakes, and small shocks as opportunities to learn and improve.

## How a Live Multiplayer Snakes Demo Tested Temporal Workflow Resilience

DevFeed: [How a Live Multiplayer Snakes Demo Tested Temporal Workflow Resilience](<https://devfeed.tech/articles/snakes-chaos-and-the-resilience-of-temporal-a-live-demo-breakdown-35982.md>)

Original publisher: [Read original article](<https://temporal.io/blog/snakes-live-demo-on-resilience>)

Author: Rob Holland

Published: 2024-12-05T08:00:00Z

Content type: article

Language: en

Sources: [Temporal Blog](<https://devfeed.tech/sources/temporal-blog.md>)

Topics: [Resilience](<https://devfeed.tech/topics/resilience.md>), [distributed-systems](<https://devfeed.tech/topics/distributed-systems.md>), [real-time](<https://devfeed.tech/topics/real-time.md>), [Svelte](<https://devfeed.tech/topics/svelte.md>), [WebSocket](<https://devfeed.tech/topics/websocket.md>), [Cloud](<https://devfeed.tech/topics/cloud.md>)

Tags: [chaos](<https://devfeed.tech/tags/chaos.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [distributed-systems](<https://devfeed.tech/tags/distributed-systems.md>), [how-to](<https://devfeed.tech/tags/how-to.md>), [live](<https://devfeed.tech/tags/live.md>), [multiplayer](<https://devfeed.tech/tags/multiplayer.md>), [real-time](<https://devfeed.tech/tags/real-time.md>), [resilience](<https://devfeed.tech/tags/resilience.md>), [svelte](<https://devfeed.tech/tags/svelte.md>), [websocket](<https://devfeed.tech/tags/websocket.md>)

### AI overview

This blog post explains how Temporal powered a live multiplayer Snakes game with independent workflows for the game, rounds, and snakes. The demo introduced Worker restarts to test recovery from induced outages while the game continued running.

### Source excerpt

We used a live, multiplayer snake game powered by Temporal workflows to showcase the power of Durable Execution. Learn all about it in this blog post.

## Ethereum Network Threat Involving RPC Size Limits Between the Merge and Dencun

DevFeed: [Ethereum Network Threat Involving RPC Size Limits Between the Merge and Dencun](<https://devfeed.tech/articles/sepolia-incident-17097.md>)

Original publisher: [Read original article](<https://blog.ethereum.org/en/2024/03/21/sepolia-incident>)

Author: Marius van der Wijden; Toni Wahrstätter; Parithosh Jayanthi

Published: 2024-03-21T00:00:00Z

Content type: article

Language: en

Sources: [Ethereum Foundation Blog](<https://devfeed.tech/sources/ethereum-foundation-blog.md>)

Topics: [incident](<https://devfeed.tech/topics/incident.md>), [Ethereum](<https://devfeed.tech/topics/ethereum.md>), [Network](<https://devfeed.tech/topics/network.md>), [Remote Procedure Call (RPC)](<https://devfeed.tech/topics/rpc.md>), [HTTP](<https://devfeed.tech/topics/http.md>)

Tags: [attacks](<https://devfeed.tech/tags/attacks.md>), [chaos](<https://devfeed.tech/tags/chaos.md>), [ethereum](<https://devfeed.tech/tags/ethereum.md>), [http](<https://devfeed.tech/tags/http.md>), [incident](<https://devfeed.tech/tags/incident.md>), [issue](<https://devfeed.tech/tags/issue.md>), [network](<https://devfeed.tech/tags/network.md>), [rpc](<https://devfeed.tech/tags/rpc.md>), [security](<https://devfeed.tech/tags/security.md>)

### AI overview

The article discloses an Ethereum network threat caused by differing RPC message-size limits carried into the engine API. This could allow blocks that some clients accepted while others rejected, potentially causing blocks to be forked away. The article also reports that lowering limits to 5MB did not fully resolve the issue because blocks could reach that limit through many sub-128KB transactions.

### Source excerpt

This blog post discloses a threat against the Ethereum network that was present from the Merge up until the Dencun hard fork....

## Scaling Findmypast for 1921 Census release - Part 3

DevFeed: [Scaling Findmypast for 1921 Census release - Part 3](<https://devfeed.tech/articles/scaling-findmypast-for-1921-census-release-part-3-19748.md>)

Original publisher: [Read original article](<https://tech.findmypast.com/scaling-fmp-part-3/>)

Author: Mike Thomas

Published: 2022-04-26T00:00:00Z

Content type: article

Language: en

Sources: [Findmypast](<https://devfeed.tech/sources/findmypast.md>)

Topics: [Deployment](<https://devfeed.tech/topics/deployment.md>), [Hardware](<https://devfeed.tech/topics/hardware.md>), [Kubernetes](<https://devfeed.tech/topics/kubernetes.md>), [Network](<https://devfeed.tech/topics/network.md>), [virtual machines](<https://devfeed.tech/topics/virtual-machines.md>), [Amazon Web Services](<https://devfeed.tech/topics/aws.md>), [Linux](<https://devfeed.tech/topics/linux.md>), [sql-server](<https://devfeed.tech/topics/sql-server.md>), [Windows](<https://devfeed.tech/topics/windows.md>), [ASP.NET](<https://devfeed.tech/topics/aspnet.md>), [C#](<https://devfeed.tech/topics/csharp.md>), [.NET](<https://devfeed.tech/topics/net.md>)

Tags: [aws](<https://devfeed.tech/tags/aws.md>), [blog-post](<https://devfeed.tech/tags/blog-post.md>), [c-sharp](<https://devfeed.tech/tags/c-sharp.md>), [chaos](<https://devfeed.tech/tags/chaos.md>), [deployment](<https://devfeed.tech/tags/deployment.md>), [findmypast](<https://devfeed.tech/tags/findmypast.md>), [hardware](<https://devfeed.tech/tags/hardware.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [k8s](<https://devfeed.tech/tags/k8s.md>), [kubernetes](<https://devfeed.tech/tags/kubernetes.md>), [load](<https://devfeed.tech/tags/load.md>), [net](<https://devfeed.tech/tags/net.md>), [network](<https://devfeed.tech/tags/network.md>), [release](<https://devfeed.tech/tags/release.md>), [sql-server](<https://devfeed.tech/tags/sql-server.md>), [stress](<https://devfeed.tech/tags/stress.md>), [testing](<https://devfeed.tech/tags/testing.md>), [virtual-machines](<https://devfeed.tech/tags/virtual-machines.md>), [windows](<https://devfeed.tech/tags/windows.md>)

### AI overview

Findmypast describes the infrastructure challenges involved in scaling its service for the 1921 Census of England and Wales release. The third post in the series examines network timeouts and throughput problems during stress testing, leading the team to investigate hardware, Hyper-V, virtualized networking, Kubernetes, AWS, and related service infrastructure.

### Source excerpt

Findmypast (FMP) released the 1921 Census of England & Wales at midnight on Jan 6th 2022 to an eager community of genealogists. The preparation of the census - preservation, digitisation and transcription took three years of hard work. Aside from the data preparation, we also had technical challenges to address. Specifically, could our services deal with the projected increase in users for the first few days of 1921 launch period? This post is the third in a series that details how we approached scaling our service to deal with a projected day one 12x increase of users. Digging deeper into the infrastructure We left the second blog post in a poor position, our scaling efforts were not going well; network timeouts and throughput issues dogged each and every one of our stress test game days. At this point, we started to suspect the issue was related to hardware rather than the services. We started to dig deeper into all layers of the network stack, along with looking at the new Hyper-V hardware. Before we dig too deeply into the things, lets recap on our existing hardware setup. For the search backend, we use SOLR with a cluster of bare metal machines to index and retrieve search results. This was one area where we knew in advance that we needed more resources, so we scaled the cluster into AWS for extra compute. The majority of the Findmypast service runs on K8s, but not on bare metal. The K8s cluster is composed of Linux virtual machines running on Hyper-V. Our Hyper-V setup consisted of two clusters of bare metal machines, however both clusters were running different versions of Hyper-V We also have some areas of the site still running Windows servers with a Microsoft .Net/C# stack with SQL Server backend. These are also virtual machines running on the same Hyper-V hardware. The internal services, staging environments, etc also run as virtual machines under Hyper-V. So, SOLR aside, Hyper-V managed pretty much everything, from our internal deployment process to the

## Scaling Findmypast for 1921 Census release - Part 2

DevFeed: [Scaling Findmypast for 1921 Census release - Part 2](<https://devfeed.tech/articles/scaling-findmypast-for-1921-census-release-part-2-19747.md>)

Original publisher: [Read original article](<https://tech.findmypast.com/scaling-fmp-part-2/>)

Author: Mike Thomas

Published: 2022-03-02T00:00:00Z

Content type: article

Language: en

Sources: [Findmypast](<https://devfeed.tech/sources/findmypast.md>)

Topics: [systems](<https://devfeed.tech/topics/systems.md>), [Testing](<https://devfeed.tech/topics/testing.md>), [GraphQL](<https://devfeed.tech/topics/graphql.md>), [Kubernetes](<https://devfeed.tech/topics/kubernetes.md>), [Deployment](<https://devfeed.tech/topics/deployment.md>)

Tags: [automated](<https://devfeed.tech/tags/automated.md>), [chaos](<https://devfeed.tech/tags/chaos.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [findmypast](<https://devfeed.tech/tags/findmypast.md>), [graphql](<https://devfeed.tech/tags/graphql.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [k8s](<https://devfeed.tech/tags/k8s.md>), [kubernetes](<https://devfeed.tech/tags/kubernetes.md>), [load](<https://devfeed.tech/tags/load.md>), [performance](<https://devfeed.tech/tags/performance.md>), [stress](<https://devfeed.tech/tags/stress.md>), [testing](<https://devfeed.tech/tags/testing.md>)

### AI overview

Findmypast describes the second phase of scaling its service for the 1921 Census of England and Wales release. The engineering teams ran coordinated load-test game days against production, targeting up to four times normal January load. The first test day kept the site generally responsive but exposed failures in a GraphQL service and other performance and Kubernetes configuration issues.

### Source excerpt

Findmypast (FMP) released the 1921 Census of England & Wales at midnight on Jan 6th 2022 to an eager community of genealogists. The preparation of the census - preservation, digitisation and transcription took three years of hard work. Aside from the data preparation, we also had technical challenges to address. Specifically, could our services deal with the projected increase in users for the first few days of 1921 launch period? This post is the second in a series that details how we approached scaling our service to deal with a projected day one 12x increase of users. If you haven't already, I'd suggest you read the first post which details the initial scaling steps taken by our engineering teams. Stress testing "game" days We ended the first post detailing that the load tests created by the teams were effective at testing the service in isolation. The next step was to get all the teams together and run their load tests at the same time against our production service. While still not a realistic example of a user journey, it would still test how our systems as a whole responded to the increased load. Game day # 1 Initially, we set a simple schedule of a load test "game day" each month. Scheduled for Thursday 8th April and spread over two hours we planned to run 3 load tests. Each was 20 mins in duration and we aimed for the first test to add 2x our normal load, the next test was 3x and the final test was 4x the load. Note that "normal load" means the load we would normally expect in a typical January. (1921 Census was launched in January). We have seasonal visitor patterns with more visitors during (the northern hemisphere) winter months than summer. When we set our targets for the load test, we aimed for 4x the normal January load, not the current April load. Overall, the first stress test game day was a success - the site stayed alive and generally responsive. But we did learn a few things from the first day: One of our key services that deals with GraphQL quer

## Scaling Findmypast for 1921 Census release - Part 1

DevFeed: [Scaling Findmypast for 1921 Census release - Part 1](<https://devfeed.tech/articles/scaling-findmypast-for-1921-census-release-part-1-19746.md>)

Original publisher: [Read original article](<https://tech.findmypast.com/scaling-fmp-part-1/>)

Author: Mike Thomas

Published: 2022-01-31T00:00:00Z

Content type: article

Language: en

Sources: [Findmypast](<https://devfeed.tech/sources/findmypast.md>)

Topics: [data](<https://devfeed.tech/topics/data.md>), [Kubernetes](<https://devfeed.tech/topics/kubernetes.md>), [Testing](<https://devfeed.tech/topics/testing.md>), [Back end](<https://devfeed.tech/topics/backend.md>), [telemetry](<https://devfeed.tech/topics/telemetry.md>)

Tags: [chaos](<https://devfeed.tech/tags/chaos.md>), [data](<https://devfeed.tech/tags/data.md>), [findmypast](<https://devfeed.tech/tags/findmypast.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [k8s](<https://devfeed.tech/tags/k8s.md>), [kubernetes](<https://devfeed.tech/tags/kubernetes.md>), [load](<https://devfeed.tech/tags/load.md>), [metrics](<https://devfeed.tech/tags/metrics.md>), [performance](<https://devfeed.tech/tags/performance.md>), [release](<https://devfeed.tech/tags/release.md>), [stress](<https://devfeed.tech/tags/stress.md>), [testing](<https://devfeed.tech/tags/testing.md>)

### AI overview

Findmypast describes how its engineering team prepared for the 1921 Census of England & Wales release, anticipating a 12-fold increase in users. The article covers a scaling working group, performance improvements, service metrics, chaos and stress testing, and mapping dependencies across nearly 140 Kubernetes-based microservices.

### Source excerpt

Findmypast (FMP) released the 1921 Census of England & Wales at midnight on Jan 6th 2022 to an eager community of genealogists. The preparation of the census - preservation, digitisation and transcription took three years of hard work. Aside from the data preparation, we also had technical challenges to address. Specifically, could our services deal with the projected increase in users for the first few days of 1921 launch period? This post is the first in a series that details how we approached scaling our service to deal with a projected day one 12x increase of users. In the beginning... ..way back in Nov 2020 the release of the 1921 census was just over 13 months away. Within engineering, our thoughts started to turn to ensuring that Findmypast was fit to meeting the anticipated increase in load. We knew from previous incidents that the reliability and performance of some of the services can cause issues when under load, so, within engineering, we started a scaling working group to address not only these specific issues, but to prove to management and our external stakeholders that the system as a whole was resilient and could perform well under load. The expectations for the working group was to: Champion performance improvements within the team. That is, evangelise scaling to the team, make sure that backend improvement work was added to backlogs, scheduled into sprints, etc. Ensure that the key metrics for each service owned by the team were correctly instrumented, visualised and alerted upon. Schedule time for chaos and stress testing team services. The initial meeting set the goals and expectations for the team and also discussed a few incidents that caused a site outages due to cascading failures and the like. Essentially, we highlighted the things that have gone wrong in the past and the discussed areas that would have limited the blast impact of the outage. We also re-visited our documentation, we have almost 140 micro services running in Kubernetes (K8s) a

## Edge of Chaos and Hyper Productive Software Development Teams

DevFeed: [Edge of Chaos and Hyper Productive Software Development Teams](<https://devfeed.tech/articles/edge-of-chaos-and-hyper-productive-software-development-teams-30321.md>)

Original publisher: [Read original article](<https://www.mdubakov.com/posts/edge-of-chaos-hyper-productive-teams/>)

Published: 2008-12-02T15:41:57Z

Content type: article

Language: en

Sources: [Blog by Michael Dubakov](<https://devfeed.tech/sources/blog-by-michael-dubakov.md>)

Topics: [Development](<https://devfeed.tech/topics/development.md>), [Agile](<https://devfeed.tech/topics/agile.md>)

Tags: [adaptive](<https://devfeed.tech/tags/adaptive.md>), [agile](<https://devfeed.tech/tags/agile.md>), [chaos](<https://devfeed.tech/tags/chaos.md>), [development](<https://devfeed.tech/tags/development.md>), [productivity](<https://devfeed.tech/tags/productivity.md>), [scrum](<https://devfeed.tech/tags/scrum.md>), [software-development](<https://devfeed.tech/tags/software-development.md>), [teams](<https://devfeed.tech/tags/teams.md>)

### AI overview

This article applies the "edge of chaos" concept from complex adaptive systems to software development teams. It argues that teams may be most effective when they balance structure and flexibility: excessive process can impede creativity, while too little process can prevent useful work. The article connects this balance with the idea of hyper-productive teams and Agile practices.

### Source excerpt

Many researchers tend to think that CAS may work efficiently at the edge of chaos. A system with a very high order can't solve complex creative tasks, as well as a pure chaotic system. It seems the system should balance between chaos and order to survive and adapt. The "edge of chaos" term was emerged during self-reproductive cellular automata researches in 1990. It appeared that at a particular noise level the system can reproduce its state, while at a low noise levels states are random. This edge value was mathematically analyzed and it was shown that in such a state the system has a maximum of information. The "Edge of Chaos" term spread and applied to complex adaptive systems. For example, we may suggest that evolution drives systems to the edge of chaos, and it is one of the reasons of life origin. Life is a very interesting and complex thing in itself. Most likely it could not appear neither in order state and in chaos state. Life exists if there is a stable base, but with possibilities of changes. Edge of Chaos and Software Development How can we apply the edge of chaos concept to software development? Obviously, edge of chaos is a state where the development team works with maximum efficiency. Complex processes and explicit rules impede creativity. The team becomes too ordered to think :). On the other hand, no processes and no rules lead to chaos. The development team can't complete anything useful, and that is something we call hack & fix process. As a software development folks, we are particularly interested in two questions: How can we drive the team to the edge of chaos? How we will ensure that the team is at the edge of chaos? The human who knows the answers to these questions most certainly will be rich and famous. However I am not sure whether you can find such a man. It is obvious that edge of chaos = hyper productivity. This term often used in Scrum. However it is quite hard to define hyper productivity. Hyper-productive teams are hard to define.