# experiments

Datadog Experiments is a product experimentation tool for running and analyzing experiments and A/B tests, with experiment-health guardrails and real-time performance monitoring.

This is one page of public article previews, not the complete archive. Follow Next page to continue. Summaries are not the original full articles.

## Analyze your experiments in ChatGPT with the Datadog Experiments plugin

DevFeed: [Analyze your experiments in ChatGPT with the Datadog Experiments plugin](<https://devfeed.tech/articles/analyze-your-experiments-in-chatgpt-with-the-datadog-experiments-plugin-2239.md>)

Original publisher: [Read original article](<https://www.datadoghq.com/blog/chatgpt-datadog-experiments/>)

Author: Jonathan Fulton; Amy Zhou; Uday Tennety

Published: 2026-09-11T00:00:00Z

Content type: article

Language: en

Sources: [Datadog | The Monitor blog](<https://devfeed.tech/sources/datadog-the-monitor-blog.md>)

Topics: [experiments](<https://devfeed.tech/topics/experiments.md>), [data](<https://devfeed.tech/topics/data.md>), [Latency](<https://devfeed.tech/topics/latency.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [analysis](<https://devfeed.tech/tags/analysis.md>), [chatgpt](<https://devfeed.tech/tags/chatgpt.md>), [experimentation](<https://devfeed.tech/tags/experimentation.md>), [experiments](<https://devfeed.tech/tags/experiments.md>), [openai](<https://devfeed.tech/tags/openai.md>)

### AI overview

Datadog introduces a ChatGPT Work plugin that lets teams query current experiment results in plain language while retaining Datadog's statistical methods, guardrails, and connected context.

### Source excerpt

Learn how the Datadog Experiments OpenAI Data plugin lets your team read, question, and act on experiment results directly in ChatGPT.

## How we built Datadog Experiments

DevFeed: [How we built Datadog Experiments](<https://devfeed.tech/articles/how-we-built-datadog-experiments-2283.md>)

Original publisher: [Read original article](<https://www.datadoghq.com/blog/how-we-built-datadog-experiments/>)

Author: Chas DeVeas; Aaron Silverman; Tyler Buffington; Jonathan Fulton; Taylor Overturf

Published: 2026-09-10T00:00:00Z

Content type: article

Language: en

Sources: [Datadog | The Monitor blog](<https://devfeed.tech/sources/datadog-the-monitor-blog.md>)

Topics: [experiments](<https://devfeed.tech/topics/experiments.md>), [real user monitoring](<https://devfeed.tech/topics/real-user-monitoring.md>)

Tags: [a-b-testing](<https://devfeed.tech/tags/a-b-testing.md>), [acquisition](<https://devfeed.tech/tags/acquisition.md>), [data-analytics](<https://devfeed.tech/tags/data-analytics.md>), [data-engineering](<https://devfeed.tech/tags/data-engineering.md>), [experimentation](<https://devfeed.tech/tags/experimentation.md>), [experiments](<https://devfeed.tech/tags/experiments.md>), [metrics](<https://devfeed.tech/tags/metrics.md>), [observability](<https://devfeed.tech/tags/observability.md>), [product-analytics](<https://devfeed.tech/tags/product-analytics.md>), [real-user-monitoring](<https://devfeed.tech/tags/real-user-monitoring.md>)

### AI overview

Datadog describes rebuilding its experimentation platform to speed up confident A/B-test decisions. The article explains a flexible CUPED approach that reduces metric variance and can be applied to segments.

### Source excerpt

Datadog Experiments shortens the time from result to decision with CUPED on percentiles, verifiable warehouse results, and near real-time RUM metrics.

## Coordinate product launches with Datadog

DevFeed: [Coordinate product launches with Datadog](<https://devfeed.tech/articles/coordinate-product-launches-with-datadog-2244.md>)

Original publisher: [Read original article](<https://www.datadoghq.com/blog/coordinate-product-launches-with-datadog/>)

Author: Milene Darnis; Adam Virani

Published: 2026-09-08T00:00:00Z

Content type: tutorial

Language: en

Sources: [Datadog | The Monitor blog](<https://devfeed.tech/sources/datadog-the-monitor-blog.md>)

Topics: [Instrumentation](<https://devfeed.tech/topics/instrumentation.md>), [experiments](<https://devfeed.tech/topics/experiments.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [bits-ai](<https://devfeed.tech/tags/bits-ai.md>), [digital-experience-monitoring](<https://devfeed.tech/tags/digital-experience-monitoring.md>), [experimentation](<https://devfeed.tech/tags/experimentation.md>), [experiments](<https://devfeed.tech/tags/experiments.md>), [feature-flags](<https://devfeed.tech/tags/feature-flags.md>), [instrumentation](<https://devfeed.tech/tags/instrumentation.md>), [launch](<https://devfeed.tech/tags/launch.md>), [product-analytics](<https://devfeed.tech/tags/product-analytics.md>), [session-replay](<https://devfeed.tech/tags/session-replay.md>)

### AI overview

A tutorial on using Datadog Product Analytics Launches to plan product releases, define measurement questions, create tracking plans, and identify missing events and properties before rollout.

### Source excerpt

Learn how to turn a product brief and feature flag into a connected launch workflow for instrumentation, experimentation, QA, and reporting.

## Visualize how CUPED adjusts experiment results with Datadog

DevFeed: [Visualize how CUPED adjusts experiment results with Datadog](<https://devfeed.tech/articles/visualize-how-cuped-adjusts-experiment-results-with-datadog-2246.md>)

Original publisher: [Read original article](<https://www.datadoghq.com/blog/cuped-adjustments-visualization/>)

Author: Tyler Buffington; Lukas Goetz-Weiss; Ryan Lucht

Published: 2026-09-01T00:00:00Z

Content type: tutorial

Language: en

Sources: [Datadog | The Monitor blog](<https://devfeed.tech/sources/datadog-the-monitor-blog.md>)

Topics: [experiments](<https://devfeed.tech/topics/experiments.md>)

Tags: [analysis](<https://devfeed.tech/tags/analysis.md>), [digital-experience-monitoring](<https://devfeed.tech/tags/digital-experience-monitoring.md>), [experiment](<https://devfeed.tech/tags/experiment.md>), [experiments](<https://devfeed.tech/tags/experiments.md>), [how-to](<https://devfeed.tech/tags/how-to.md>), [metrics](<https://devfeed.tech/tags/metrics.md>), [product-analytics](<https://devfeed.tech/tags/product-analytics.md>), [visualization](<https://devfeed.tech/tags/visualization.md>)

### AI overview

This tutorial explains Datadog Experiments' CUPED adjustments visualization, which breaks the difference between raw and CUPED-adjusted experiment lift into individual covariate contributions.

### Source excerpt

Learn how Datadog visualizes CUPED adjustments so you can trace which covariates change experiment lift estimates and improve precision.

## Two ways to measure the cumulative impact of experiments

DevFeed: [Two ways to measure the cumulative impact of experiments](<https://devfeed.tech/articles/two-ways-to-measure-the-cumulative-impact-of-experiments-2316.md>)

Original publisher: [Read original article](<https://www.datadoghq.com/blog/two-ways-to-measure-cumulative-impact/>)

Author: Lukas Goetz-Weiss; Eddie Cai

Published: 2026-08-18T00:00:00Z

Content type: opinion

Language: en

Sources: [Datadog | The Monitor blog](<https://devfeed.tech/sources/datadog-the-monitor-blog.md>)

Topics: [experiments](<https://devfeed.tech/topics/experiments.md>), [Statistics](<https://devfeed.tech/topics/statistics.md>), [Data Science](<https://devfeed.tech/topics/data-science.md>)

Tags: [analysis](<https://devfeed.tech/tags/analysis.md>), [experimentation](<https://devfeed.tech/tags/experimentation.md>), [experiments](<https://devfeed.tech/tags/experiments.md>), [product](<https://devfeed.tech/tags/product.md>), [statistics](<https://devfeed.tech/tags/statistics.md>), [testing](<https://devfeed.tech/tags/testing.md>)

### AI overview

This article explains why summing the observed lift from winning experiments overstates cumulative impact because of the winner's curse. It presents two established approaches: a randomized holdout that measures the combined effect directly, and a statistical correction that aggregates existing experiment estimates. Datadog's Cumulative Impact feature applies the correction without requiring a quarter-long holdout and can analyze an entire experimentation program or a filtered team subset.

### Source excerpt

Summing individual wins overstates true impact. See two accurate methods, holdouts and Datadog's Cumulative Impact, and how to choose between them.

## Implementing dynamic feature flags with AWS AppConfig on AWS Lambda

DevFeed: [Implementing dynamic feature flags with AWS AppConfig on AWS Lambda](<https://devfeed.tech/articles/implementing-dynamic-feature-flags-with-aws-appconfig-on-aws-lambda-4664.md>)

Original publisher: [Read original article](<https://aws.amazon.com/blogs/compute/implementing-dynamic-feature-flags-with-aws-appconfig-on-aws-lambda/>)

Author: Daniel Abib

Published: 2026-08-14T16:25:51Z

Content type: tutorial

Language: en

Sources: [AWS Compute Blog](<https://devfeed.tech/sources/aws-compute-blog.md>)

Topics: [AWS Lambda](<https://devfeed.tech/topics/aws-lambda.md>), [Serverless](<https://devfeed.tech/topics/serverless.md>), [experiments](<https://devfeed.tech/topics/experiments.md>), [configuration](<https://devfeed.tech/topics/configuration.md>), [Latency](<https://devfeed.tech/topics/latency.md>)

Tags: [a-b-testing](<https://devfeed.tech/tags/a-b-testing.md>), [advanced-300](<https://devfeed.tech/tags/advanced-300.md>), [aws](<https://devfeed.tech/tags/aws.md>), [aws-lambda](<https://devfeed.tech/tags/aws-lambda.md>), [feature](<https://devfeed.tech/tags/feature.md>), [feature-flags](<https://devfeed.tech/tags/feature-flags.md>), [how-to](<https://devfeed.tech/tags/how-to.md>), [latency](<https://devfeed.tech/tags/latency.md>), [real-time](<https://devfeed.tech/tags/real-time.md>), [serverless](<https://devfeed.tech/tags/serverless.md>), [technical-how-to](<https://devfeed.tech/tags/technical-how-to.md>)

### AI overview

A tutorial on implementing dynamic feature flags with AWS AppConfig on AWS Lambda. It explains how feature flags support safe deployments, experiments, gradual rollouts, and changes without redeploying functions, using the AWS AppConfig Lambda extension to cache configuration locally and reduce latency.

### Source excerpt

Feature toggles allow you to change application behavior in real time without deploying new code. Learn how to implement dynamic feature flags with AWS AppConfig on AWS Lambda for safe deployments, gradual rollouts, and instant rollback.

## LaunchDarkly is now available on the Vercel Marketplace

DevFeed: [LaunchDarkly is now available on the Vercel Marketplace](<https://devfeed.tech/articles/launchdarkly-is-now-available-on-the-vercel-marketplace-997.md>)

Original publisher: [Read original article](<https://vercel.com/changelog/launchdarkly-is-now-available-on-the-vercel-marketplace>)

Author: Sam Halstead

Published: 2026-08-11T00:00:00Z

Content type: release

Language: en

Sources: [Vercel News](<https://devfeed.tech/sources/vercel-news.md>)

Topics: [feature flags](<https://devfeed.tech/topics/feature-flags.md>), [Vercel](<https://devfeed.tech/topics/vercel.md>), [experiments](<https://devfeed.tech/topics/experiments.md>), [SDKs](<https://devfeed.tech/topics/sdks.md>)

Tags: [agent](<https://devfeed.tech/tags/agent.md>), [experiments](<https://devfeed.tech/tags/experiments.md>), [feature-flags](<https://devfeed.tech/tags/feature-flags.md>), [sdk](<https://devfeed.tech/tags/sdk.md>), [vercel](<https://devfeed.tech/tags/vercel.md>)

### AI overview

LaunchDarkly is now available through the Vercel Marketplace, providing feature flags with local evaluation, release targeting, experiments, metrics, rollback controls, and Vercel Toolbar management.

### Source excerpt

LaunchDarkly is now available on the Vercel Marketplace, allowing you to quickly get started with feature flags without additional setup. You can: Sync flags into Global Config and evaluate them locally Target releases by user, attribute, or segment Run experiments with metrics, and roll back with a kill switch View and override flags from the Vercel Toolbar with the Flags Explorer To get started, run vercel install launchdarkly, add the @flags-sdk/launchdarkly adapter, and declare a flag with the Flags SDK: Using a coding agent? Hand it this prompt: Add LaunchDarkly from the Vercel Marketplace, or read the adapter docs. Read more

## Tailscale's Remy Guercio on what comes after token maxing

DevFeed: [Tailscale's Remy Guercio on what comes after token maxing](<https://devfeed.tech/articles/tailscale-s-remy-guercio-on-what-comes-after-token-maxing-16050.md>)

Original publisher: [Read original article](<https://workos.com/blog/remy-guercio-tailscale-token-maxing-roi-aie-2026>)

Author: WorkOS

Published: 2026-08-05T23:49:35Z

Content type: article

Language: en

Sources: [WorkOS Blog](<https://devfeed.tech/sources/workos-blog.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [AI-assisted coding](<https://devfeed.tech/topics/ai-assisted-coding.md>), [experiments](<https://devfeed.tech/topics/experiments.md>)

Tags: [agents](<https://devfeed.tech/tags/agents.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-gateway](<https://devfeed.tech/tags/ai-gateway.md>), [analysis](<https://devfeed.tech/tags/analysis.md>), [cost](<https://devfeed.tech/tags/cost.md>), [experiments](<https://devfeed.tech/tags/experiments.md>), [model](<https://devfeed.tech/tags/model.md>), [tailscale](<https://devfeed.tech/tags/tailscale.md>), [token](<https://devfeed.tech/tags/token.md>)

### AI overview

An interview with Tailscale's Remy Guercio examines the shift from maximizing token usage to maximizing return on investment in AI development. It covers coding-agent adoption, rising bills across multiple AI tools, cost per task, model experimentation, and the visibility an AI gateway can provide.

### Source excerpt

Tailscale's Remy Guercio tells Michael Grinich why the next 12 months of AI are ROI maxing: cost per task, model experiments, and what a gateway can see.

## Building self-healing feature releases with Harness FME metric alerts and Event Relay

DevFeed: [Building self-healing feature releases with Harness FME metric alerts and Event Relay](<https://devfeed.tech/articles/when-metrics-scream-your-flags-hit-mute-13496.md>)

Original publisher: [Read original article](<https://www.harness.io/blog/when-metrics-scream-your-flags-hit-mute>)

Author: Joshua Klein

Published: 2026-07-24T00:00:00Z

Content type: tutorial

Language: en

Sources: [Harness Blog](<https://devfeed.tech/sources/harness-blog.md>)

Topics: [releases](<https://devfeed.tech/topics/releases.md>), [feature flags](<https://devfeed.tech/topics/feature-flags.md>), [Automation](<https://devfeed.tech/topics/automation.md>), [experiments](<https://devfeed.tech/topics/experiments.md>), [Resilience](<https://devfeed.tech/topics/resilience.md>), [Monitoring](<https://devfeed.tech/topics/monitoring.md>)

Tags: [automation](<https://devfeed.tech/tags/automation.md>), [event](<https://devfeed.tech/tags/event.md>), [feature](<https://devfeed.tech/tags/feature.md>), [feature-flags](<https://devfeed.tech/tags/feature-flags.md>), [metrics](<https://devfeed.tech/tags/metrics.md>), [monitoring](<https://devfeed.tech/tags/monitoring.md>), [payload](<https://devfeed.tech/tags/payload.md>), [pipelines](<https://devfeed.tech/tags/pipelines.md>), [resilience](<https://devfeed.tech/tags/resilience.md>), [self-healing](<https://devfeed.tech/tags/self-healing.md>), [webhooks](<https://devfeed.tech/tags/webhooks.md>)

### AI overview

This tutorial explains how to connect Harness FME metric alerts to Event Relay triggers so a pipeline can automatically mitigate risky feature releases through actions such as killing a feature flag. The pattern separates signal emission, webhook handling, and controlled remediation.

### Source excerpt

Learn how to build self-healing feature releases with Harness Feature Management & Experimentation. Connect metric-alert webhooks to Event Relay triggers and au | Blog

## Harness AI Configs for Runtime Controls of AI Behavior

DevFeed: [Harness AI Configs for Runtime Controls of AI Behavior](<https://devfeed.tech/articles/harness-ai-configs-for-runtime-controls-of-ai-behavior-13364.md>)

Original publisher: [Read original article](<https://www.harness.io/blog/announcing-ai-config-management>)

Author: Nico Zelaya

Published: 2026-07-21T00:00:00Z

Content type: article

Language: en

Sources: [Harness Blog](<https://devfeed.tech/sources/harness-blog.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [configuration](<https://devfeed.tech/topics/configuration.md>), [feature flags](<https://devfeed.tech/topics/feature-flags.md>), [experiments](<https://devfeed.tech/topics/experiments.md>), [Model Routing](<https://devfeed.tech/topics/model-routing.md>), [Inference](<https://devfeed.tech/topics/inference.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [blog](<https://devfeed.tech/tags/blog.md>), [config](<https://devfeed.tech/tags/config.md>), [experimentation](<https://devfeed.tech/tags/experimentation.md>), [inference](<https://devfeed.tech/tags/inference.md>), [policy](<https://devfeed.tech/tags/policy.md>), [routing](<https://devfeed.tech/tags/routing.md>)

### AI overview

Harness AI Config Management provides a governed runtime configuration layer for changing prompts, models, routing, inference parameters, and other AI behavior without redeploying code. It supports targeting, experimentation, approvals, policy controls, versioning, and audit trails.

### Source excerpt

Harness AI Config Management helps teams change prompts, models, and AI behavior at runtime with targeting, experimentation, approvals, policy, and audit trails | Blog

## A Recap of the 2026 Experimentation Conference at Booking.com

DevFeed: [A Recap of the 2026 Experimentation Conference at Booking.com](<https://devfeed.tech/articles/a-recap-of-the-2026-experimentation-conference-at-booking-com-30447.md>)

Original publisher: [Read original article](<https://booking.ai/a-recap-of-the-2026-experimentation-conference-at-booking-com-f43d48698fcd?source=rss----4d265f07defc---4>)

Author: Mel JI Mueller

Published: 2026-07-16T08:18:03Z

Content type: article

Language: en

Sources: [Booking.com Data Science](<https://devfeed.tech/sources/booking-com-data-science.md>)

Topics: [experiments](<https://devfeed.tech/topics/experiments.md>), [decision-making](<https://devfeed.tech/topics/decision-making.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Tooling](<https://devfeed.tech/topics/tooling.md>)

Tags: [2025](<https://devfeed.tech/tags/2025.md>), [ai](<https://devfeed.tech/tags/ai.md>), [analysis](<https://devfeed.tech/tags/analysis.md>), [data-science](<https://devfeed.tech/tags/data-science.md>), [events](<https://devfeed.tech/tags/events.md>), [experimentation](<https://devfeed.tech/tags/experimentation.md>), [organizational](<https://devfeed.tech/tags/organizational.md>), [parallel](<https://devfeed.tech/tags/parallel.md>), [recap](<https://devfeed.tech/tags/recap.md>), [themes](<https://devfeed.tech/tags/themes.md>)

### AI overview

A recap of Booking.com's 2026 Experimentation Conference, which brought together more than 150 experimentation practitioners from 49 companies. The article summarizes survey findings, conference themes, and sessions on AI-assisted experimentation, experimentation quality and velocity, and organizational culture.

### Source excerpt

By Kevin Anderson, Angelica Goetzen, Jorden Lentze, and Melanie Mueller On May 18, 2026, we hosted the third annual Experimentation Conference at Booking.com on our Amsterdam campus. What started in 2024 as an experiment itself -- would large-scale experimentation practitioners come together to learn from each other? -- has grown into an event which brings together over 150 practitioners from 49 companies which run experiments at scale. About one third of attendees came back a second or third time. The room collectively ran 56,000 experiments per year. It's a unique crowd, and that's exactly the point. The day opened with sharing the results of the survey data we collected from the participating companies on the state of experimentation across the room, revealing some interesting findings: most teams operate a centre of excellence model, roughly a third release over 90% of features through controlled experiments, and the top challenges are scaling, coordination, platform tooling, and culture. These shared experiences helped shape the programme. We had three sessions, grouped by the three conference themes: AI and experimentation: AI-assisted analysis, no-code experimentation Quality / velocity tradeoff: High-quality vs high-speed experimentation Experimentation culture: Build organizational buy-in and data-driven decision-making Each session followed the same format: two talks, then a panel discussion on the same topic. We closed with nine parallel breakout groups for deeper conversation. Below is a recap of the key sessions. Read the recap of 2025 | Read the recap of 2024 Session 1: AI and experimentation The conference started off with the hot topic of AI in experimentation. AI is changing how we experiment and how we support experimenters. How Experimentation Protects Decisions in an AI-Written World -- Marcel Toben Marcel Toben, Head of Engineering at Zalando, opened with a provocation he'd recently heard from software engineers in Berlin: nobody on his team had wr

## GLM-5.2: Considerations for enterprise teams starting out with open-weight models

DevFeed: [GLM-5.2: Considerations for enterprise teams starting out with open-weight models](<https://devfeed.tech/articles/glm-5-2-considerations-for-enterprise-teams-starting-out-with-open-weight-models-33584.md>)

Original publisher: [Read original article](<https://blog.scottlogic.com/2026/07/08/glm-5-2-considerations-for-enterprise-teams-starting-out-with-open-weight-models.html>)

Author: Robat Williams

Published: 2026-07-08T10:00:00Z

Content type: article

Language: en

Sources: [Scott Logic](<https://devfeed.tech/sources/scott-logic.md>)

Topics: [Agentic development](<https://devfeed.tech/topics/agentic-development.md>), [developer setup](<https://devfeed.tech/topics/developer-setup.md>), [AI Development](<https://devfeed.tech/topics/ai-development.md>), [experiments](<https://devfeed.tech/topics/experiments.md>), [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [Claude](<https://devfeed.tech/topics/claude.md>)

Tags: [agent](<https://devfeed.tech/tags/agent.md>), [agentic](<https://devfeed.tech/tags/agentic.md>), [agentic-development](<https://devfeed.tech/tags/agentic-development.md>), [artificial-intelligence](<https://devfeed.tech/tags/artificial-intelligence.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [claude](<https://devfeed.tech/tags/claude.md>), [developer-setup](<https://devfeed.tech/tags/developer-setup.md>), [experiments](<https://devfeed.tech/tags/experiments.md>), [open-source](<https://devfeed.tech/tags/open-source.md>)

### AI overview

The article describes how a small team prepared an open-weight agentic development setup for AI productivity experiments. It explains the team's model-selection criteria, distinguishes open-weight models from locally run models, and reports choosing GLM-5.2 after comparing DeepSeek V4, Kimi K-2.6, and other models using benchmarks and practical trials.

### Source excerpt

Setting yourself up to try out open-weight models for agentic development isn't difficult, but it isn't as straightforward as downloading a coding agent from one of the handful of well-known AI vendors. In preparation for the latest round of our AI productivity experiments, we've recently been through this process. Read on for the choices we made, the considerations at play, and what made our situation unusual.

## The effect distribution: The missing piece in experimentation programs

DevFeed: [The effect distribution: The missing piece in experimentation programs](<https://devfeed.tech/articles/the-effect-distribution-the-missing-piece-in-experimentation-programs-2266.md>)

Original publisher: [Read original article](<https://www.datadoghq.com/blog/effect-distribution-in-experimentation/>)

Author: Tyler Buffington

Published: 2026-07-02T00:00:00Z

Content type: article

Language: en

Sources: [Datadog | The Monitor blog](<https://devfeed.tech/sources/datadog-the-monitor-blog.md>)

Topics: [experiments](<https://devfeed.tech/topics/experiments.md>)

Tags: [a-b-testing](<https://devfeed.tech/tags/a-b-testing.md>), [experiment](<https://devfeed.tech/tags/experiment.md>), [experimentation](<https://devfeed.tech/tags/experimentation.md>), [experiments](<https://devfeed.tech/tags/experiments.md>), [testing](<https://devfeed.tech/tags/testing.md>)

### AI overview

The article explains why effect distributions are essential for interpreting results across experimentation programs. It shows how statistically significant results can all be false positives when true effects are concentrated at zero, and introduces the challenge of estimating true effects from noisy observed effects.

### Source excerpt

Learn about the importance of considering the effect distribution when running experiments.

## Variance Reduction Below the Randomization Grain

DevFeed: [Variance Reduction Below the Randomization Grain](<https://devfeed.tech/articles/variance-reduction-below-the-randomization-grain-20111.md>)

Original publisher: [Read original article](<https://tech.instacart.com/variance-reduction-below-the-randomization-grain-31719f87a7d2?source=rss----587883b5d2ee---4>)

Author: Tilman Drerup

Published: 2026-07-01T16:28:36Z

Content type: article

Language: en

Sources: [Instacart](<https://devfeed.tech/sources/instacart.md>)

Topics: [experiments](<https://devfeed.tech/topics/experiments.md>)

Tags: [capacity](<https://devfeed.tech/tags/capacity.md>), [causal-inference](<https://devfeed.tech/tags/causal-inference.md>), [economics](<https://devfeed.tech/tags/economics.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [experiment](<https://devfeed.tech/tags/experiment.md>), [experimentation](<https://devfeed.tech/tags/experimentation.md>), [experiments](<https://devfeed.tech/tags/experiments.md>), [machine-learning](<https://devfeed.tech/tags/machine-learning.md>), [marketplaces](<https://devfeed.tech/tags/marketplaces.md>), [metrics](<https://devfeed.tech/tags/metrics.md>), [reduce](<https://devfeed.tech/tags/reduce.md>), [science](<https://devfeed.tech/tags/science.md>), [techniques](<https://devfeed.tech/tags/techniques.md>), [variance](<https://devfeed.tech/tags/variance.md>)

### AI overview

This article explains how marketplace experiments can reduce metric variance below the level at which treatment is randomized. It describes cluster-level randomization for containing interference and shows how fine-grained outcome predictability can improve statistical power and reduce experimentation time.

### Source excerpt

Sergio Camelo, Caitlin Kearns, Matias Cersosimo, and Tilman Drerup As artificial intelligence increases the velocity of engineering and science teams, experimental throughput is set to become a bottleneck for many product decisions. Many companies can now build faster than they can experiment, with queues of good ideas running the risk of not being tested because of lack of experimental capacity. This problem is particularly severe in marketplaces, where the presence of spillover and cannibalization effects between experimental units requires cluster-level randomization techniques. That randomization, in turn, has the unfortunate tendency to substantially reduce statistical power and slow down experimentation. In this post, we show that the predictability of outcomes at fine grains can be exploited to reduce the variance of aggregate metrics, even when experiments themselves are run at a coarse level. Since statistical power depends on metric variability, this yields considerable reductions in experimentation time. The Interference Problem In marketplace settings, behavior and outcomes for individual participants are inherently intertwined. In a delivery marketplace like Instacart, for example, the dispatch system solves a bipartite matching problem between shoppers and customer orders. Since assignments are global and interdependent, matching an order to one shopper means that the same order cannot be matched to another shopper. As a result, changing the handling for a single order creates ripples that affect the orders around it. If an experimenter were to assign a treatment intervention to one of these orders while leaving neighboring orders as controls, the latter would evidently be contaminated. A common response to this problem is to randomize treatments at the level of a cluster, chosen so that interference can stay within it. In food and grocery delivery, that cluster is typically a geographical region. Since every order within a region sees the same treatme

## Outlier Handling at Scale in Experimentation

DevFeed: [Outlier Handling at Scale in Experimentation](<https://devfeed.tech/articles/outlier-handling-at-scale-in-experimentation-30453.md>)

Original publisher: [Read original article](<https://booking.ai/outlier-handling-at-scale-in-experimentation-a8bb140e1ab8?source=rss----4d265f07defc---4>)

Author: Margarida Moreira da Silva

Published: 2026-07-01T13:44:26Z

Content type: article

Language: en

Sources: [Booking.com Data Science](<https://devfeed.tech/sources/booking-com-data-science.md>)

Topics: [experiments](<https://devfeed.tech/topics/experiments.md>), [data](<https://devfeed.tech/topics/data.md>), [Simulation](<https://devfeed.tech/topics/simulation.md>), [Ground truth / benchmark quality](<https://devfeed.tech/topics/ground-truth-benchmark-quality.md>), [plotting](<https://devfeed.tech/topics/plotting.md>)

Tags: [data](<https://devfeed.tech/tags/data.md>), [experimentation](<https://devfeed.tech/tags/experimentation.md>), [experiments](<https://devfeed.tech/tags/experiments.md>), [false-positive](<https://devfeed.tech/tags/false-positive.md>), [outlier-detection](<https://devfeed.tech/tags/outlier-detection.md>), [simulation](<https://devfeed.tech/tags/simulation.md>), [variance](<https://devfeed.tech/tags/variance.md>)

### AI overview

The article examines how extreme values affect experimentation at Booking.com. It describes permutation tests and simulated A/A experiments for diagnosing distorted p-value distributions, and reports that increasing outlier magnitude and frequency can cause test failures.

### Source excerpt

At Booking.com, thousands of experiments run simultaneously across highly heterogeneous users, from individual travellers to large travel agencies. This means our experiment data regularly contains legitimate but extreme values. When these go unhandled, they distort the statistical conclusions we draw, leading us to scale ideas that don't create value, or to discard ones that do. So, we need outlier handling methods that are reliable, automated, and applicable across diverse metrics without manual intervention. The Problem When extreme values are present in experiment data, they can compromise the estimation of average treatment effects (ATE), leading to unreliable test results and reduced statistical power. Even a single observation can inflate variance enough to mask a real effect or produce a spurious one. In practice, this means we risk shipping changes that appear positive but are not, or killing promising features because noise masked their real effect. At Booking.com's scale, this increase in false conclusions quickly compounds into a meaningful impact on customer experience and business outcomes. A Diagnostic Tool: the Permutation Test One way to assess whether extreme values are distorting results is the permutation test. By permuting over experiment data, we generate hundreds of simulated AA experiments where we know the ground truth: there is no real effect. Plotting the resulting p-values, we expect a uniform distribution. If it instead looks skewed, the underlying data distribution is compromising the validity of results. Plot 1: P-value distributions from simulated A/A tests. Clean normally-distributed estimated effects produce a uniform distribution (left), while the presence of extreme outliers results in skewed p-values (right), indicating a distorted false positive rate.Simulation Evidence: What Drives Failure? We ran AA permutation tests across a range of simulated data distributions to understand when they fail (i.e. not show a uniform p-value di

## Debug and evaluate your AI app from your coding agent with Datadog Agent Observability

DevFeed: [Debug and evaluate your AI app from your coding agent with Datadog Agent Observability](<https://devfeed.tech/articles/debug-and-evaluate-your-ai-app-from-your-coding-agent-with-datadog-agent-observability-2263.md>)

Original publisher: [Read original article](<https://www.datadoghq.com/blog/debug-and-evaluate-your-ai-app-from-your-coding-agent/>)

Author: Michael Bevilacqua-Linn; Till W; Tanguy Renaudie; Mehul Sonowal; Gabriele Lorenzo; Alex Barksdale

Published: 2026-06-30T00:00:00Z

Content type: article

Language: en

Sources: [Datadog | The Monitor blog](<https://devfeed.tech/sources/datadog-the-monitor-blog.md>)

Topics: [agent observability](<https://devfeed.tech/topics/agent-observability.md>), [AI-assisted coding](<https://devfeed.tech/topics/ai-assisted-coding.md>), [AI Engineering](<https://devfeed.tech/topics/ai-engineering.md>), [observability](<https://devfeed.tech/topics/observability.md>), [MCP](<https://devfeed.tech/topics/mcp.md>), [MCP Server](<https://devfeed.tech/topics/mcp-server.md>), [Command-line interface](<https://devfeed.tech/topics/cli.md>), [LLM evaluation / benchmarking](<https://devfeed.tech/topics/llm-evaluation-benchmarking.md>), [datasets](<https://devfeed.tech/topics/datasets.md>), [Traces](<https://devfeed.tech/topics/traces.md>), [debug](<https://devfeed.tech/topics/debug.md>), [experiments](<https://devfeed.tech/topics/experiments.md>)

Tags: [agent-observability](<https://devfeed.tech/tags/agent-observability.md>), [agent-skills](<https://devfeed.tech/tags/agent-skills.md>), [agents](<https://devfeed.tech/tags/agents.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-engineering](<https://devfeed.tech/tags/ai-engineering.md>), [ai-observability](<https://devfeed.tech/tags/ai-observability.md>), [analysis](<https://devfeed.tech/tags/analysis.md>), [api](<https://devfeed.tech/tags/api.md>), [claude](<https://devfeed.tech/tags/claude.md>), [claude-code](<https://devfeed.tech/tags/claude-code.md>), [cli](<https://devfeed.tech/tags/cli.md>), [code](<https://devfeed.tech/tags/code.md>), [codex](<https://devfeed.tech/tags/codex.md>), [coding](<https://devfeed.tech/tags/coding.md>), [command-line](<https://devfeed.tech/tags/command-line.md>), [cursor](<https://devfeed.tech/tags/cursor.md>), [datasets](<https://devfeed.tech/tags/datasets.md>), [debug](<https://devfeed.tech/tags/debug.md>), [eval](<https://devfeed.tech/tags/eval.md>), [how-to](<https://devfeed.tech/tags/how-to.md>), [observability](<https://devfeed.tech/tags/observability.md>)

### AI overview

This article explains how to use Datadog Agent Observability from coding agents such as Claude Code, Cursor, and Codex CLI. It presents the Datadog MCP Server, Pup CLI, and Agent Skills as ways to access traces, evaluation results, experiment metrics, and other telemetry for classifying sessions, debugging production failures, creating evaluation datasets, and generating fixes.

### Source excerpt

Learn how to give your coding agent access to Datadog Agent Observability data to classify failures, run RCA, bootstrap evaluators, and generate fixes.

## How We Built an AI Agent to Clean Up Dead Code After A/B Tests

DevFeed: [How We Built an AI Agent to Clean Up Dead Code After A/B Tests](<https://devfeed.tech/articles/how-we-built-an-ai-agent-to-clean-up-dead-code-after-a-b-tests-26513.md>)

Original publisher: [Read original article](<https://medium.com/engineering-housing/how-we-built-an-ai-agent-to-clean-up-dead-code-after-a-b-tests-a5519af4892e?source=rss----3a69e32e2594---4>)

Author: Aseem Upadhyay

Published: 2026-06-23T10:54:43Z

Content type: article

Language: en

Sources: [Housing.com](<https://devfeed.tech/sources/housing-com.md>)

Topics: [AI Agent](<https://devfeed.tech/topics/ai-agent.md>), [experiments](<https://devfeed.tech/topics/experiments.md>), [Code](<https://devfeed.tech/topics/code.md>), [Pull Request](<https://devfeed.tech/topics/pull-request.md>), [context](<https://devfeed.tech/topics/context.md>), [Android](<https://devfeed.tech/topics/android.md>)

Tags: [agent](<https://devfeed.tech/tags/agent.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-agent](<https://devfeed.tech/tags/ai-agent.md>), [ai-agents-in-action](<https://devfeed.tech/tags/ai-agents-in-action.md>), [automated](<https://devfeed.tech/tags/automated.md>), [code](<https://devfeed.tech/tags/code.md>), [concurrent](<https://devfeed.tech/tags/concurrent.md>), [experiment](<https://devfeed.tech/tags/experiment.md>), [experiments](<https://devfeed.tech/tags/experiments.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [jira](<https://devfeed.tech/tags/jira.md>), [llm](<https://devfeed.tech/tags/llm.md>)

### AI overview

Housing.com describes building an AI-agent pipeline to help clean up code after A/B tests. The workflow interprets experiment tickets, reports experiment status, and applies instructions to code, while using scripts for deterministic steps and an LLM where judgment is required.

### Source excerpt

Photo by Microsoft Copilot on Unsplash At Housing.com, running product experiments is a continuous cycle. A/B tests go live, collect data, and eventually reach a conclusion. That's the exciting part. Then comes the mundane reality where someone has to clean up the code .i.e. remove a feature flag, promote a winning variant or revert the loser variant and finally raise a change request. Sounds simple? Maybe Is it tedious and quietly expensive? Yes! lifecycle of a taskThe Problem Worth Solving An experiment conclusion ticket typically lands on an engineer's desk looking something like this: Experiment: show_listing_map_widget Platform: Android Result: Negative - revert to control The job of the assigned engineer involves four distinct steps: Find every reference to the flag across the codebase. Delete the losing variant's code path. Trace every side-effect that only existed to support that variant Commit, open a PR, and comment on the Jira ticket. Step 3 is where the trap lies. Be it applying or removing a change, changing all the infrastructure code dependent on it could increase the complexity and risk of creating technical debt. But what if we automated a part of it? The AI Agent Pipelineupdated AI enabled lifecycle The problem statement became simple: Let stakeholders own the trigger. We built two agents to make it happen, 1. to interpret tickets and report experiment status 2. to take the instructions and code. Then came the hard part. Navigating Roadblocks The real complexity lies in building an AI agent that runs autonomously and serves different users across different use cases We found ourselves wrestling with questions we hadn't fully anticipated: How do we optimise on the tokens used per request? How do we handle concurrent requests? How do we ensure that the consistency in the output? So we went looking for answers.. Optimising Tokens per request Not every step needs AI. At each point in the workflow, we asked one question: is this operation deterministic

## Using Evaluation Frameworks with Agent Observability

DevFeed: [Using Evaluation Frameworks with Agent Observability](<https://devfeed.tech/articles/using-evaluation-frameworks-with-agent-observability-2318.md>)

Original publisher: [Read original article](<https://www.datadoghq.com/blog/using-evaluation-frameworks-with-agent-observability/>)

Author: Jennifer Mickel; Eddie Cai

Published: 2026-06-22T00:00:00Z

Content type: article

Language: en

Sources: [Datadog | The Monitor blog](<https://devfeed.tech/sources/datadog-the-monitor-blog.md>)

Topics: [agent observability](<https://devfeed.tech/topics/agent-observability.md>), [observability](<https://devfeed.tech/topics/observability.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [Pydantic](<https://devfeed.tech/topics/pydantic.md>), [experiments](<https://devfeed.tech/topics/experiments.md>), [Traces](<https://devfeed.tech/topics/traces.md>), [Retrieval Augmented Generation (RAG)](<https://devfeed.tech/topics/retrieval-augmented-generation-rag.md>)

Tags: [agent-observability](<https://devfeed.tech/tags/agent-observability.md>), [ai-observability](<https://devfeed.tech/tags/ai-observability.md>), [code](<https://devfeed.tech/tags/code.md>), [deployment](<https://devfeed.tech/tags/deployment.md>), [development](<https://devfeed.tech/tags/development.md>), [evals](<https://devfeed.tech/tags/evals.md>), [integration](<https://devfeed.tech/tags/integration.md>), [llm](<https://devfeed.tech/tags/llm.md>), [observability](<https://devfeed.tech/tags/observability.md>), [traces](<https://devfeed.tech/tags/traces.md>)

### AI overview

This article explains how Datadog Agent Observability integrates existing DeepEval and Pydantic Evals frameworks. It covers running evaluations in experiments, connecting evaluation scores to production traces, and continuously monitoring LLM evaluation quality across development and deployment.

### Source excerpt

Run DeepEval and Pydantic Evals natively in Datadog Agent Observability. Track regressions and connect eval scores to production traces.

## Fable 5 Launch Claims Tested Through Seven Experiments and 1,000+ Timed Runs

DevFeed: [Fable 5 Launch Claims Tested Through Seven Experiments and 1,000+ Timed Runs](<https://devfeed.tech/articles/claude-fable-5-the-ultimate-guide-for-pms-v3-39177.md>)

Original publisher: [Read original article](<https://www.productcompass.pm/p/claude-fable-5-guide>)

Author: Paweł Huryn

Published: 2026-06-11T04:28:43Z

Content type: comparison

Language: en

Sources: [The Product Compass](<https://devfeed.tech/sources/the-product-compass.md>)

Topics: [experiments](<https://devfeed.tech/topics/experiments.md>), [prompt](<https://devfeed.tech/topics/prompt.md>), [Claude](<https://devfeed.tech/topics/claude.md>)

Tags: [claude](<https://devfeed.tech/tags/claude.md>), [days](<https://devfeed.tech/tags/days.md>), [experiments](<https://devfeed.tech/tags/experiments.md>), [guide](<https://devfeed.tech/tags/guide.md>), [launch](<https://devfeed.tech/tags/launch.md>), [prompt](<https://devfeed.tech/tags/prompt.md>), [v3](<https://devfeed.tech/tags/v3.md>)

### AI overview

The article examines Fable 5 launch claims using seven experiments and more than 1,000 timed runs, covering claims that changed, the cost of producing a real finding, and an initial prompt to run.

### Source excerpt

Fable 5 is four days old. 7 experiments and 1,000+ timed runs later: the launch claims that flipped, what a real finding costs, and the first prompt you should run.

## AI-Driven Development Should Focus on Product Value, Not Just Faster Shipping

DevFeed: [AI-Driven Development Should Focus on Product Value, Not Just Faster Shipping](<https://devfeed.tech/articles/are-you-asking-the-real-questions-40032.md>)

Original publisher: [Read original article](<https://dpereira.substack.com/p/are-you-asking-the-real-questions>)

Author: David Pereira

Published: 2026-06-10T12:27:59Z

Content type: opinion

Language: en

Sources: [Untrapping Product Teams](<https://devfeed.tech/sources/untrapping-product-teams.md>)

Topics: [AI Development](<https://devfeed.tech/topics/ai-development.md>), [product analytics](<https://devfeed.tech/topics/product-analytics.md>), [session replay](<https://devfeed.tech/topics/session-replay.md>), [experiments](<https://devfeed.tech/topics/experiments.md>), [datadog](<https://devfeed.tech/topics/datadog.md>), [Software Engineering](<https://devfeed.tech/topics/software-engineering.md>), [ITMO](<https://devfeed.tech/topics/itmo.md>)

Tags: [agentic-ai](<https://devfeed.tech/tags/agentic-ai.md>), [ai](<https://devfeed.tech/tags/ai.md>), [analytics](<https://devfeed.tech/tags/analytics.md>), [datadog](<https://devfeed.tech/tags/datadog.md>), [product-analytics](<https://devfeed.tech/tags/product-analytics.md>), [session-replay](<https://devfeed.tech/tags/session-replay.md>), [software-engineering](<https://devfeed.tech/tags/software-engineering.md>)

### AI overview

The article argues that AI-assisted development should be judged by the value it creates for customers and businesses, not simply by shipping speed. It contrasts new products with existing products that have customers, legacy constraints, churn concerns, and growth goals, and suggests combining product analytics, session replay, and experiments to inform roadmaps.

### Source excerpt

The world is getting weird.

## Kotlin Multiplatform in Production: Two Real-World Use Cases from Booking.com

DevFeed: [Kotlin Multiplatform in Production: Two Real-World Use Cases from Booking.com](<https://devfeed.tech/articles/kotlin-multiplatform-in-production-two-real-world-use-cases-from-booking-com-23724.md>)

Original publisher: [Read original article](<https://medium.com/booking-com-development/kotlin-multiplatform-in-production-two-real-world-use-cases-from-booking-com-46ffe13a773d?source=rss----1c36c35f9c76---4>)

Author: Diego Gómez Olvera

Published: 2026-06-05T15:09:18Z

Content type: article

Language: en

Sources: [Booking.com Development - Medium](<https://devfeed.tech/sources/booking-com-development-medium.md>)

Topics: [Kotlin Multiplatform](<https://devfeed.tech/topics/kotlin-multiplatform.md>), [compose-multiplatform](<https://devfeed.tech/topics/compose-multiplatform.md>), [experiments](<https://devfeed.tech/topics/experiments.md>), [A/B Testing](<https://devfeed.tech/topics/a-b-testing.md>), [Android](<https://devfeed.tech/topics/android.md>), [iOS](<https://devfeed.tech/topics/ios.md>), [Design system](<https://devfeed.tech/topics/design-system.md>), [Mobile](<https://devfeed.tech/topics/mobile.md>), [Development](<https://devfeed.tech/topics/development.md>)

Tags: [a-b-testing](<https://devfeed.tech/tags/a-b-testing.md>), [android](<https://devfeed.tech/tags/android.md>), [booking](<https://devfeed.tech/tags/booking.md>), [bookingcom](<https://devfeed.tech/tags/bookingcom.md>), [compose](<https://devfeed.tech/tags/compose.md>), [compose-multiplatform](<https://devfeed.tech/tags/compose-multiplatform.md>), [concepts](<https://devfeed.tech/tags/concepts.md>), [consistency](<https://devfeed.tech/tags/consistency.md>), [data](<https://devfeed.tech/tags/data.md>), [development](<https://devfeed.tech/tags/development.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [experiment](<https://devfeed.tech/tags/experiment.md>), [experimentation](<https://devfeed.tech/tags/experimentation.md>), [experiments](<https://devfeed.tech/tags/experiments.md>), [ios](<https://devfeed.tech/tags/ios.md>), [java](<https://devfeed.tech/tags/java.md>), [jetpack-compose](<https://devfeed.tech/tags/jetpack-compose.md>), [kotlin](<https://devfeed.tech/tags/kotlin.md>), [kotlin-multiplatform](<https://devfeed.tech/tags/kotlin-multiplatform.md>), [mobile](<https://devfeed.tech/tags/mobile.md>), [multiplatform](<https://devfeed.tech/tags/multiplatform.md>), [objective-c](<https://devfeed.tech/tags/objective-c.md>)

### AI overview

This article describes two Booking.com engineering use cases for Kotlin Multiplatform and Compose Multiplatform: a shared experimentation library for consistent experiment assignments across Android and iOS, and hosting an Android design system in a web browser.

### Source excerpt

Introduction For the majority of Booking.com travelers, mobile is the primary channel for researching, planning, and booking trips. Recent data shows that over 80% of travelers rely on a mobile app during the research phase, with more than half of all bookings occurring on mobile devices. Consequently, the Android and iOS platforms are critical to the company's product strategy; engineering choices made here have significant repercussions for the entire organisation. To maintain agility at this scale, two elements must function in unison: Strict decision validation: At any time, Booking.com manages over 1,000 simultaneous experiments across its product suite, with hundreds active on mobile. Every minor adjustment undergoes A/B testing via our proprietary experimentation library before reaching the user. A unified design system ensures product consistency and makes design goals transparent to all contributors, not just maintenance engineers. This article examines two specific engineering challenges solved using Kotlin Multiplatform (KMP) and Compose Multiplatform (CMP): Developing a shared experimentation library to ensure uniform experiment assignments across Android and iOS. Using Compose Multiplatform to host our Android design system in a web browser, bridging the gap between design concepts and implementation. While both cases use the same underlying technology, each provides unique insights into multiplatform development. Use case 1: shared experimentation library on Android and iOSThe problem with two implementations Historically, our internal experimentation library, responsible for managing experiment assignments, evaluations, and tracking on mobile, was maintained as two distinct codebases: a mix of Java and Kotlin for Android and Objective-C for iOS. While intended to be identical, managing two languages with fluctuating team resources inevitably led to logic drift. Discrepancies in event-tracking and experiment-fetching behaviours emerged, though they wer

## How Braintrust turns customer requests into code with Codex

DevFeed: [How Braintrust turns customer requests into code with Codex](<https://devfeed.tech/articles/how-braintrust-turns-customer-requests-into-code-with-codex-6314.md>)

Original publisher: [Read original article](<https://openai.com/index/braintrust>)

Published: 2026-05-29T12:00:00Z

Content type: article

Language: en

Sources: [OpenAI News](<https://devfeed.tech/sources/openai-news.md>)

Topics: [codex](<https://devfeed.tech/topics/codex.md>), [observability](<https://devfeed.tech/topics/observability.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Code](<https://devfeed.tech/topics/code.md>), [experiments](<https://devfeed.tech/topics/experiments.md>), [Command-line interface](<https://devfeed.tech/topics/cli.md>), [Terminal](<https://devfeed.tech/topics/terminal.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [codex](<https://devfeed.tech/tags/codex.md>), [coding](<https://devfeed.tech/tags/coding.md>), [customer](<https://devfeed.tech/tags/customer.md>), [development](<https://devfeed.tech/tags/development.md>), [eval](<https://devfeed.tech/tags/eval.md>), [experimentation](<https://devfeed.tech/tags/experimentation.md>), [experiments](<https://devfeed.tech/tags/experiments.md>), [feature](<https://devfeed.tech/tags/feature.md>), [observability](<https://devfeed.tech/tags/observability.md>), [real-time](<https://devfeed.tech/tags/real-time.md>), [sandbox](<https://devfeed.tech/tags/sandbox.md>), [speed](<https://devfeed.tech/tags/speed.md>), [tool](<https://devfeed.tech/tags/tool.md>), [workflow](<https://devfeed.tech/tags/workflow.md>)

### AI overview

Braintrust engineers use Codex to turn customer feature requests into preview branches and working ideas within minutes. The article describes faster terminal performance, real-time customer iteration, and sandboxed experiments driven by tests.

### Source excerpt

How Braintrust engineers use Codex with GPT-5.5 to run experiments and code faster.

## Scaling Experimentation Quality at Booking.com

DevFeed: [Scaling Experimentation Quality at Booking.com](<https://devfeed.tech/articles/scaling-experimentation-quality-at-booking-com-30454.md>)

Original publisher: [Read original article](<https://booking.ai/scaling-experimentation-quality-at-booking-com-726152ee4ef0?source=rss----4d265f07defc---4>)

Author: Edgar Cano

Published: 2026-03-24T11:44:21Z

Content type: article

Language: en

Sources: [Booking.com Data Science](<https://devfeed.tech/sources/booking-com-data-science.md>)

Topics: [experiments](<https://devfeed.tech/topics/experiments.md>), [Development](<https://devfeed.tech/topics/development.md>), [decision-making](<https://devfeed.tech/topics/decision-making.md>), [human review](<https://devfeed.tech/topics/human-review.md>)

Tags: [best-practices](<https://devfeed.tech/tags/best-practices.md>), [data-science](<https://devfeed.tech/tags/data-science.md>), [decision-making](<https://devfeed.tech/tags/decision-making.md>), [experimentation](<https://devfeed.tech/tags/experimentation.md>), [featured](<https://devfeed.tech/tags/featured.md>), [human-review](<https://devfeed.tech/tags/human-review.md>), [metrics](<https://devfeed.tech/tags/metrics.md>), [product-development](<https://devfeed.tech/tags/product-development.md>), [quality](<https://devfeed.tech/tags/quality.md>)

### AI overview

Booking.com describes how it addressed declining experimentation quality as experiment volume grew. The article discusses arbitrary test durations, significance-seeking, and the trade-offs between enforcing standards and educating product teams.

### Source excerpt

Authors: Edgar Cano, Daisy Duursma, Nils Skotara, Melanie Mueller Figure 1. Three-pillar components of Booking.com's Experimentation Quality Experimentation is at the core of product development in Booking.com, powered by our in-house platform, "ET" (Experiment Tool). At any given moment, we run approximately 1,000 parallel experiments to evaluate product changes. These experiments or A/B tests allow teams to directly compare a new version of the website against the existing one, validating hypotheses about how specific changes impact important metrics. In our organization, these pitfalls became more evident as our experiment volume grew. We observed that experimenters might set an arbitrary "two-week" duration without thinking about sufficient power, or extend a test until results "became significant" or "trended positive." Knowing that this lack of consistency leads to flawed decision-making we dedicated significant effort to increasing the quality of our experimentation process, making Experimentation Quality a key KPI for our program. However, identifying the problem was only the start; the greater challenge is how to implement these standards across a large organization. Enforcement vs. Education When deciding how to scale quality, we faced a fundamental choice: Do we enforce strict controls or we rely on education. Ultimately, we left it to product teams to decide how to conduct their experiments. This choice entailed several trade-offs: Enforcement: Ensures comparability, consistency, and reliability. However, it comes at the cost of flexibility. There is a risk that people follow "rules" blindly without understanding the rationale. Education: Aims for a culture where experimenters understand the why behind best practices. This leads to better buy-in, allows teams to challenge methods, and highlights individual responsibility. It also prevents bottlenecking. If enforcement requires human review, it slows down development. However, education requires a massive

## Gremlin Release Roundup 2025: Reliability across AI, on-prem, and applications

DevFeed: [Gremlin Release Roundup 2025: Reliability across AI, on-prem, and applications](<https://devfeed.tech/articles/gremlin-release-roundup-2025-reliability-across-ai-on-prem-and-applications-11684.md>)

Original publisher: [Read original article](<https://www.gremlin.com/blog/release-roundup-2025>)

Author: Andre Newman

Published: 2025-12-15T00:00:00Z

Content type: release

Language: en

Sources: [Gremlin Blog](<https://devfeed.tech/sources/gremlin-blog.md>)

Topics: [on-prem](<https://devfeed.tech/topics/on-prem.md>), [Failure Flags](<https://devfeed.tech/topics/failure-flags.md>), [experiments](<https://devfeed.tech/topics/experiments.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [MCP Server](<https://devfeed.tech/topics/mcp-server.md>), [telemetry](<https://devfeed.tech/topics/telemetry.md>), [SRE](<https://devfeed.tech/topics/sre.md>)

Tags: [2025](<https://devfeed.tech/tags/2025.md>), [ai](<https://devfeed.tech/tags/ai.md>), [analysis](<https://devfeed.tech/tags/analysis.md>), [experiments](<https://devfeed.tech/tags/experiments.md>), [failure-flags](<https://devfeed.tech/tags/failure-flags.md>), [features](<https://devfeed.tech/tags/features.md>), [gremlin](<https://devfeed.tech/tags/gremlin.md>), [mcp](<https://devfeed.tech/tags/mcp.md>), [mcp-server](<https://devfeed.tech/tags/mcp-server.md>), [no-code](<https://devfeed.tech/tags/no-code.md>), [on-prem](<https://devfeed.tech/tags/on-prem.md>), [outages](<https://devfeed.tech/tags/outages.md>), [release](<https://devfeed.tech/tags/release.md>), [releases](<https://devfeed.tech/tags/releases.md>), [sre](<https://devfeed.tech/tags/sre.md>), [telemetry](<https://devfeed.tech/tags/telemetry.md>)

### AI overview

Gremlin's 2025 release roundup describes improvements aimed at preventing outages and making reliability testing easier. Highlights include Reliability Intelligence for analyzing failures and recommending remediations, a Gremlin MCP server for querying environments through an LLM, new experiments and Failure Flags capabilities, expanded platform support, streamlined onboarding, and web UI refinements.

### Source excerpt

This year's release roundup covers our new on-prem offering, new Failure Flags features, intelligent analysis of failed experiments, and much more.

[Next page](<https://devfeed.tech/topics/experiments.md?cursor=WyIyMDI1LTEyLTE1VDAwOjAwOjAwKzAwOjAwIiwgIjZjNTIyYjBkLTI3OWYtNDc1ZS1hZTQ0LTljZDVkNmZjYjk4ZCJd>)