# experiments

Published articles for experiments.

This is one page of public article previews, not the complete archive. Follow Next page to continue. Summaries are not the original full articles.

## Agent and Model Evaluations in Gemini Enterprise Agent Platform are now GA

DevFeed: [Agent and Model Evaluations in Gemini Enterprise Agent Platform are now GA](<https://devfeed.tech/articles/agent-and-model-evaluations-in-gemini-enterprise-agent-platform-are-now-ga-4202.md>)

Original publisher: [Read original article](<https://developers.googleblog.com/agent-and-model-evaluations-in-gemini-enterprise-agent-platform-are-now-ga/>)

Author: Alex Martin; Dima Melnyk

Published: 2026-09-12T11:04:33.891311Z

Content type: release

Language: en

Sources: [Google Developers Blog](<https://devfeed.tech/sources/google-developers-blog.md>)

Topics: [ai observability](<https://devfeed.tech/topics/ai-observability.md>), [SDKs](<https://devfeed.tech/topics/sdks.md>), [ci](<https://devfeed.tech/topics/ci.md>)

Tags: [agent](<https://devfeed.tech/tags/agent.md>), [agents](<https://devfeed.tech/tags/agents.md>), [ci](<https://devfeed.tech/tags/ci.md>), [cli](<https://devfeed.tech/tags/cli.md>), [development](<https://devfeed.tech/tags/development.md>), [enterprise](<https://devfeed.tech/tags/enterprise.md>), [evaluation](<https://devfeed.tech/tags/evaluation.md>), [experiments](<https://devfeed.tech/tags/experiments.md>), [gemini](<https://devfeed.tech/tags/gemini.md>), [llm](<https://devfeed.tech/tags/llm.md>), [metrics](<https://devfeed.tech/tags/metrics.md>), [model](<https://devfeed.tech/tags/model.md>), [platform](<https://devfeed.tech/tags/platform.md>), [production](<https://devfeed.tech/tags/production.md>), [telemetry](<https://devfeed.tech/tags/telemetry.md>), [testing](<https://devfeed.tech/tags/testing.md>), [traces](<https://devfeed.tech/tags/traces.md>)

### AI overview

Gemini Enterprise Agent Platform's evaluation service is generally available. It provides consistent evaluation of agents and models across local experiments and production traffic, with pre-built metrics, adaptive rubrics, custom metrics, simulators, and workflow integrations.

### Source excerpt

Agent Platform's evaluation service is now generally available, providing developers with a unified engine to measure agent quality consistently across local development experiments and live production traffic. You can evaluate agents using over 20 pre-built metrics, DeepMind-backed adaptive rubrics, or custom code-based and LLM-as-a-judge metrics stored in a centralized, versioned registry. The service integrates directly into existing workflows via the Agent Platform SDK, agents-cli, and ADK, offering built-in user and environment simulators to automate complex multi-turn testing and streamline CI pipelines.

## Analyze your experiments in ChatGPT with the Datadog Experiments plugin

DevFeed: [Analyze your experiments in ChatGPT with the Datadog Experiments plugin](<https://devfeed.tech/articles/analyze-your-experiments-in-chatgpt-with-the-datadog-experiments-plugin-2239.md>)

Original publisher: [Read original article](<https://www.datadoghq.com/blog/chatgpt-datadog-experiments/>)

Author: Jonathan Fulton; Amy Zhou; Uday Tennety

Published: 2026-09-11T00:00:00Z

Content type: article

Language: en

Sources: [Datadog | The Monitor blog](<https://devfeed.tech/sources/datadog-the-monitor-blog.md>)

Topics: [experiments](<https://devfeed.tech/topics/experiments.md>), [data](<https://devfeed.tech/topics/data.md>), [Latency](<https://devfeed.tech/topics/latency.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [analysis](<https://devfeed.tech/tags/analysis.md>), [chatgpt](<https://devfeed.tech/tags/chatgpt.md>), [experimentation](<https://devfeed.tech/tags/experimentation.md>), [experiments](<https://devfeed.tech/tags/experiments.md>), [openai](<https://devfeed.tech/tags/openai.md>)

### AI overview

Datadog introduces a ChatGPT Work plugin that lets teams query current experiment results in plain language while retaining Datadog's statistical methods, guardrails, and connected context.

### Source excerpt

Learn how the Datadog Experiments OpenAI Data plugin lets your team read, question, and act on experiment results directly in ChatGPT.

## DiscoSign: Discourse-Aware Text to Sign Language Gloss Translation

DevFeed: [DiscoSign: Discourse-Aware Text to Sign Language Gloss Translation](<https://devfeed.tech/articles/discosign-discourse-aware-text-to-sign-language-gloss-translation-6728.md>)

Original publisher: [Read original article](<https://machinelearning.apple.com/research/discosign-gloss-translation>)

Published: 2026-09-11T00:00:00Z

Content type: article

Language: en

Sources: [Apple Machine Learning Research](<https://devfeed.tech/sources/apple-machine-learning-research.md>)

Topics: [AI, ML & Data Engineering](<https://devfeed.tech/topics/ai-ml-data-engineering.md>), [Language models](<https://devfeed.tech/topics/language-models.md>)

Tags: [evaluation](<https://devfeed.tech/tags/evaluation.md>), [experiments](<https://devfeed.tech/tags/experiments.md>), [framework](<https://devfeed.tech/tags/framework.md>), [llm](<https://devfeed.tech/tags/llm.md>), [metrics](<https://devfeed.tech/tags/metrics.md>)

### AI overview

DiscoSign is an LLM-based framework for translating text to sign-language gloss while preserving discourse-level coherence. It targets spatial coreference, Question-Answer Clauses, and consistent English-concept-to-ASL-sign mappings, with evaluation metrics for these dimensions.

### Source excerpt

Sign language processing systems have traditionally operated at the sentence level, ignoring critical discourse phenomena fundamental to sign language comprehension. We introduce DiscoSign, a computational approach for discourse-aware text to sign language gloss translation grounded in linguistic research. We address three key phenomena within our modular Large Language Model (LLM)-based translation framework: (i) spatial coreference resolution, where entities maintain consistent spatial locations throughout discourse; (ii) Question-Answer Clauses (QACs), pseudocleft structures serving...

## Inside the First Three.js Conference in Paris

DevFeed: [Inside the First Three.js Conference in Paris](<https://devfeed.tech/articles/inside-the-first-three-js-conference-in-paris-4348.md>)

Original publisher: [Read original article](<https://tympanus.net/codrops/2026/09/10/inside-the-first-three-js-conference-in-paris/>)

Author: Manoela Ilic

Published: 2026-09-10T08:30:11Z

Content type: news

Language: en

Sources: [Codrops](<https://devfeed.tech/sources/codrops.md>)

Topics: [Three.js community](<https://devfeed.tech/topics/three-js-community.md>), [webgpu](<https://devfeed.tech/topics/webgpu.md>)

Tags: [3d-web-development](<https://devfeed.tech/tags/3d-web-development.md>), [ai](<https://devfeed.tech/tags/ai.md>), [articles](<https://devfeed.tech/tags/articles.md>), [browser](<https://devfeed.tech/tags/browser.md>), [coding](<https://devfeed.tech/tags/coding.md>), [codrops](<https://devfeed.tech/tags/codrops.md>), [community](<https://devfeed.tech/tags/community.md>), [conference](<https://devfeed.tech/tags/conference.md>), [creative-coding](<https://devfeed.tech/tags/creative-coding.md>), [creative-development](<https://devfeed.tech/tags/creative-development.md>), [creative-technology](<https://devfeed.tech/tags/creative-technology.md>), [creativity](<https://devfeed.tech/tags/creativity.md>), [creators](<https://devfeed.tech/tags/creators.md>), [developers](<https://devfeed.tech/tags/developers.md>), [digital-experiences](<https://devfeed.tech/tags/digital-experiences.md>), [event](<https://devfeed.tech/tags/event.md>), [experiments](<https://devfeed.tech/tags/experiments.md>), [frontend-development](<https://devfeed.tech/tags/frontend-development.md>), [generative-art](<https://devfeed.tech/tags/generative-art.md>), [immersive-web](<https://devfeed.tech/tags/immersive-web.md>), [interactive-web-experiences](<https://devfeed.tech/tags/interactive-web-experiences.md>), [javascript-3d](<https://devfeed.tech/tags/javascript-3d.md>), [shaders](<https://devfeed.tech/tags/shaders.md>), [three-js](<https://devfeed.tech/tags/three-js.md>), [three-js-community](<https://devfeed.tech/tags/three-js-community.md>), [three-js-conference](<https://devfeed.tech/tags/three-js-conference.md>), [three-js-conference-paris](<https://devfeed.tech/tags/three-js-conference-paris.md>), [three-js-demos](<https://devfeed.tech/tags/three-js-demos.md>), [three-js-developers](<https://devfeed.tech/tags/three-js-developers.md>), [three-js-projects](<https://devfeed.tech/tags/three-js-projects.md>), [three-js-talks](<https://devfeed.tech/tags/three-js-talks.md>), [three-js-web-development](<https://devfeed.tech/tags/three-js-web-development.md>), [three-js-workshops](<https://devfeed.tech/tags/three-js-workshops.md>), [web-3d](<https://devfeed.tech/tags/web-3d.md>), [web-design](<https://devfeed.tech/tags/web-design.md>), [webgl](<https://devfeed.tech/tags/webgl.md>), [webgpu](<https://devfeed.tech/tags/webgpu.md>)

### AI overview

A live look at the first Three.js Conference in Paris, covering community moments, opening talks, a Zelda rebuild moving from WebGL to WebGPU, and a discussion about AI and coding.

### Source excerpt

A look inside the first-ever Three.js Conference, straight from Paris.

## Creating an AI Platform for classic ML online inference

DevFeed: [Creating an AI Platform for classic ML online inference](<https://devfeed.tech/articles/creating-an-ai-platform-for-classic-ml-online-inference-22589.md>)

Original publisher: [Read original article](<https://medium.com/amex-gbt-technology/creating-an-ai-platform-for-classic-ml-online-inference-e2165d68e18a?source=rss----60a0578f4096---4>)

Author: Rohith Leeladharan

Published: 2026-09-10T07:26:46Z

Content type: tutorial

Language: en

Sources: [Amex GBT Technology](<https://devfeed.tech/sources/amex-gbt-technology.md>)

Topics: [AI Platform](<https://devfeed.tech/topics/ai-platform.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [MLOps](<https://devfeed.tech/topics/mlops.md>), [Machine learning](<https://devfeed.tech/topics/machine-learning.md>), [model-deployment](<https://devfeed.tech/topics/model-deployment.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-platform](<https://devfeed.tech/tags/ai-platform.md>), [ai-platform-engineering](<https://devfeed.tech/tags/ai-platform-engineering.md>), [deploy](<https://devfeed.tech/tags/deploy.md>), [experiments](<https://devfeed.tech/tags/experiments.md>), [feature-store](<https://devfeed.tech/tags/feature-store.md>), [inference](<https://devfeed.tech/tags/inference.md>), [machine-learning](<https://devfeed.tech/tags/machine-learning.md>), [ml](<https://devfeed.tech/tags/ml.md>), [mlops](<https://devfeed.tech/tags/mlops.md>), [monitoring](<https://devfeed.tech/tags/monitoring.md>), [predictions](<https://devfeed.tech/tags/predictions.md>)

### AI overview

This article describes how American Express Global Business Travel built an AI platform for deploying classic machine-learning systems and supporting online inference. It explains the platform's requirements--simplicity, self-service, experimentation, and continuous improvement--and details the pre-process, predict, post-process pattern used by inference engines.

### Source excerpt

Introduction In 2021, we were given the mission to have AI Systems running in production. The team, instead of just following a classical MLOps process, that involves transforming a Jupyter notebook into a product running in production, decided to go further by creating a platform to deploy AI systems in production. The team decided the platform should respect these requirements: Simplicity: The code powering AI systems should be simple, readable, and easy to maintain -- less intricacy means fewer bugs in production and greater reliability. Self-service: Anyone should be able to build and deploy AI systems autonomously, without depending on a central team. Experimentation: The platform should make it easy to run and iterate on experiments. Continuous improvement: Data related to events and interactions within AI systems must be captured, enabling monitoring and continuous improvement over time. In this article, we will walk through the work done to build a platform that fulfills these four requirements. Background At American Express Global Business Travel, we use machine learning (ML) models for a variety of user experiences like ranking hotel and flight search results. Our ML models are wrapped in inference engines that handle both pre-processing of input data before we run a prediction with the model, and post-processing of output data before returning the output to the caller. The overall flow looks something like this: Figure 1: Handling an inference request A client service that would like the ML model's predictions provides necessary context about the request like which user the request is for. Then, optionally, the inference engine fetches any necessary features for inference from our feature store [part 1][part 2]. Finally, it pre-processes the data, runs the predictions using the trained ML model, and does any necessary post-processing of the model output before returning the response to the caller. We call this the pre-process, predict, post-process patter

## How we built Datadog Experiments

DevFeed: [How we built Datadog Experiments](<https://devfeed.tech/articles/how-we-built-datadog-experiments-2283.md>)

Original publisher: [Read original article](<https://www.datadoghq.com/blog/how-we-built-datadog-experiments/>)

Author: Chas DeVeas; Aaron Silverman; Tyler Buffington; Jonathan Fulton; Taylor Overturf

Published: 2026-09-10T00:00:00Z

Content type: article

Language: en

Sources: [Datadog | The Monitor blog](<https://devfeed.tech/sources/datadog-the-monitor-blog.md>)

Topics: [experiments](<https://devfeed.tech/topics/experiments.md>), [real user monitoring](<https://devfeed.tech/topics/real-user-monitoring.md>)

Tags: [a-b-testing](<https://devfeed.tech/tags/a-b-testing.md>), [acquisition](<https://devfeed.tech/tags/acquisition.md>), [data-analytics](<https://devfeed.tech/tags/data-analytics.md>), [data-engineering](<https://devfeed.tech/tags/data-engineering.md>), [experimentation](<https://devfeed.tech/tags/experimentation.md>), [experiments](<https://devfeed.tech/tags/experiments.md>), [metrics](<https://devfeed.tech/tags/metrics.md>), [observability](<https://devfeed.tech/tags/observability.md>), [product-analytics](<https://devfeed.tech/tags/product-analytics.md>), [real-user-monitoring](<https://devfeed.tech/tags/real-user-monitoring.md>)

### AI overview

Datadog describes rebuilding its experimentation platform to speed up confident A/B-test decisions. The article explains a flexible CUPED approach that reduces metric variance and can be applied to segments.

### Source excerpt

Datadog Experiments shortens the time from result to decision with CUPED on percentiles, verifiable warehouse results, and near real-time RUM metrics.

## Testing application resilience with Amazon SQS and AWS Fault Injection Service

DevFeed: [Testing application resilience with Amazon SQS and AWS Fault Injection Service](<https://devfeed.tech/articles/testing-application-resilience-with-amazon-sqs-and-aws-fault-injection-service-4651.md>)

Original publisher: [Read original article](<https://aws.amazon.com/blogs/architecture/testing-application-resilience-with-amazon-sqs-and-aws-fault-injection-service/>)

Author: Richard Whitworth

Published: 2026-09-09T21:33:59Z

Content type: tutorial

Language: en

Sources: [AWS Architecture Blog](<https://devfeed.tech/sources/aws-architecture-blog.md>)

Topics: [Amazon Simple Queue Service (SQS)](<https://devfeed.tech/topics/amazon-simple-queue-service-sqs.md>), [AWS Fault Injection Service (FIS)](<https://devfeed.tech/topics/aws-fault-injection-service-fis.md>), [Chaos Engineering](<https://devfeed.tech/topics/chaos-engineering.md>), [AWS IAM](<https://devfeed.tech/topics/aws-iam.md>), [AWS Identity and Access Management (IAM)](<https://devfeed.tech/topics/aws-identity-and-access-management-iam.md>)

Tags: [advanced-300](<https://devfeed.tech/tags/advanced-300.md>), [amazon-cloudwatch](<https://devfeed.tech/tags/amazon-cloudwatch.md>), [amazon-simple-queue-service-sqs](<https://devfeed.tech/tags/amazon-simple-queue-service-sqs.md>), [amazon-sqs](<https://devfeed.tech/tags/amazon-sqs.md>), [automation](<https://devfeed.tech/tags/automation.md>), [aws](<https://devfeed.tech/tags/aws.md>), [aws-fault-injection-service-fis](<https://devfeed.tech/tags/aws-fault-injection-service-fis.md>), [experiments](<https://devfeed.tech/tags/experiments.md>), [how-to](<https://devfeed.tech/tags/how-to.md>), [iam](<https://devfeed.tech/tags/iam.md>), [observability](<https://devfeed.tech/tags/observability.md>), [resilience](<https://devfeed.tech/tags/resilience.md>), [technical-how-to](<https://devfeed.tech/tags/technical-how-to.md>), [testing](<https://devfeed.tech/tags/testing.md>)

### AI overview

A tutorial on testing application resilience when Amazon SQS data-plane operations fail. It uses AWS Fault Injection Service and Systems Manager Automation to progressively deny queue access, evaluate recovery and observability with CloudWatch metrics, and avoid IAM deny-policy lockouts.

### Source excerpt

Learn how to use AWS Fault Injection Service and AWS Systems Manager Automation to run progressive chaos experiments against Amazon SQS queues. Validate that your retry logic, circuit breakers, and dead-letter queues actually work under failure before a real outage hits production.

## How we built data-driven AI Golden Paths at Datadog

DevFeed: [How we built data-driven AI Golden Paths at Datadog](<https://devfeed.tech/articles/how-we-built-data-driven-ai-golden-paths-at-datadog-2226.md>)

Original publisher: [Read original article](<https://www.datadoghq.com/blog/ai-development-golden-paths/>)

Author: Addie Beach; Rui Martins Lacerda

Published: 2026-09-09T00:00:00Z

Content type: article

Language: en

Sources: [Datadog | The Monitor blog](<https://devfeed.tech/sources/datadog-the-monitor-blog.md>)

Topics: [AI Bots](<https://devfeed.tech/topics/ai-bots.md>)

Tags: [agent](<https://devfeed.tech/tags/agent.md>), [agent-observability](<https://devfeed.tech/tags/agent-observability.md>), [ai](<https://devfeed.tech/tags/ai.md>), [cli](<https://devfeed.tech/tags/cli.md>), [cost](<https://devfeed.tech/tags/cost.md>), [dashboards](<https://devfeed.tech/tags/dashboards.md>), [development](<https://devfeed.tech/tags/development.md>), [documentation](<https://devfeed.tech/tags/documentation.md>), [experiments](<https://devfeed.tech/tags/experiments.md>), [performance](<https://devfeed.tech/tags/performance.md>), [security](<https://devfeed.tech/tags/security.md>), [skills](<https://devfeed.tech/tags/skills.md>), [workflows](<https://devfeed.tech/tags/workflows.md>)

### AI overview

Datadog describes a data-driven process for establishing AI Golden Paths: standardized AI-agent workflows governed by controls such as skills, hooks, and tests. The approach evaluates controls for their effects on code quality, security, token usage, cost, and agent performance.

### Source excerpt

See how a Datadog guild achieved 13% faster agent runs by building Golden Paths for AI-assisted development using controls, experiments, and dashboards.

## How GPT-5.6 Sol helps run quantum computing experiments

DevFeed: [How GPT-5.6 Sol helps run quantum computing experiments](<https://devfeed.tech/articles/how-gpt-5-6-sol-helps-run-quantum-computing-experiments-6352.md>)

Original publisher: [Read original article](<https://openai.com/index/codex-quantum-computing-experiments>)

Published: 2026-09-08T17:00:00Z

Content type: article

Language: en

Sources: [OpenAI News](<https://devfeed.tech/sources/openai-news.md>)

Topics: [AI research agents](<https://devfeed.tech/topics/ai-research-agents.md>)

Tags: [agents](<https://devfeed.tech/tags/agents.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-agents](<https://devfeed.tech/tags/ai-agents.md>), [analysis](<https://devfeed.tech/tags/analysis.md>), [applied-ai](<https://devfeed.tech/tags/applied-ai.md>), [codex](<https://devfeed.tech/tags/codex.md>), [experiments](<https://devfeed.tech/tags/experiments.md>), [gpt](<https://devfeed.tech/tags/gpt.md>), [quantum](<https://devfeed.tech/tags/quantum.md>), [quantum-computing](<https://devfeed.tech/tags/quantum-computing.md>)

### AI overview

An MIT researcher connected GPT-5.6 Sol through Codex to laboratory software for superconducting-qubit experiments. The system ran routine measurements, analyzed results, and selected next steps, reducing the need for constant supervision during calibration workflows.

### Source excerpt

See how an MIT researcher uses GPT-5.6 Sol with Codex to autonomously run quantum computing experiments, analyze results, and calibrate qubits.

## Safety for Whom? Refusing the Right Subset of a Topic, Not the Whole Topic

DevFeed: [Safety for Whom? Refusing the Right Subset of a Topic, Not the Whole Topic](<https://devfeed.tech/articles/safety-for-whom-refusing-the-right-subset-of-a-topic-not-the-whole-topic-7025.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/MultiverseComputingCAI/safety-for-whom>)

Author: Antonio Tiene; Alejo Lopez Avila; Iker García-Ferrero

Published: 2026-09-08T14:23:07Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [AI Chat](<https://devfeed.tech/topics/ai-chat.md>)

Tags: [blog](<https://devfeed.tech/tags/blog.md>), [deployment](<https://devfeed.tech/tags/deployment.md>), [experiments](<https://devfeed.tech/tags/experiments.md>), [hugging-face](<https://devfeed.tech/tags/hugging-face.md>), [llm](<https://devfeed.tech/tags/llm.md>), [model](<https://devfeed.tech/tags/model.md>), [policy](<https://devfeed.tech/tags/policy.md>), [safety](<https://devfeed.tech/tags/safety.md>), [training](<https://devfeed.tech/tags/training.md>)

### AI overview

The article examines LLM safety policies that refuse harmful requests within a topic while continuing to answer benign requests in that same topic.

### Source excerpt

A Blog post by Multiverse Computing on Hugging Face

## Coordinate product launches with Datadog

DevFeed: [Coordinate product launches with Datadog](<https://devfeed.tech/articles/coordinate-product-launches-with-datadog-2244.md>)

Original publisher: [Read original article](<https://www.datadoghq.com/blog/coordinate-product-launches-with-datadog/>)

Author: Milene Darnis; Adam Virani

Published: 2026-09-08T00:00:00Z

Content type: tutorial

Language: en

Sources: [Datadog | The Monitor blog](<https://devfeed.tech/sources/datadog-the-monitor-blog.md>)

Topics: [Instrumentation](<https://devfeed.tech/topics/instrumentation.md>), [experiments](<https://devfeed.tech/topics/experiments.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [bits-ai](<https://devfeed.tech/tags/bits-ai.md>), [digital-experience-monitoring](<https://devfeed.tech/tags/digital-experience-monitoring.md>), [experimentation](<https://devfeed.tech/tags/experimentation.md>), [experiments](<https://devfeed.tech/tags/experiments.md>), [feature-flags](<https://devfeed.tech/tags/feature-flags.md>), [instrumentation](<https://devfeed.tech/tags/instrumentation.md>), [launch](<https://devfeed.tech/tags/launch.md>), [product-analytics](<https://devfeed.tech/tags/product-analytics.md>), [session-replay](<https://devfeed.tech/tags/session-replay.md>)

### AI overview

A tutorial on using Datadog Product Analytics Launches to plan product releases, define measurement questions, create tracking plans, and identify missing events and properties before rollout.

### Source excerpt

Learn how to turn a product brief and feature flag into a connected launch workflow for instrumentation, experimentation, QA, and reporting.

## Research acceleration: The view inside OpenAI

DevFeed: [Research acceleration: The view inside OpenAI](<https://devfeed.tech/articles/research-acceleration-the-view-inside-openai-6628.md>)

Original publisher: [Read original article](<https://openai.com/index/research-acceleration-view-inside-openai>)

Published: 2026-09-06T08:00:00Z

Content type: article

Language: en

Sources: [OpenAI News](<https://devfeed.tech/sources/openai-news.md>)

Topics: [AI Research](<https://devfeed.tech/topics/ai-research.md>), [AI-assisted coding](<https://devfeed.tech/topics/ai-assisted-coding.md>), [coding](<https://devfeed.tech/topics/coding.md>)

Tags: [agents](<https://devfeed.tech/tags/agents.md>), [ai-research](<https://devfeed.tech/tags/ai-research.md>), [coding](<https://devfeed.tech/tags/coding.md>), [deep-learning](<https://devfeed.tech/tags/deep-learning.md>), [experiments](<https://devfeed.tech/tags/experiments.md>), [frontier-ai](<https://devfeed.tech/tags/frontier-ai.md>), [openai](<https://devfeed.tech/tags/openai.md>), [research](<https://devfeed.tech/tags/research.md>)

### AI overview

OpenAI describes how coding agents are being used throughout its AI research workflow, with reported increases in code contribution, experiment execution, task complexity, and success rates. It frames this progress as a step toward supervised automated AI research while emphasizing human control over research priorities and deployment decisions.

### Source excerpt

Inside OpenAI, coding agents are reshaping AI research. Explore early data on agent usage, experiment velocity, task complexity, and research acceleration.

## How we make AI coding more cost efficient without sacrificing task quality

DevFeed: [How we make AI coding more cost efficient without sacrificing task quality](<https://devfeed.tech/articles/how-we-make-ai-coding-more-cost-efficient-without-sacrificing-task-quality-79.md>)

Original publisher: [Read original article](<https://github.blog/ai-and-ml/github-copilot/how-we-make-ai-coding-more-cost-efficient-without-sacrificing-task-quality/>)

Author: Erik Kristensen

Published: 2026-09-02T18:00:00Z

Content type: article

Language: en

Sources: [GitHub](<https://devfeed.tech/sources/github.md>), [GitHub Engineering](<https://devfeed.tech/sources/github-engineering.md>)

Topics: [GitHub Copilot](<https://devfeed.tech/topics/github-copilot.md>), [AI-assisted coding](<https://devfeed.tech/topics/ai-assisted-coding.md>), [GitHub Copilot CLI](<https://devfeed.tech/topics/github-copilot-cli.md>), [coding](<https://devfeed.tech/topics/coding.md>), [GitHub Copilot app](<https://devfeed.tech/topics/github-copilot-app.md>), [GitHub Copilot code review](<https://devfeed.tech/topics/github-copilot-code-review.md>)

Tags: [agentic-ai](<https://devfeed.tech/tags/agentic-ai.md>), [agentic-coding](<https://devfeed.tech/tags/agentic-coding.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-coding](<https://devfeed.tech/tags/ai-coding.md>), [ai-ml](<https://devfeed.tech/tags/ai-ml.md>), [architecture-optimization](<https://devfeed.tech/tags/architecture-optimization.md>), [benchmarks](<https://devfeed.tech/tags/benchmarks.md>), [compression](<https://devfeed.tech/tags/compression.md>), [cost](<https://devfeed.tech/tags/cost.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [experiments](<https://devfeed.tech/tags/experiments.md>), [github-copilot](<https://devfeed.tech/tags/github-copilot.md>), [github-copilot-cli](<https://devfeed.tech/tags/github-copilot-cli.md>), [github-copilot-code-review](<https://devfeed.tech/tags/github-copilot-code-review.md>), [llms](<https://devfeed.tech/tags/llms.md>), [prompt-engineering](<https://devfeed.tech/tags/prompt-engineering.md>)

### AI overview

GitHub Copilot's efficiency work focuses on total task cost and duration rather than minimizing tokens in individual tool responses. The article describes evaluating changes with coding benchmarks and controlled experiments, and explains that overly compressed output can cause agents to repeat work.

### Source excerpt

Why shorter outputs can cost more, and how GitHub Copilot reduces wasted work across the complete coding task. The post How we make AI coding more cost efficient without sacrificing task quality appeared first on The GitHub Blog.

## From traces to experiments: A loop for improving AI agents

DevFeed: [From traces to experiments: A loop for improving AI agents](<https://devfeed.tech/articles/from-traces-to-experiments-a-loop-for-improving-ai-agents-2276.md>)

Original publisher: [Read original article](<https://www.datadoghq.com/blog/from-traces-to-experiments-a-loop-for-improving-ai-agents/>)

Author: Adam Virani; Lukas Goetz-Weiss; Natasha Silva

Published: 2026-09-01T00:00:00Z

Content type: tutorial

Language: en

Sources: [Datadog | The Monitor blog](<https://devfeed.tech/sources/datadog-the-monitor-blog.md>)

Topics: [ai observability](<https://devfeed.tech/topics/ai-observability.md>), [AI Bots](<https://devfeed.tech/topics/ai-bots.md>)

Tags: [agent-observability](<https://devfeed.tech/tags/agent-observability.md>), [agents](<https://devfeed.tech/tags/agents.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-agents](<https://devfeed.tech/tags/ai-agents.md>), [ai-observability](<https://devfeed.tech/tags/ai-observability.md>), [evaluation](<https://devfeed.tech/tags/evaluation.md>), [experiments](<https://devfeed.tech/tags/experiments.md>), [latency](<https://devfeed.tech/tags/latency.md>), [optimization](<https://devfeed.tech/tags/optimization.md>), [production](<https://devfeed.tech/tags/production.md>), [telemetry](<https://devfeed.tech/tags/telemetry.md>), [traces](<https://devfeed.tech/tags/traces.md>)

### AI overview

The article explains how teams can use AI-agent trace data, evaluations, and production experiments to identify performance issues and test whether changes improve outcomes.

### Source excerpt

Learn how to read AI agent traces as a roadmap and how to run production experiments that measure whether improvements hold in production.

## Visualize how CUPED adjusts experiment results with Datadog

DevFeed: [Visualize how CUPED adjusts experiment results with Datadog](<https://devfeed.tech/articles/visualize-how-cuped-adjusts-experiment-results-with-datadog-2246.md>)

Original publisher: [Read original article](<https://www.datadoghq.com/blog/cuped-adjustments-visualization/>)

Author: Tyler Buffington; Lukas Goetz-Weiss; Ryan Lucht

Published: 2026-09-01T00:00:00Z

Content type: tutorial

Language: en

Sources: [Datadog | The Monitor blog](<https://devfeed.tech/sources/datadog-the-monitor-blog.md>)

Topics: [experiments](<https://devfeed.tech/topics/experiments.md>)

Tags: [analysis](<https://devfeed.tech/tags/analysis.md>), [digital-experience-monitoring](<https://devfeed.tech/tags/digital-experience-monitoring.md>), [experiment](<https://devfeed.tech/tags/experiment.md>), [experiments](<https://devfeed.tech/tags/experiments.md>), [how-to](<https://devfeed.tech/tags/how-to.md>), [metrics](<https://devfeed.tech/tags/metrics.md>), [product-analytics](<https://devfeed.tech/tags/product-analytics.md>), [visualization](<https://devfeed.tech/tags/visualization.md>)

### AI overview

This tutorial explains Datadog Experiments' CUPED adjustments visualization, which breaks the difference between raw and CUPED-adjusted experiment lift into individual covariate contributions.

### Source excerpt

Learn how Datadog visualizes CUPED adjustments so you can trace which covariates change experiment lift estimates and improve precision.

## Harness RT Agents Detect Resilience Risks and Generate Tests for CD Pipelines and Kubernetes Workloads

DevFeed: [Harness RT Agents Detect Resilience Risks and Generate Tests for CD Pipelines and Kubernetes Workloads](<https://devfeed.tech/articles/automate-resilience-testing-with-agents-13398.md>)

Original publisher: [Read original article](<https://www.harness.io/blog/find-resilience-risks-automatically-then-confirm-them>)

Author: Uma Mukkara

Published: 2026-08-24T00:00:00Z

Content type: release

Language: en

Sources: [Harness Blog](<https://devfeed.tech/sources/harness-blog.md>)

Topics: [Resilience](<https://devfeed.tech/topics/resilience.md>), [Testing](<https://devfeed.tech/topics/testing.md>), [Continuous Delivery (CD)](<https://devfeed.tech/topics/continuous-delivery.md>), [Kubernetes](<https://devfeed.tech/topics/kubernetes.md>), [dashboards](<https://devfeed.tech/topics/dashboards.md>)

Tags: [agentic](<https://devfeed.tech/tags/agentic.md>), [agents](<https://devfeed.tech/tags/agents.md>), [blog](<https://devfeed.tech/tags/blog.md>), [chaos](<https://devfeed.tech/tags/chaos.md>), [configuration](<https://devfeed.tech/tags/configuration.md>), [continuous-delivery](<https://devfeed.tech/tags/continuous-delivery.md>), [dashboards](<https://devfeed.tech/tags/dashboards.md>), [deployment](<https://devfeed.tech/tags/deployment.md>), [experiments](<https://devfeed.tech/tags/experiments.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [insights](<https://devfeed.tech/tags/insights.md>), [kubernetes](<https://devfeed.tech/tags/kubernetes.md>), [load](<https://devfeed.tech/tags/load.md>), [pipelines](<https://devfeed.tech/tags/pipelines.md>), [product](<https://devfeed.tech/tags/product.md>), [resilience](<https://devfeed.tech/tags/resilience.md>), [services](<https://devfeed.tech/tags/services.md>), [teams](<https://devfeed.tech/tags/teams.md>), [testing](<https://devfeed.tech/tags/testing.md>), [tests](<https://devfeed.tech/tags/tests.md>)

### AI overview

Harness announces an update to Resilience Testing called RT Agents. The agents analyze CD pipelines and Kubernetes workloads for resilience risks, recommend the testing needed to confirm those risks, and can generate and run chaos experiments or load tests and interpret the results.

### Source excerpt

RT Agents detect resilience risk in your CD pipelines and Kubernetes workloads, then generate and run chaos experiments or load tests to confirm it. | Blog

## Two ways to measure the cumulative impact of experiments

DevFeed: [Two ways to measure the cumulative impact of experiments](<https://devfeed.tech/articles/two-ways-to-measure-the-cumulative-impact-of-experiments-2316.md>)

Original publisher: [Read original article](<https://www.datadoghq.com/blog/two-ways-to-measure-cumulative-impact/>)

Author: Lukas Goetz-Weiss; Eddie Cai

Published: 2026-08-18T00:00:00Z

Content type: opinion

Language: en

Sources: [Datadog | The Monitor blog](<https://devfeed.tech/sources/datadog-the-monitor-blog.md>)

Topics: [experiments](<https://devfeed.tech/topics/experiments.md>), [Statistics](<https://devfeed.tech/topics/statistics.md>), [Data Science](<https://devfeed.tech/topics/data-science.md>)

Tags: [analysis](<https://devfeed.tech/tags/analysis.md>), [experimentation](<https://devfeed.tech/tags/experimentation.md>), [experiments](<https://devfeed.tech/tags/experiments.md>), [product](<https://devfeed.tech/tags/product.md>), [statistics](<https://devfeed.tech/tags/statistics.md>), [testing](<https://devfeed.tech/tags/testing.md>)

### AI overview

This article explains why summing the observed lift from winning experiments overstates cumulative impact because of the winner's curse. It presents two established approaches: a randomized holdout that measures the combined effect directly, and a statistical correction that aggregates existing experiment estimates. Datadog's Cumulative Impact feature applies the correction without requiring a quarter-long holdout and can analyze an entire experimentation program or a filtered team subset.

### Source excerpt

Summing individual wins overstates true impact. See two accurate methods, holdouts and Datadog's Cumulative Impact, and how to choose between them.

## LaunchDarkly is now available on the Vercel Marketplace

DevFeed: [LaunchDarkly is now available on the Vercel Marketplace](<https://devfeed.tech/articles/launchdarkly-is-now-available-on-the-vercel-marketplace-997.md>)

Original publisher: [Read original article](<https://vercel.com/changelog/launchdarkly-is-now-available-on-the-vercel-marketplace>)

Author: Sam Halstead

Published: 2026-08-11T00:00:00Z

Content type: release

Language: en

Sources: [Vercel News](<https://devfeed.tech/sources/vercel-news.md>)

Topics: [feature flags](<https://devfeed.tech/topics/feature-flags.md>), [Vercel](<https://devfeed.tech/topics/vercel.md>), [experiments](<https://devfeed.tech/topics/experiments.md>), [SDKs](<https://devfeed.tech/topics/sdks.md>)

Tags: [agent](<https://devfeed.tech/tags/agent.md>), [experiments](<https://devfeed.tech/tags/experiments.md>), [feature-flags](<https://devfeed.tech/tags/feature-flags.md>), [sdk](<https://devfeed.tech/tags/sdk.md>), [vercel](<https://devfeed.tech/tags/vercel.md>)

### AI overview

LaunchDarkly is now available through the Vercel Marketplace, providing feature flags with local evaluation, release targeting, experiments, metrics, rollback controls, and Vercel Toolbar management.

### Source excerpt

LaunchDarkly is now available on the Vercel Marketplace, allowing you to quickly get started with feature flags without additional setup. You can: Sync flags into Global Config and evaluate them locally Target releases by user, attribute, or segment Run experiments with metrics, and roll back with a kill switch View and override flags from the Vercel Toolbar with the Flags Explorer To get started, run vercel install launchdarkly, add the @flags-sdk/launchdarkly adapter, and declare a flag with the Flags SDK: Using a coding agent? Hand it this prompt: Add LaunchDarkly from the Vercel Marketplace, or read the adapter docs. Read more

## Tailscale's Remy Guercio on what comes after token maxing

DevFeed: [Tailscale's Remy Guercio on what comes after token maxing](<https://devfeed.tech/articles/tailscale-s-remy-guercio-on-what-comes-after-token-maxing-16050.md>)

Original publisher: [Read original article](<https://workos.com/blog/remy-guercio-tailscale-token-maxing-roi-aie-2026>)

Author: WorkOS

Published: 2026-08-05T23:49:35Z

Content type: article

Language: en

Sources: [WorkOS Blog](<https://devfeed.tech/sources/workos-blog.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [AI-assisted coding](<https://devfeed.tech/topics/ai-assisted-coding.md>), [experiments](<https://devfeed.tech/topics/experiments.md>)

Tags: [agents](<https://devfeed.tech/tags/agents.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-gateway](<https://devfeed.tech/tags/ai-gateway.md>), [analysis](<https://devfeed.tech/tags/analysis.md>), [cost](<https://devfeed.tech/tags/cost.md>), [experiments](<https://devfeed.tech/tags/experiments.md>), [model](<https://devfeed.tech/tags/model.md>), [tailscale](<https://devfeed.tech/tags/tailscale.md>), [token](<https://devfeed.tech/tags/token.md>)

### AI overview

An interview with Tailscale's Remy Guercio examines the shift from maximizing token usage to maximizing return on investment in AI development. It covers coding-agent adoption, rising bills across multiple AI tools, cost per task, model experimentation, and the visibility an AI gateway can provide.

### Source excerpt

Tailscale's Remy Guercio tells Michael Grinich why the next 12 months of AI are ROI maxing: cost per task, model experiments, and what a gateway can see.

## How Speechify serves 500,000 dynamic pages to 60 million users on Vercel

DevFeed: [How Speechify serves 500,000 dynamic pages to 60 million users on Vercel](<https://devfeed.tech/articles/how-speechify-serves-500-000-dynamic-pages-to-60-million-users-on-vercel-749.md>)

Original publisher: [Read original article](<https://vercel.com/blog/how-speechify-serves-50000-dynamic-pages-to-60-million-users-on-vercel>)

Author: Susan Aziz

Published: 2026-07-15T04:00:00Z

Content type: article

Language: en

Sources: [Vercel News](<https://devfeed.tech/sources/vercel-news.md>)

Topics: [Vercel](<https://devfeed.tech/topics/vercel.md>), [Caching](<https://devfeed.tech/topics/caching.md>), [Next.js](<https://devfeed.tech/topics/next-js.md>), [Cache](<https://devfeed.tech/topics/cache.md>), [Security](<https://devfeed.tech/topics/security.md>), [incident](<https://devfeed.tech/topics/incident.md>)

Tags: [cache](<https://devfeed.tech/tags/cache.md>), [experiments](<https://devfeed.tech/tags/experiments.md>), [incident](<https://devfeed.tech/tags/incident.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [next-js](<https://devfeed.tech/tags/next-js.md>), [scale](<https://devfeed.tech/tags/scale.md>), [security](<https://devfeed.tech/tags/security.md>), [vercel](<https://devfeed.tech/tags/vercel.md>)

### AI overview

A case study of how Speechify rebuilt its infrastructure on Next.js and Vercel to serve more than 500,000 dynamic pages across 40-plus languages to 60 million users. It describes caching, auto-scaling, cost reduction, uptime, and Instant Rollbacks for safer frequent deployments.

### Source excerpt

Speechify on Vercel 500,000+ pages served across 40+ languages Cut costs 50% by auto-scaling with Fluid compute Zero user impact on bad deploys with Instant Rollbacks Speechify started as a tool for people with dyslexia. Cliff Weitzman, Founder & CEO, built it because reading was challenging and audio made it much easier. That initial use case led Speechify to tens of millions of users and an Apple Design Award for inclusion. The product has since grown into something much larger: an AI work platform where 60 million people listen to documents, delegate tasks to agents, and complete work entirely through voice. But the journey wasn't easy. Early on, they got hacked, and for half a day, every visitor was redirected to a casino website. After that, Denis Chernobai, Head of Growth Engineering, knew it was time to re-evaluate their infrastructure stack. He decided to rebuild from scratch on Next.js and Vercel, and they ended up serving 40 times more pages, reaching a global audience three times larger, and achieving a 50% cost reduction. Since migrating to Vercel, Speechify has maintained 99.99% uptime and hasn't had a single security incident. The cost of serving 500,000 dynamic pages Speechify's website is its primary growth engine, with 10,000 base pages translated into 40+ languages across onboarding funnels, localized landing pages, and pricing experiments that change constantly. Static generation isn't an option because the content changes too frequently. But serving everything dynamically means every visitor is a potential database read, which compounds fast at hundreds of thousands of visits every day. Vercel's Data Cache, ISR, and Next.js Cache Components resolve both problems at once. Pages render dynamically on first visit, cache immediately, and serve from the closest point of presence until something changes. The result is a global growth engine that scales without the infrastructure costs scaling with it. The risk of shipping to 60 million usersInstant Rol

## GLM-5.2: Considerations for enterprise teams starting out with open-weight models

DevFeed: [GLM-5.2: Considerations for enterprise teams starting out with open-weight models](<https://devfeed.tech/articles/glm-5-2-considerations-for-enterprise-teams-starting-out-with-open-weight-models-33584.md>)

Original publisher: [Read original article](<https://blog.scottlogic.com/2026/07/08/glm-5-2-considerations-for-enterprise-teams-starting-out-with-open-weight-models.html>)

Author: Robat Williams

Published: 2026-07-08T10:00:00Z

Content type: article

Language: en

Sources: [Scott Logic](<https://devfeed.tech/sources/scott-logic.md>)

Topics: [Agentic development](<https://devfeed.tech/topics/agentic-development.md>), [developer setup](<https://devfeed.tech/topics/developer-setup.md>), [AI Development](<https://devfeed.tech/topics/ai-development.md>), [experiments](<https://devfeed.tech/topics/experiments.md>), [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [Claude](<https://devfeed.tech/topics/claude.md>)

Tags: [agent](<https://devfeed.tech/tags/agent.md>), [agentic](<https://devfeed.tech/tags/agentic.md>), [agentic-development](<https://devfeed.tech/tags/agentic-development.md>), [artificial-intelligence](<https://devfeed.tech/tags/artificial-intelligence.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [claude](<https://devfeed.tech/tags/claude.md>), [developer-setup](<https://devfeed.tech/tags/developer-setup.md>), [experiments](<https://devfeed.tech/tags/experiments.md>), [open-source](<https://devfeed.tech/tags/open-source.md>)

### AI overview

The article describes how a small team prepared an open-weight agentic development setup for AI productivity experiments. It explains the team's model-selection criteria, distinguishes open-weight models from locally run models, and reports choosing GLM-5.2 after comparing DeepSeek V4, Kimi K-2.6, and other models using benchmarks and practical trials.

### Source excerpt

Setting yourself up to try out open-weight models for agentic development isn't difficult, but it isn't as straightforward as downloading a coding agent from one of the handful of well-known AI vendors. In preparation for the latest round of our AI productivity experiments, we've recently been through this process. Read on for the choices we made, the considerations at play, and what made our situation unusual.

## The effect distribution: The missing piece in experimentation programs

DevFeed: [The effect distribution: The missing piece in experimentation programs](<https://devfeed.tech/articles/the-effect-distribution-the-missing-piece-in-experimentation-programs-2266.md>)

Original publisher: [Read original article](<https://www.datadoghq.com/blog/effect-distribution-in-experimentation/>)

Author: Tyler Buffington

Published: 2026-07-02T00:00:00Z

Content type: article

Language: en

Sources: [Datadog | The Monitor blog](<https://devfeed.tech/sources/datadog-the-monitor-blog.md>)

Topics: [experiments](<https://devfeed.tech/topics/experiments.md>)

Tags: [a-b-testing](<https://devfeed.tech/tags/a-b-testing.md>), [experiment](<https://devfeed.tech/tags/experiment.md>), [experimentation](<https://devfeed.tech/tags/experimentation.md>), [experiments](<https://devfeed.tech/tags/experiments.md>), [testing](<https://devfeed.tech/tags/testing.md>)

### AI overview

The article explains why effect distributions are essential for interpreting results across experimentation programs. It shows how statistically significant results can all be false positives when true effects are concentrated at zero, and introduces the challenge of estimating true effects from noisy observed effects.

### Source excerpt

Learn about the importance of considering the effect distribution when running experiments.

## Variance Reduction Below the Randomization Grain

DevFeed: [Variance Reduction Below the Randomization Grain](<https://devfeed.tech/articles/variance-reduction-below-the-randomization-grain-20111.md>)

Original publisher: [Read original article](<https://tech.instacart.com/variance-reduction-below-the-randomization-grain-31719f87a7d2?source=rss----587883b5d2ee---4>)

Author: Tilman Drerup

Published: 2026-07-01T16:28:36Z

Content type: article

Language: en

Sources: [Instacart](<https://devfeed.tech/sources/instacart.md>)

Topics: [experiments](<https://devfeed.tech/topics/experiments.md>)

Tags: [capacity](<https://devfeed.tech/tags/capacity.md>), [causal-inference](<https://devfeed.tech/tags/causal-inference.md>), [economics](<https://devfeed.tech/tags/economics.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [experiment](<https://devfeed.tech/tags/experiment.md>), [experimentation](<https://devfeed.tech/tags/experimentation.md>), [experiments](<https://devfeed.tech/tags/experiments.md>), [machine-learning](<https://devfeed.tech/tags/machine-learning.md>), [marketplaces](<https://devfeed.tech/tags/marketplaces.md>), [metrics](<https://devfeed.tech/tags/metrics.md>), [reduce](<https://devfeed.tech/tags/reduce.md>), [science](<https://devfeed.tech/tags/science.md>), [techniques](<https://devfeed.tech/tags/techniques.md>), [variance](<https://devfeed.tech/tags/variance.md>)

### AI overview

This article explains how marketplace experiments can reduce metric variance below the level at which treatment is randomized. It describes cluster-level randomization for containing interference and shows how fine-grained outcome predictability can improve statistical power and reduce experimentation time.

### Source excerpt

Sergio Camelo, Caitlin Kearns, Matias Cersosimo, and Tilman Drerup As artificial intelligence increases the velocity of engineering and science teams, experimental throughput is set to become a bottleneck for many product decisions. Many companies can now build faster than they can experiment, with queues of good ideas running the risk of not being tested because of lack of experimental capacity. This problem is particularly severe in marketplaces, where the presence of spillover and cannibalization effects between experimental units requires cluster-level randomization techniques. That randomization, in turn, has the unfortunate tendency to substantially reduce statistical power and slow down experimentation. In this post, we show that the predictability of outcomes at fine grains can be exploited to reduce the variance of aggregate metrics, even when experiments themselves are run at a coarse level. Since statistical power depends on metric variability, this yields considerable reductions in experimentation time. The Interference Problem In marketplace settings, behavior and outcomes for individual participants are inherently intertwined. In a delivery marketplace like Instacart, for example, the dispatch system solves a bipartite matching problem between shoppers and customer orders. Since assignments are global and interdependent, matching an order to one shopper means that the same order cannot be matched to another shopper. As a result, changing the handling for a single order creates ripples that affect the orders around it. If an experimenter were to assign a treatment intervention to one of these orders while leaving neighboring orders as controls, the latter would evidently be contaminated. A common response to this problem is to randomize treatments at the level of a cluster, chosen so that interference can stay within it. In food and grocery delivery, that cluster is typically a geographical region. Since every order within a region sees the same treatme

## Outlier Handling at Scale in Experimentation

DevFeed: [Outlier Handling at Scale in Experimentation](<https://devfeed.tech/articles/outlier-handling-at-scale-in-experimentation-30453.md>)

Original publisher: [Read original article](<https://booking.ai/outlier-handling-at-scale-in-experimentation-a8bb140e1ab8?source=rss----4d265f07defc---4>)

Author: Margarida Moreira da Silva

Published: 2026-07-01T13:44:26Z

Content type: article

Language: en

Sources: [Booking.com Data Science](<https://devfeed.tech/sources/booking-com-data-science.md>)

Topics: [experiments](<https://devfeed.tech/topics/experiments.md>), [data](<https://devfeed.tech/topics/data.md>), [Simulation](<https://devfeed.tech/topics/simulation.md>), [Ground truth / benchmark quality](<https://devfeed.tech/topics/ground-truth-benchmark-quality.md>), [plotting](<https://devfeed.tech/topics/plotting.md>)

Tags: [data](<https://devfeed.tech/tags/data.md>), [experimentation](<https://devfeed.tech/tags/experimentation.md>), [experiments](<https://devfeed.tech/tags/experiments.md>), [false-positive](<https://devfeed.tech/tags/false-positive.md>), [outlier-detection](<https://devfeed.tech/tags/outlier-detection.md>), [simulation](<https://devfeed.tech/tags/simulation.md>), [variance](<https://devfeed.tech/tags/variance.md>)

### AI overview

The article examines how extreme values affect experimentation at Booking.com. It describes permutation tests and simulated A/A experiments for diagnosing distorted p-value distributions, and reports that increasing outlier magnitude and frequency can cause test failures.

### Source excerpt

At Booking.com, thousands of experiments run simultaneously across highly heterogeneous users, from individual travellers to large travel agencies. This means our experiment data regularly contains legitimate but extreme values. When these go unhandled, they distort the statistical conclusions we draw, leading us to scale ideas that don't create value, or to discard ones that do. So, we need outlier handling methods that are reliable, automated, and applicable across diverse metrics without manual intervention. The Problem When extreme values are present in experiment data, they can compromise the estimation of average treatment effects (ATE), leading to unreliable test results and reduced statistical power. Even a single observation can inflate variance enough to mask a real effect or produce a spurious one. In practice, this means we risk shipping changes that appear positive but are not, or killing promising features because noise masked their real effect. At Booking.com's scale, this increase in false conclusions quickly compounds into a meaningful impact on customer experience and business outcomes. A Diagnostic Tool: the Permutation Test One way to assess whether extreme values are distorting results is the permutation test. By permuting over experiment data, we generate hundreds of simulated AA experiments where we know the ground truth: there is no real effect. Plotting the resulting p-values, we expect a uniform distribution. If it instead looks skewed, the underlying data distribution is compromising the validity of results. Plot 1: P-value distributions from simulated A/A tests. Clean normally-distributed estimated effects produce a uniform distribution (left), while the presence of extreme outliers results in skewed p-values (right), indicating a distorted false positive rate.Simulation Evidence: What Drives Failure? We ran AA permutation tests across a range of simulated data distributions to understand when they fail (i.e. not show a uniform p-value di

[Next page](<https://devfeed.tech/tags/experiments.md?cursor=WyIyMDI2LTA3LTAxVDEzOjQ0OjI2KzAwOjAwIiwgIjI4YmU0ODI3LWQ5ZjItNDRjOS04ZmQ5LTRlMDMwMDY4NGZmZiJd>)