# Scaling Experimentation Quality at Booking.com

DevFeed: [Scaling Experimentation Quality at Booking.com](<https://devfeed.tech/articles/scaling-experimentation-quality-at-booking-com-30454.md>)

Original publisher: [Read original article](<https://booking.ai/scaling-experimentation-quality-at-booking-com-726152ee4ef0?source=rss----4d265f07defc---4>)

Author: Edgar Cano

Published: 2026-03-24T11:44:21Z

Content type: article

Language: en

Sources: [Booking.com Data Science](<https://devfeed.tech/sources/booking-com-data-science.md>)

Topics: [experiments](<https://devfeed.tech/topics/experiments.md>), [Development](<https://devfeed.tech/topics/development.md>), [decision-making](<https://devfeed.tech/topics/decision-making.md>), [human review](<https://devfeed.tech/topics/human-review.md>)

Tags: [best-practices](<https://devfeed.tech/tags/best-practices.md>), [data-science](<https://devfeed.tech/tags/data-science.md>), [decision-making](<https://devfeed.tech/tags/decision-making.md>), [experimentation](<https://devfeed.tech/tags/experimentation.md>), [featured](<https://devfeed.tech/tags/featured.md>), [human-review](<https://devfeed.tech/tags/human-review.md>), [metrics](<https://devfeed.tech/tags/metrics.md>), [product-development](<https://devfeed.tech/tags/product-development.md>), [quality](<https://devfeed.tech/tags/quality.md>)

## AI overview

Booking.com describes how it addressed declining experimentation quality as experiment volume grew. The article discusses arbitrary test durations, significance-seeking, and the trade-offs between enforcing standards and educating product teams.

## Source excerpt

Authors: Edgar Cano, Daisy Duursma, Nils Skotara, Melanie Mueller Figure 1. Three-pillar components of Booking.com's Experimentation Quality Experimentation is at the core of product development in Booking.com, powered by our in-house platform, "ET" (Experiment Tool). At any given moment, we run approximately 1,000 parallel experiments to evaluate product changes. These experiments or A/B tests allow teams to directly compare a new version of the website against the existing one, validating hypotheses about how specific changes impact important metrics. In our organization, these pitfalls became more evident as our experiment volume grew. We observed that experimenters might set an arbitrary "two-week" duration without thinking about sufficient power, or extend a test until results "became significant" or "trended positive." Knowing that this lack of consistency leads to flawed decision-making we dedicated significant effort to increasing the quality of our experimentation process, making Experimentation Quality a key KPI for our program. However, identifying the problem was only the start; the greater challenge is how to implement these standards across a large organization. Enforcement vs. Education When deciding how to scale quality, we faced a fundamental choice: Do we enforce strict controls or we rely on education. Ultimately, we left it to product teams to decide how to conduct their experiments. This choice entailed several trade-offs: Enforcement: Ensures comparability, consistency, and reliability. However, it comes at the cost of flexibility. There is a risk that people follow "rules" blindly without understanding the rationale. Education: Aims for a culture where experimenters understand the why behind best practices. This leads to better buy-in, allows teams to challenge methods, and highlights individual responsibility. It also prevents bottlenecking. If enforcement requires human review, it slows down development. However, education requires a massive