# Building better AI benchmarks: How many raters are enough?

DevFeed: [Building better AI benchmarks: How many raters are enough?](<https://devfeed.tech/articles/building-better-ai-benchmarks-how-many-raters-are-enough-6752.md>)

Original publisher: [Read original article](<https://research.google/blog/building-better-ai-benchmarks-how-many-raters-are-enough/>)

Published: 2026-03-31T16:16:00Z

Content type: article

Language: en

Sources: [The latest research from Google](<https://devfeed.tech/sources/the-latest-research-from-google.md>)

Topics: [LLM evaluation / benchmarking](<https://devfeed.tech/topics/llm-evaluation-benchmarking.md>), [Human-AI evaluation](<https://devfeed.tech/topics/human-ai-evaluation.md>), [Ground truth / benchmark quality](<https://devfeed.tech/topics/ground-truth-benchmark-quality.md>), [benchmarking](<https://devfeed.tech/topics/benchmarking.md>), [data](<https://devfeed.tech/topics/data.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [algorithms-theory](<https://devfeed.tech/tags/algorithms-theory.md>), [benchmarking](<https://devfeed.tech/tags/benchmarking.md>), [benchmarks](<https://devfeed.tech/tags/benchmarks.md>), [cost](<https://devfeed.tech/tags/cost.md>), [evaluation](<https://devfeed.tech/tags/evaluation.md>), [machine-intelligence](<https://devfeed.tech/tags/machine-intelligence.md>), [machine-learning](<https://devfeed.tech/tags/machine-learning.md>), [research](<https://devfeed.tech/tags/research.md>)

## AI overview

Google Research presents an evaluation framework for machine-learning models that balances the number of rated items with the number of human raters per item. The research addresses reproducibility, human disagreement, benchmark quality, and the cost of collecting evaluation data, arguing that the common practice of using one to five raters per item can miss meaningful disagreement.

## Source excerpt

Algorithms & Theory