# Separating signal from noise in coding evaluations

DevFeed: [Separating signal from noise in coding evaluations](<https://devfeed.tech/articles/separating-signal-from-noise-in-coding-evaluations-6648.md>)

Original publisher: [Read original article](<https://openai.com/index/separating-signal-from-noise-coding-evaluations>)

Published: 2026-07-08T13:00:00Z

Content type: article

Language: en

Sources: [OpenAI News](<https://devfeed.tech/sources/openai-news.md>)

Topics: [AI-assisted coding](<https://devfeed.tech/topics/ai-assisted-coding.md>), [Developer Tools](<https://devfeed.tech/topics/developer-tools.md>)

Tags: [agentic-coding](<https://devfeed.tech/tags/agentic-coding.md>), [ai-models](<https://devfeed.tech/tags/ai-models.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [data-quality](<https://devfeed.tech/tags/data-quality.md>), [evaluation](<https://devfeed.tech/tags/evaluation.md>), [research](<https://devfeed.tech/tags/research.md>), [testing](<https://devfeed.tech/tags/testing.md>)

## AI overview

OpenAI audits SWE-Bench Pro and estimates that about 30% of its tasks are broken, identifying strict, underspecified, low-coverage, and misleading tests as sources of unreliable coding-evaluation results.

## Source excerpt

A new analysis from OpenAI reveals issues in SWE-Bench Pro, a popular coding benchmark, raising concerns about reliability and accuracy in evaluating AI models.