# Claude performed best on a new benchmark for 'agents that build agents'. But it passed fewer than a quarter of the tests.

DevFeed: [Claude performed best on a new benchmark for 'agents that build agents'. But it passed fewer than a quarter of the tests.](<https://devfeed.tech/articles/claude-performed-best-on-a-new-benchmark-for-agents-that-build-agents-but-it-passed-fewer-than-a-quarter-of-the-tests-8472.md>)

Original publisher: [Read original article](<https://thenewstack.io/claude-build-agents-benchmark/>)

Author: Paul Sawers

Published: 2026-09-09T20:14:09Z

Content type: news

Language: en

Sources: [The New Stack](<https://devfeed.tech/sources/the-new-stack.md>)

Topics: [Claude](<https://devfeed.tech/topics/claude.md>), [coding](<https://devfeed.tech/topics/coding.md>), [anthropic](<https://devfeed.tech/topics/anthropic.md>), [Software](<https://devfeed.tech/topics/software.md>)

Tags: [agents](<https://devfeed.tech/tags/agents.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-agents](<https://devfeed.tech/tags/ai-agents.md>), [ai-engineering](<https://devfeed.tech/tags/ai-engineering.md>), [ai-models](<https://devfeed.tech/tags/ai-models.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [claude](<https://devfeed.tech/tags/claude.md>), [claude-code](<https://devfeed.tech/tags/claude-code.md>), [codex](<https://devfeed.tech/tags/codex.md>), [coding](<https://devfeed.tech/tags/coding.md>), [large-language-models](<https://devfeed.tech/tags/large-language-models.md>), [leaderboard](<https://devfeed.tech/tags/leaderboard.md>), [testing](<https://devfeed.tech/tags/testing.md>)

## AI overview

Hyper-𝜏-bench evaluates whether AI developer agents can build customer-service agents from simulated business materials. Claude Opus 5 in Claude Code led the six tested configurations at 23.9%, while none exceeded 25%.

## Source excerpt

AI models now power all manner of agents, from coding assistants that write and debug software to customer service systems The post Claude performed best on a new benchmark for 'agents that build agents'. But it passed fewer than a quarter of the tests. appeared first on The New Stack.