# StateM Reports 95.3% Raw Accuracy on Terminal-Bench 2.1 by Scaling the Evaluation Harness

DevFeed: [StateM Reports 95.3% Raw Accuracy on Terminal-Bench 2.1 by Scaling the Evaluation Harness](<https://devfeed.tech/articles/terminal-bench-shows-harness-scaling-is-the-coding-agent-benchmark-now-56205.md>)

Original publisher: [Read original article](<https://www.developersdigest.tech/blog/long-horizon-terminal-bench-agent-evals>)

Author: Developers Digest

Published: 2026-07-14T00:00:00Z

Content type: opinion

Language: en

Sources: [Developers Digest](<https://devfeed.tech/sources/developers-digest.md>)

Topics: [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [Terminal](<https://devfeed.tech/topics/terminal.md>)

Tags: [agent](<https://devfeed.tech/tags/agent.md>), [ai-agents](<https://devfeed.tech/tags/ai-agents.md>), [ai-coding](<https://devfeed.tech/tags/ai-coding.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [benchmarks](<https://devfeed.tech/tags/benchmarks.md>), [coding-agent](<https://devfeed.tech/tags/coding-agent.md>), [developer-workflow](<https://devfeed.tech/tags/developer-workflow.md>), [evals](<https://devfeed.tech/tags/evals.md>), [harness](<https://devfeed.tech/tags/harness.md>), [model](<https://devfeed.tech/tags/model.md>), [recovery](<https://devfeed.tech/tags/recovery.md>), [scaling](<https://devfeed.tech/tags/scaling.md>), [state](<https://devfeed.tech/tags/state.md>)

## AI overview

StateM reports 95.3% raw accuracy on Terminal-Bench 2.1 after scaling the evaluation harness around the model. The article argues that runbooks, state, and recovery loops are increasingly important for coding-agent teams, alongside model choice.

## Source excerpt

StateM pushes Terminal-Bench 2.1 to 95.3% raw accuracy by scaling the harness around the model. The lesson for coding-agent teams is that runbooks, state, and recovery loops now matter as much as model choice.