# Terminal-Bench 4.0

Published articles for Terminal-Bench 4.0.

This is one page of public article previews, not the complete archive. Follow Next page to continue. Summaries are not the original full articles.

## Anthropic Releases Claude Sonnet 5.5: 70.6% on Terminal-Bench 4.0 at the Same $2/$10 Price

DevFeed: [Anthropic Releases Claude Sonnet 5.5: 70.6% on Terminal-Bench 4.0 at the Same $2/$10 Price](<https://devfeed.tech/articles/anthropic-releases-claude-sonnet-5-5-70-6-on-terminal-bench-4-0-at-the-same-2-10-price-61432.md>)

Original publisher: [Read original article](<https://www.marktechpost.com/2026/09/28/anthropic-releases-claude-sonnet-5-5-70-6-on-terminal-bench-4-0-at-the-same-2-10-price/>)

Author: Michal Sutter

Published: 2026-09-29T04:20:59Z

Content type: news

Language: en

Sources: [MarkTechPost](<https://devfeed.tech/sources/marktechpost.md>)

Topics: [Claude](<https://devfeed.tech/topics/claude.md>), [anthropic](<https://devfeed.tech/topics/anthropic.md>), [releases](<https://devfeed.tech/topics/releases.md>), [Terminal](<https://devfeed.tech/topics/terminal.md>)

Tags: [agentic-ai](<https://devfeed.tech/tags/agentic-ai.md>), [ai-agents](<https://devfeed.tech/tags/ai-agents.md>), [ai-shorts](<https://devfeed.tech/tags/ai-shorts.md>), [anthropic](<https://devfeed.tech/tags/anthropic.md>), [applications](<https://devfeed.tech/tags/applications.md>), [artificial-intelligence](<https://devfeed.tech/tags/artificial-intelligence.md>), [aws](<https://devfeed.tech/tags/aws.md>), [azure](<https://devfeed.tech/tags/azure.md>), [benchmarks](<https://devfeed.tech/tags/benchmarks.md>), [claude](<https://devfeed.tech/tags/claude.md>), [claude-api](<https://devfeed.tech/tags/claude-api.md>), [claude-sonnet-5](<https://devfeed.tech/tags/claude-sonnet-5.md>), [claude-sonnet-5-5](<https://devfeed.tech/tags/claude-sonnet-5-5.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [editors-pick](<https://devfeed.tech/tags/editors-pick.md>), [faster](<https://devfeed.tech/tags/faster.md>), [google-cloud](<https://devfeed.tech/tags/google-cloud.md>), [language-model](<https://devfeed.tech/tags/language-model.md>), [large-language-model](<https://devfeed.tech/tags/large-language-model.md>), [new-releases](<https://devfeed.tech/tags/new-releases.md>), [releases](<https://devfeed.tech/tags/releases.md>), [tech-news](<https://devfeed.tech/tags/tech-news.md>), [technology](<https://devfeed.tech/tags/technology.md>), [terminal](<https://devfeed.tech/tags/terminal.md>), [terminal-bench-4-0](<https://devfeed.tech/tags/terminal-bench-4-0.md>), [token](<https://devfeed.tech/tags/token.md>), [tokens](<https://devfeed.tech/tags/tokens.md>), [uncategorized](<https://devfeed.tech/tags/uncategorized.md>)

### AI overview

Anthropic has released Claude Sonnet 5.5, a model for everyday tasks, bug fixing, and document work. It scores 70.6% on Terminal-Bench 4.0, generates output more than 30% faster than Sonnet 5, and retains the same $2 per million input-token and $10 per million output-token pricing. Anthropic says token efficiency can reduce cost per task by up to 30%. The model is available through the Claude Platform, AWS, Google Cloud, and Microsoft Azure.

### Source excerpt

Anthropic has released Claude Sonnet 5.5, the second model in its Claude 5.5 family. It scores 70.6% on Terminal-Bench 4.0 and lands within 2 points of Opus 5.5 on GDPval-AA. It also generates output 30%+ faster than Sonnet 5 and keeps the same $2/$10 per million token price. Anthropic says cost per task falls by up to 30% because the model uses fewer tokens. You can deploy it today via the Claude API, AWS, Google Cloud and Azure. The post Anthropic Releases Claude Sonnet 5.5: 70.6% on Terminal-Bench 4.0 at the Same $2/$10 Price appeared first on MarkTechPost.

## Claude Sonnet 5.5 adds cyber safeguards and may route higher-risk requests to Sonnet 5

DevFeed: [Claude Sonnet 5.5 adds cyber safeguards and may route higher-risk requests to Sonnet 5](<https://devfeed.tech/articles/you-picked-claude-sonnet-5-5-but-anthropic-may-send-your-request-to-sonnet-5-in-higher-risk-situations-61395.md>)

Original publisher: [Read original article](<https://thenewstack.io/claude-sonnet-cyber-safeguards/>)

Author: Amanda Caswell

Published: 2026-09-28T21:34:23Z

Content type: news

Language: en

Sources: [The New Stack](<https://devfeed.tech/sources/the-new-stack.md>)

Topics: [Claude](<https://devfeed.tech/topics/claude.md>), [AI software supply chain security](<https://devfeed.tech/topics/ai-software-supply-chain-security.md>), [anthropic](<https://devfeed.tech/topics/anthropic.md>), [Vulnerabilities](<https://devfeed.tech/topics/vulnerabilities.md>)

Tags: [ai-agents](<https://devfeed.tech/tags/ai-agents.md>), [ai-models](<https://devfeed.tech/tags/ai-models.md>), [anthropic](<https://devfeed.tech/tags/anthropic.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [claude-sonnet-5](<https://devfeed.tech/tags/claude-sonnet-5.md>), [claude-sonnet-5-5](<https://devfeed.tech/tags/claude-sonnet-5-5.md>), [cybersecurity](<https://devfeed.tech/tags/cybersecurity.md>), [exploit-development](<https://devfeed.tech/tags/exploit-development.md>), [security](<https://devfeed.tech/tags/security.md>), [terminal-bench-4-0](<https://devfeed.tech/tags/terminal-bench-4-0.md>)

### AI overview

Anthropic launched Claude Sonnet 5.5 with cyber safeguards and fallback routing that may send higher-risk requests to Sonnet 5. The model scores well on the cited coding benchmark and has improved substantially on offensive security tasks, prompting safeguards similar to those used for more capable models. Its protections may also lead to more refusals on legitimate cybersecurity work.

### Source excerpt

Anthropic's Claude Sonnet 5.5, released Monday, is the first Sonnet model to launch with cyber safeguards and model fallbacks like The post You picked Claude Sonnet 5.5 -- but Anthropic may send your request to Sonnet 5 in "higher-risk" situations appeared first on The New Stack.

## GPT-6 Astra Release Guide: Benchmarks, $10/$50 Pricing, and How to Run It in Codex and OpenCode

DevFeed: [GPT-6 Astra Release Guide: Benchmarks, $10/$50 Pricing, and How to Run It in Codex and OpenCode](<https://devfeed.tech/articles/gpt-6-astra-release-guide-benchmarks-10-50-pricing-and-how-to-run-it-in-codex-and-opencode-61122.md>)

Original publisher: [Read original article](<https://www.developersdigest.tech/blog/gpt-6-astra-release-guide-2026>)

Author: Developers Digest

Published: 2026-09-28T00:00:00Z

Content type: tutorial

Language: en

Sources: [Developers Digest](<https://devfeed.tech/sources/developers-digest.md>)

Topics: [gpt-6-astra](<https://devfeed.tech/topics/gpt-6-astra.md>), [OpenAI](<https://devfeed.tech/topics/openai.md>), [LLM evaluation / benchmarking](<https://devfeed.tech/topics/llm-evaluation-benchmarking.md>), [codex](<https://devfeed.tech/topics/codex.md>)

Tags: [ai-models](<https://devfeed.tech/tags/ai-models.md>), [benchmarks](<https://devfeed.tech/tags/benchmarks.md>), [codex](<https://devfeed.tech/tags/codex.md>), [gpt-6](<https://devfeed.tech/tags/gpt-6.md>), [gpt-6-astra](<https://devfeed.tech/tags/gpt-6-astra.md>), [news](<https://devfeed.tech/tags/news.md>), [openai](<https://devfeed.tech/tags/openai.md>), [opencode](<https://devfeed.tech/tags/opencode.md>), [terminal-bench-4-0](<https://devfeed.tech/tags/terminal-bench-4-0.md>), [tokens](<https://devfeed.tech/tags/tokens.md>)

### AI overview

GPT-6 Astra is described as OpenAI's max-capability model, with a 1.05M-token context window, pricing of $10/$50 per million tokens, a Critical cyber rating, and new highs on Terminal-Bench 4.0 and OSWorld 2.0. The guide covers what the model is, what its benchmarks do and do not show, and verified commands for trying it in Codex and OpenCode.

### Source excerpt

GPT-6 Astra is OpenAI's max-capability model: 1.05M context, $10/$50 per million tokens, the first OpenAI model rated Critical for cyber, and new highs on Terminal-Bench 4.0 and OSWorld 2.0. What it is, what the benchmarks do and do not say, and the verified commands to try it.

## Claude Sonnet 5.5: Pricing, Benchmarks, API Changes, and Availability

DevFeed: [Claude Sonnet 5.5: Pricing, Benchmarks, API Changes, and Availability](<https://devfeed.tech/articles/claude-sonnet-5-5-developer-guide-pricing-benchmarks-and-the-five-api-changes-61480.md>)

Original publisher: [Read original article](<https://www.developersdigest.tech/blog/claude-sonnet-5-5-release-guide-2026>)

Author: Developers Digest

Published: 2026-09-28T00:00:00Z

Content type: article

Language: en

Sources: [Developers Digest](<https://devfeed.tech/sources/developers-digest.md>)

Topics: [Claude](<https://devfeed.tech/topics/claude.md>), [LLM evaluation / benchmarking](<https://devfeed.tech/topics/llm-evaluation-benchmarking.md>)

Tags: [ai-models](<https://devfeed.tech/tags/ai-models.md>), [anthropic](<https://devfeed.tech/tags/anthropic.md>), [api](<https://devfeed.tech/tags/api.md>), [benchmarks](<https://devfeed.tech/tags/benchmarks.md>), [claude](<https://devfeed.tech/tags/claude.md>), [claude-code](<https://devfeed.tech/tags/claude-code.md>), [claude-sonnet](<https://devfeed.tech/tags/claude-sonnet.md>), [claude-sonnet-5-5](<https://devfeed.tech/tags/claude-sonnet-5-5.md>), [developer](<https://devfeed.tech/tags/developer.md>), [developer-guide](<https://devfeed.tech/tags/developer-guide.md>), [github-copilot](<https://devfeed.tech/tags/github-copilot.md>), [news](<https://devfeed.tech/tags/news.md>), [opus-5](<https://devfeed.tech/tags/opus-5.md>), [opus-5-5](<https://devfeed.tech/tags/opus-5-5.md>), [pricing](<https://devfeed.tech/tags/pricing.md>), [terminal-bench-4-0](<https://devfeed.tech/tags/terminal-bench-4-0.md>), [tokens](<https://devfeed.tech/tags/tokens.md>), [vercel-ai-gateway](<https://devfeed.tech/tags/vercel-ai-gateway.md>)

### AI overview

Anthropic's Claude Sonnet 5.5 is described as a mid-tier model priced at $2/$10 per million tokens, with a 1M context window and a 70.6% Terminal-Bench 4.0 score. The guide covers pricing, five breaking API changes, and its positioning relative to Opus 5.5 and Sonnet 5; it also notes availability in GitHub Copilot and Vercel AI Gateway.

### Source excerpt

Claude Sonnet 5.5 (claude-sonnet-5-5) is Anthropic's new mid-tier model: $2/$10 per million tokens, 70.6% on Terminal-Bench 4.0, 1M context, now GA in GitHub Copilot and on Vercel AI Gateway. The pricing math, the five breaking API changes, and where it fits next to Opus 5.5 and Sonnet 5.