# How enabling two settings tripled our scores on the ARC-AGI-3 benchmark

DevFeed: [How enabling two settings tripled our scores on the ARC-AGI-3 benchmark](<https://devfeed.tech/articles/how-enabling-two-settings-tripled-our-scores-on-the-arc-agi-3-benchmark-6461.md>)

Original publisher: [Read original article](<https://openai.com/index/how-two-settings-tripled-our-arc-agi-3-scores>)

Published: 2026-07-29T15:00:00Z

Content type: article

Language: en

Sources: [OpenAI News](<https://devfeed.tech/sources/openai-news.md>)

Topics: [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [AI Models](<https://devfeed.tech/topics/ai-models.md>), [API](<https://devfeed.tech/topics/api.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [ChatGPT](<https://devfeed.tech/topics/chatgpt.md>), [codex](<https://devfeed.tech/topics/codex.md>)

Tags: [agents](<https://devfeed.tech/tags/agents.md>), [ai-agents](<https://devfeed.tech/tags/ai-agents.md>), [ai-models](<https://devfeed.tech/tags/ai-models.md>), [analysis](<https://devfeed.tech/tags/analysis.md>), [api](<https://devfeed.tech/tags/api.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [chatgpt](<https://devfeed.tech/tags/chatgpt.md>), [codex](<https://devfeed.tech/tags/codex.md>), [comparisons](<https://devfeed.tech/tags/comparisons.md>), [efficiency](<https://devfeed.tech/tags/efficiency.md>), [performance](<https://devfeed.tech/tags/performance.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>), [research](<https://devfeed.tech/tags/research.md>), [tokens](<https://devfeed.tech/tags/tokens.md>)

## AI overview

The article explains how enabling retained reasoning and compaction in the harness tripled GPT-5.6 Sol's scores on the ARC-AGI-3 benchmark while reducing output tokens sixfold. It argues that benchmark results depend not only on the model but also on API settings, harness design, and prompting.

## Source excerpt

How two API settings improved GPT-5.6 performance on ARC-AGI-3, boosting scores and efficiency by retaining reasoning and enabling compaction.