# Evaluating chain-of-thought monitorability

DevFeed: [Evaluating chain-of-thought monitorability](<https://devfeed.tech/articles/evaluating-chain-of-thought-monitorability-6396.md>)

Original publisher: [Read original article](<https://openai.com/index/evaluating-chain-of-thought-monitorability>)

Published: 2025-12-18T12:00:00Z

Content type: article

Language: en

Sources: [OpenAI News](<https://devfeed.tech/sources/openai-news.md>)

Topics: [Chain-of-thought](<https://devfeed.tech/topics/chain-of-thought.md>), [LLM evaluation / benchmarking](<https://devfeed.tech/topics/llm-evaluation-benchmarking.md>), [Language models](<https://devfeed.tech/topics/language-models.md>), [OpenAI](<https://devfeed.tech/topics/openai.md>), [Reinforcement learning](<https://devfeed.tech/topics/reinforcement-learning.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [chain-of-thought](<https://devfeed.tech/tags/chain-of-thought.md>), [evaluation](<https://devfeed.tech/tags/evaluation.md>), [language-models](<https://devfeed.tech/tags/language-models.md>), [openai](<https://devfeed.tech/tags/openai.md>), [reinforcement-learning](<https://devfeed.tech/tags/reinforcement-learning.md>), [research](<https://devfeed.tech/tags/research.md>)

## AI overview

OpenAI introduces a framework and 13-evaluation suite spanning 24 environments to measure chain-of-thought monitorability. The study finds that monitoring internal reasoning is generally more effective than monitoring actions and final outputs alone, and that monitorability often improves when models reason for longer.

## Source excerpt

OpenAI introduces a new framework and evaluation suite for chain-of-thought monitorability, covering 13 evaluations across 24 environments. Our findings show that monitoring a model's internal reasoning is far more effective than monitoring outputs alone, offering a promising path toward scalable control as AI systems grow more capable.