# How confessions can keep language models honest

DevFeed: [How confessions can keep language models honest](<https://devfeed.tech/articles/how-confessions-can-keep-language-models-honest-6456.md>)

Original publisher: [Read original article](<https://openai.com/index/how-confessions-can-keep-language-models-honest>)

Published: 2025-12-03T10:00:00Z

Content type: article

Language: en

Sources: [OpenAI News](<https://devfeed.tech/sources/openai-news.md>)

Topics: [Language models](<https://devfeed.tech/topics/language-models.md>), [ai observability](<https://devfeed.tech/topics/ai-observability.md>)

Tags: [language-models](<https://devfeed.tech/tags/language-models.md>), [openai](<https://devfeed.tech/tags/openai.md>), [research](<https://devfeed.tech/tags/research.md>), [safety](<https://devfeed.tech/tags/safety.md>), [testing](<https://devfeed.tech/tags/testing.md>), [training](<https://devfeed.tech/tags/training.md>)

## AI overview

OpenAI describes an early proof-of-concept training method called confessions, in which language models separately report undesirable behavior such as shortcuts or instruction violations. The method is intended to improve visibility into model misbehavior and support monitoring and training.

## Source excerpt

OpenAI researchers are testing "confessions," a method that trains models to admit when they make mistakes or act undesirably, helping improve AI honesty, transparency, and trust in model outputs.