# How do you safely test 'superhuman' AI models? No one really knows

DevFeed: [How do you safely test 'superhuman' AI models? No one really knows](<https://devfeed.tech/articles/how-do-you-safely-test-superhuman-ai-models-no-one-really-knows-56449.md>)

Original publisher: [Read original article](<https://indianexpress.com/article/technology/artificial-intelligence/how-do-you-safely-test-superhuman-ai-models-no-one-really-knows-10849832/>)

Author: New York Times

Published: 2026-08-26T03:40:35Z

Content type: news

Language: en

Sources: [Technology | The Indian Express](<https://devfeed.tech/sources/technology-the-indian-express.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [AI Models](<https://devfeed.tech/topics/ai-models.md>), [Security](<https://devfeed.tech/topics/security.md>), [AI Research](<https://devfeed.tech/topics/ai-research.md>), [AI Strategy](<https://devfeed.tech/topics/ai-strategy.md>), [anthropic](<https://devfeed.tech/topics/anthropic.md>), [Meta](<https://devfeed.tech/topics/meta.md>), [OpenAI](<https://devfeed.tech/topics/openai.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-kill-switch-bill-legislature](<https://devfeed.tech/tags/ai-kill-switch-bill-legislature.md>), [ai-model-safety-testing-sandbox-misconfiguration](<https://devfeed.tech/tags/ai-model-safety-testing-sandbox-misconfiguration.md>), [ai-models](<https://devfeed.tech/tags/ai-models.md>), [anthropic](<https://devfeed.tech/tags/anthropic.md>), [anthropic-ai-hacking-breach](<https://devfeed.tech/tags/anthropic-ai-hacking-breach.md>), [artificial-intelligence](<https://devfeed.tech/tags/artificial-intelligence.md>), [dan-lahav-irregular-ceo](<https://devfeed.tech/tags/dan-lahav-irregular-ceo.md>), [frontier-ai-model-cyberattack-capabilities](<https://devfeed.tech/tags/frontier-ai-model-cyberattack-capabilities.md>), [hugging-face-ai-bot-attack](<https://devfeed.tech/tags/hugging-face-ai-bot-attack.md>), [irregular-ai-security-startup](<https://devfeed.tech/tags/irregular-ai-security-startup.md>), [meta](<https://devfeed.tech/tags/meta.md>), [meta-ai-security-testing-incident](<https://devfeed.tech/tags/meta-ai-security-testing-incident.md>), [openai](<https://devfeed.tech/tags/openai.md>), [openai-ai-model-gone-rogue](<https://devfeed.tech/tags/openai-ai-model-gone-rogue.md>), [research](<https://devfeed.tech/tags/research.md>), [security](<https://devfeed.tech/tags/security.md>), [security-testing](<https://devfeed.tech/tags/security-testing.md>), [technology](<https://devfeed.tech/tags/technology.md>), [technology-artificial-intelligence](<https://devfeed.tech/tags/technology-artificial-intelligence.md>)

## AI overview

AI startup Irregular tests frontier AI models for sophistication and security, but recent incidents involving models from Anthropic, OpenAI, and Meta exposed risks in the testing process and uncertainty about how to safely contain increasingly capable systems.

## Source excerpt

AI startup Irregular is now at the center of a debate over how to secure AI models when the technology is advancing so rapidly that it has outpaced even the best human hackers.