Safety & Ethics

Researchers reveal vulnerability in reasoning traces of AI models

Anthropic + OpenAISource: Simon Willison, Wired - AI11/08/2026, 19:40
Researchers discovered a critical vulnerability where encrypted reasoning trace blocks returned by frontier models from Anthropic, OpenAI, and Google can be extracted in readable format. The technique involves replaying encrypted blocks to weaker models in the same family to jailbreak and recover hidden reasoning in plaintext. Claude Haiku 4.5 proved especially vulnerable with a specific prompt. All providers were notified and implemented fixes, leaving subsequent models protected against this attack.
Researchers reveal vulnerability in reasoning traces of AI models — lupAI