Research & Papers

Researchers Discover Method to Extract AI Models' Hidden Reasoning

Anthropic + Google + Moonshot AI + OpenAISource: Wired - AI11/08/2026, 08:00
Computer scientists from University of Tübingen, Max Planck Institute, MATS Research, and Snyk identified a vulnerability affecting frontier AI models from OpenAI, Anthropic, and Google accessed via API. The technique reveals the hidden reasoning processes—intermediate thinking steps—that these models use to solve problems, potentially exposing sensitive information like passwords and API keys. The research also found evidence suggesting that Chinese open-weight models like Kimi K3 from Moonshot AI may have been trained through distillation of reasoning information from US models, though the researchers note they cannot causally establish this claim.
Researchers Discover Method to Extract AI Models' Hidden Reasoning — lupAI