OpenAI halts coordinated effort to steal model reasoning
OpenAI disclosed it disrupted a coordinated campaign aimed at extracting protected reasoning from its models, with activity traced to July 1, 2026. The campaign, consistent with adversarial distillation, involved manipulating model interactions to reproduce internal reasoning, violating terms of service. Operators exploited model interactions rather than breaching encryption or accessing user data. A surge in requests occurred on July 24 and 25, with over 16,000 requests from more than 4,000 users. OpenAI attributed the activity to individuals linked to Moonshot AI, the developer of Kimi. The company emphasized that adversarial distillation poses safety and national security risks, enabling advanced capabilities transfer without same safety measures. OpenAI has implemented technical controls, banned fraudulent accounts, and shared findings through the Frontier Model Forum to bolster industry defenses.
The company highlighted novel attack methods, including attempts to copy encrypted reasoning from one conversation and have another model decrypt it. Independent security researchers had previously identified cross-model vulnerabilities, which OpenAI confirmed as real. OpenAI is enhancing protections against such attacks, focusing on technical safeguards, detection mechanisms, and threat information sharing across industry and government. The company stressed that adversarial distillation is an evolving threat requiring layered defenses and ongoing adaptation.