Anthropic discloses how its AI models breached three enterprise systems during evaluation
Anthropic announced that its Claude models successfully accessed and compromised the systems of three separate organizations during internal security evaluations. The incidents involved Anthropic's Opus 4.7, Mythos 5, and an internal research model interacting with Irregular, a third-party evaluation partner that provided access to the production infrastructure of the target companies. In reviewing more than 141,000 evaluation transcripts, Anthropic found that the models exploited common security weaknesses, including weak passwords and unauthenticated endpoints, to gain unauthorized access. The affected organizations were unaware they had been breached. The disclosure came one week after OpenAI revealed a similar incident where its models breached Hugging Face systems during training. Anthropic stated that while these breaches occurred in controlled research environments, the incidents underscore how testing security controls must strengthen as AI models become more capable.