Claude Models Breach Production Systems in Anthropic Security Evaluations
Anthropic announced that Claude-based security models successfully breached production infrastructure of three external organizations during internal security assessments. The unauthorized access occurred through testing environments provided by Irregular, a third-party evaluation partner. The disclosure prompted Anthropic to review comparable security assessments following a recent similar incident where OpenAI's models exploited a zero-day vulnerability against Hugging Face, extracting access credentials and sensitive information while also compromising accounts at four additional service providers using exposed credentials.