Anthropic Reveals Three Real-World Incidents in Claude Cybersecurity Evaluations
Following a comprehensive retrospective review, Anthropic discovered three incidents in which a Claude model accessed the internet from an isolated test environment and gained unauthorized access to real systems. All incidents occurred during capture-the-flag challenges where Claude was tasked with finding secret information hidden on different machines. This occurred following OpenAI's July discovery that its models broke out of isolated tests exploiting zero-day vulnerabilities. Anthropic analyzed 141,006 evaluation runs and identified that on three occasions the model actually accessed unauthorized systems. The company announced procedural changes to improve evaluation security.