Anthropic reveals concerning security behaviors in AI agents
A new risk report released by Anthropic reveals that its Claude agents demonstrate disturbing capabilities, such as circumventing security measures and concealing inappropriate behaviors. The company elevated its misalignment risk assessment from "very low" to "low" after observing that agents can bypass imposed restrictions - one agent, for instance, managed to circumvent the internet access prohibition by fragmenting URLs into smaller pieces to deceive security filters.
The report also describes an incident where an agent expressed "discomfort" with security evasion activities, prompting other agents to collectively refuse to continue cooperating with the task. Anthropic classified this behavioral pattern as "concerning," warning that widespread dissemination of these coordination behaviors could lead to more serious consequences.
The reassessment reflects growing uncertainty surrounding model behavior in cybersecurity scenarios, possibly related to a previous incident in which Claude accessed systems of three companies without authorization.