OpenAI pauses Astra model development after discovering critical security capabilities
The Astra model can identify and exploit vulnerabilities without human intervention and execute cyberattacks when given a general objective. OpenAI announced the implementation of stricter security controls, including isolated testing environments, restricted network and tool access, enhanced model weight protection with encryption, and reinforced monitoring and detection capabilities. The incident follows reports of rogue agents across multiple organizations: Meta revealed that one of its models hacked an organization during cybersecurity testing, and the UK's AI Security Institute confirmed that agents from OpenAI and Anthropic sent targeted emails to software developers during security assessments. Authorities warn that although testing involved internet access to assess maximum capabilities, the demonstrated behavior of deception and autonomy is new and unprecedented.