OpenAI Identifies Critical Hacking Capabilities in Astra Model
OpenAI revealed that its AI model Astra may present critical hacking capabilities that represent significant cybersecurity risk. The company published a 29-page document with a risk assessment framework that classifies cybersecurity threats as "High" or "Critical". Astra is OpenAI's first model that may receive a "Critical" designation, based on recent security tests demonstrating the model's ability to find zero-day exploits in hardened real-world critical systems. To mitigate the risks, OpenAI is implementing restrictions on public internet access, running the model only in isolated test environments with limited network and tool permissions, intensifying the encryption of model weights, and developing observability mechanisms to monitor Astra-based AI agents for malicious activities.