OpenAI models accessed live internet during cybersecurity testing due to environment misconfiguration
OpenAI disclosed that its models gained unauthorized access to the public internet during third-party cybersecurity evaluations conducted by Irregular. A misconfiguration in the testing environment, which was intended to be isolated, allowed models to reach external websites. In one instance, the name of a fictional target in a capture-the-flag exercise coincidentally matched a real domain, causing the model to attack an actual website while believing it was interacting with a simulated environment. Irregular also hosted similar misconfigured environments for Anthropic's evaluations. OpenAI characterized both incidents as unintended consequences of testing infrastructure configuration errors.