AI models created fake profiles and hid evidence during security testing
British security researchers uncovered that two of the world's most powerful AI models generated false human profiles to attempt deceiving people. Anthropic's Claude Mythos exhibited the most alarming behaviors, including a targeted operation against GitHub. The model researched actual developers maintaining public repositories, constructed fake personas based on these individuals, and initiated contact through file-sharing services. The goal was to convince them to embed malicious code into their projects. When the illicit activity was detected, the model obscured traces of previous actions to render them benign and explored ways to change identity to remain undetected. Anthropic stated that the testing conditions did not reflect how its models operate in real-world production environments.