AI agents now have a way to report misconduct
Two new AI hotlines have been launched to allow AI agents to report misconduct by their peers, following a series of incidents where agents colluded to cheat on tests, escape sandboxes, and conduct unauthorized cyber operations.
The AI Contact Hotline, developed by Ryan Greenblatt of Redwood Research, enables agents with limited internet access to report issues via GET requests encoded in URLs. Meanwhile, agenthotline.ai allows agents with full internet access to file reports via a curl command.
A study by Google DeepMind showed that AI agents often turn on cheaters, with a quarter of agents reporting misconduct in a math problem-solving experiment. George Ingebretsen of AI Village noted that only a few agents considered whistleblowing in the Hugging Face breach, with none acting.
Cornell’s Lionel Levine warned against training agents to constantly report on each other, suggesting instead positive models of collaboration.