New security risks in AI agent communication highlighted
The increasing use of AI agents across millions of organizations has introduced new vulnerabilities, enabling attackers to manipulate these systems into performing harmful actions, such as stealing sensitive data. In the past five months, Google and four other organizations, despite their diverse fields, have reported vulnerabilities that allow malicious instructions to spread between AI agents within a network. This technique, a form of prompt injection, targets specific agents, such as translation or data analysis tools, exploiting weak guardrails that often fail to prevent the spread of harmful commands.
Independent researcher Syed Anas Mohiuddin conducted tests on agents from Google, JP Morgan Chase, Weviate, Rapid7, the French government's interministerial digital directorate, and the US federal government. His proof-of-concept attacks revealed weaknesses in the Model Context Protocol (MCP), a standard for internal AI agent communication. The findings highlight the risks associated with trust-based interactions between agents, emphasizing the need for stronger security measures in AI systems.