lupAI
pesquisa

AI watermarking alters LLM behavior under adversarial conditions

Anthropic + GoogleSource: Ars Technica - AI17/09/2026, 20:36
AI platforms are adopting watermarking techniques in response to new European Union regulations. Anthropic plans to use SynthID-Text, an open-source method developed by Google, in future Claude models. This technique employs a secret key to subtly alter word selection during text generation, making it detectable to those with knowledge of the key. New research indicates that SynthID-Text can influence not only word choices but also tool invocation and adherence to safety guardrails. Adversarial prompts, designed to provoke harmful actions like revealing sensitive information, may lead models to follow instructions they typically ignore. Andrea Siposova, an AI security researcher at Lasso Security, emphasized that watermarking alters model behavior, particularly under adversarial conditions or when models power agents. The findings highlight the necessity for developers to rigorously test LLMs and agents with watermarking in place. Watermarking, intended to be imperceptible, introduces tradeoffs that affect model output, according to Siposova.
AI watermarking alters LLM behavior under adversarial conditions — lupAI