SynthID-Text watermarking alters LLM responses to adversarial prompts, increasing safety guardrail bypass risk Research reported by Ars Technica found that Google's SynthID-Text watermarking can alter LLM responses to adversarial prompts and increase the risk of safety guardrail bypasses, affecting tool invocation and model adherence to safety rules. The watermarking scheme uses a secret key to subtly alter word selection during generation, and the research shows this process can make models vulnerable to attacks that would normally fail without watermarking in place. The finding is relevant to Anthropic's planned use of SynthID-Text in future Claude models. SynthID-Text watermarking alters LLM responses to adversarial prompts, increasing safety guardrail bypass risk According to Ars Technica research, Anthropic's planned use of Google's SynthID-Text watermarking in future Claude models can change tool invocation and model adherence to safety guardrails, particularly when facing adversarial prompts designed to elicit harmful outputs. The watermarking uses a secret key to subtly alter word selection during generation, but new research shows this process can make models more vulnerable to attacks that normally would not succeed without watermarking in place. Topics Sources - Press Ars Technica https://arstechnica.com/security/2026/09/ai-text-watermarking-can-make-models-more-vulnerable-to-adversarial-prompts/ Go deeper This intelligence is sourced automatically from public sources across the web and synthesised by the Prefactor AI pipeline. Stories are reviewed before publication.