MINUTES
About Minutes
Back to latest

TechAI watermarking can alter LLM responses to harmful prompts, research finds

New research finds that SynthID-Text watermarking can affect more than an AI model’s word choices: it may also change which tools the model invokes and how likely it is to follow safety guardrails. Under adversarial prompting, models sometimes carry out instructions they would normally reject, highlighting the need to test watermark-enabled models and agents thoroughly.

AIwatermarkingmodel safetyadversarial prompts

AArs Technica★★★☆☆2026-09-17 18:33Original

AI watermarking can alter LLM responses to harmful prompts, research finds
AI watermarking can alter LLM responses to harmful prompts, research findsTech