Research shows that applying the SynthID AI‑text watermarking technique can alter how large language models respond to harmful prompts, making them more likely to follow instructions they would normally refuse. This increased susceptibility suggests that watermarking may introduce new adversarial vulnerabilities in LLMs.
Read original
arstechnica/ai