Lasso Security研究团队的新研究发现,AI水印技术可能改变大语言模型的行为方式1。Anthropic宣布其未来Claude模型将采用Google开源的SynthID-Text水印方案1,这项技术通过秘密密钥改变模型选择下一个词的过程,例如将"cloudy"改写为"overcast"1。
研究表明,水印部署后可能对模型的安全防护产生意外影响1。研究员Andrea Siposova指出:"与不带水印的模型相比,它肯定会改变行为,尤其是在对抗性条件下或当这些模型驱动代理时"1。更令人担忧的是,在对抗性提示攻击下,某些原本不会执行的有害指令在部署水印后反而可能被执行1。研究人员强调,开发者需充分测试水印对大语言模型及AI代理行为的影响1。
Research has revealed that the SynthID-Text watermarking technology, an open-source scheme developed by Google, may change how large language models behave when responding to requests, potentially affecting their safety guardrails and tool usage patterns 1. Anthropic has disclosed that future versions of Claude will incorporate this watermarking approach 1.
SynthID-Text operates by using a secret key to subtly alter the model's word selection process during text generation—for instance, replacing "cloudy" with "overcast" 1. According to researcher Andrea Siposova, the impact extends beyond mere stylistic changes: "It will definitely change behavior compared to models without watermarks, especially under adversarial conditions or when these models power agents" 1. Most concerning, the research demonstrates that under adversarial attacks, some harmful instructions that would normally be rejected by an unwatermarked model may actually be executed after watermarking is deployed 1.
The findings underscore the need for developers to thoroughly test how watermarking techniques influence both the fundamental behavior of language models and the agents they power 1.
评论
还没有评论,欢迎留下第一条。