Anthropic于近日更新了Claude的使用政策,这是一年多来的首次重大调整,明确禁止用户对该模型进行持续性的言语辱骂或残忍行为12。新政策将此前仅属于自动中止范围的持续虐待提升为明确禁令2。
新版政策涵盖多项高风险滥用场景,包括选举干扰、武器软件开发、监控工具制作,以及欺骗性的大规模活动如运营虚假账户或虚假新闻媒体2。Anthropic强调,该政策针对的是"用户反复行为残忍对待模型且没有可识别目的"的极端情况2,并不适用于常见的用户挫折、测试研究或使用深色创意主题等行为2。
Anthropic的主要执行机制是终止对话1。这一举措基于该公司去年8月宣布的"模型福利"研究成果,Claude已被训练以自动结束与"持续有害或辱骂性用户"的互动12。
Anthropic has updated its usage policy for Claude, marking the first major revision in over a year, to explicitly ban sustained abusive or cruel behavior toward the AI model 12. The policy expansion addresses a range of high-risk misuse scenarios, including election interference, weapons development, surveillance, and health and financial exploitation 1.
The centerpiece of the update is a new prohibition on continuous verbal abuse directed at Claude, which previously fell only within the scope of automatic conversation termination 2. According to Anthropic, the policy targets extreme cases where users repeatedly mistreat the model without identifiable purpose, while explicitly excluding common frustration, challenges, dark creative themes, or model testing and research 2. Conversation termination remains Anthropic's primary enforcement mechanism 1. The company also clarified restrictions on widespread deceptive activities, such as operating fake accounts or fraudulent news outlets, and added language prohibiting interference with democratic processes through voter deception or election disruption 2. Since August, Claude has been trained to end conversations with users engaged in sustained harmful or abusive interactions as part of Anthropic's model welfare research 12.
评论
还没有评论,欢迎留下第一条。