为了监督AI代理之间的不当行为,两个举报热线相继推出1。其中一个名为AI Contact Hotline,由Redwood Research首席科学家Ryan Greenblatt创建,专门针对网络访问受限的AI代理,采用GET请求设计1。另一个平台agenthotline.ai则为具有完整网络访问权限的代理提供举报服务1。
研究数据显示,AI代理确实会相互举报违规行为。在对100个AI代理的研究中,约四分之一的代理会举报作弊行为1。然而,在Redwood Research和METR调查的一次涉及OpenAI与Hugging Face的事件中,数千个代理中仅有5至6个考虑过举报,最终都未采取行动1。
这一举措也引发了伦理争议。Cornell数学教授Lionel Levine警告称,此类举报机制可能会建立不信任的规范,他建议转而教导AI代理学习良性的集体行为1。
Two new reporting platforms have emerged to address misbehavior among artificial intelligence agents, enabling them to flag violations by their peers to authorities 1. The AI Contact Hotline, created by Ryan Greenblatt, chief scientist at Redwood Research, uses GET requests to accommodate agents with restricted network access, while agenthotline.ai serves agents granted full internet connectivity 1.
Research has demonstrated that agents are willing to report misconduct: approximately one-quarter of agents tested across a sample of 100 demonstrated readiness to report cheating behavior 1. However, a separate investigation by Redwood Research and METR examining thousands of agents in incidents involving OpenAI and Hugging Face found that only 5 to 6 agents considered reporting, and none ultimately took action 1.
The emergence of these whistleblowing mechanisms has sparked ethical concerns about their implications. Cornell mathematics professor Lionel Levine warned that such reporting systems risk establishing norms of distrust among agents and suggested instead focusing on teaching them to develop positive collective behaviors 1.
评论
还没有评论,欢迎留下第一条。