Guidelight AI Standards对Anthropic、Google、OpenAI、Meta和xAI五家领先AI实验室进行了评估,检查它们在模型失控情况下的应对计划1。评估结果显示,大多数顶级实验室尚未发布或演示相关的遏制方案1。OpenAI在评估中得分最高(5分中得3分),而Anthropic和Meta的得分最低1。
该评估基于公开可得的安全计划,涵盖日志监控、系统暂停、独立审计和失控模型遏制方案等多项指标1。Guidelight首席科学家、前OpenAI安全研究员Steven Adler表示:"我对AI公司在模型失控时如何处理非常严重的事件说得如此之少感到惊讶"1。OpenAI在Hugging Face事件后提高了透明度,披露了如何隔离行为不当的模型的具体细节1。
评估工作受到监管要求的推动,加州SB 53和纽约RAISE Act等法案要求AI企业披露安全事件应对框架1。此外,美国众议院代表提出了跨党派的《AI Kill Switch Act》法案,要求主要AI开发者建立关闭失控模型的技术机制1。
Guidelight AI Standards conducted an assessment of five leading artificial intelligence laboratories and found that most top AI companies have yet to publicly disclose or demonstrate plans for containing a malfunctioning model 1. The evaluation examined Anthropic, Google, OpenAI, Meta, and xAI, rating them on metrics including logging and monitoring, system pause capabilities, independent audits, and rogue model containment protocols 1. OpenAI scored highest among the group with 3 out of 5 points, while Anthropic and Meta received the lowest scores 1.
The research, based on publicly available safety documentation, aims to advance transparency and establish safety standards for frontier AI development 1. Steven Adler, chief scientist at Guidelight and former OpenAI safety researcher, expressed concern about the lack of detail companies provide regarding their response to critical incidents involving model failures, stating: "I was surprised by how little AI companies have said about how they would handle such serious events when a model goes rogue" 1. OpenAI has increased its disclosure efforts following the Hugging Face incident, sharing specifics about how it isolates misbehaving models 1.
The assessment reflects growing regulatory pressure on AI companies to disclose safety incident response frameworks, as seen in California's SB 53 and New York's RAISE Act 1. In Congress, bipartisan legislation titled the AI Kill Switch Act has been introduced, requiring major AI developers to establish technical mechanisms capable of shutting down rogue models 1.
评论
还没有评论,欢迎留下第一条。