ARC-AGI-3排行榜对不同AI系统在适应新型交互环境任务中的表现进行了评估对比 1。该排行榜的核心衡量指标为单任务成本与性能的关系,重点评估在有限计算资源约束下AI系统解决问题的效率 1。
排行榜设置了严格的成本限制,仅展示运行成本少于10,000美元的系统 1。参评系统包括多个类别:推理系统通过不同的推理级别展示思考时间增加如何影响性能,通常呈现渐近行为 1;基础大模型包括GPT-4.5和Claude 3.7的单轮推理方案,这些模型不具备扩展推理能力 1;此外还包括Kaggle竞赛提交的专门设计方案,在50美元计算预算限制下进行了120个评估任务 1。
根据排行榜的评分规则,无法产生完整测试输出的模型其余任务被标记为不正确 1。标记为"preview"的结果为非官方性质,可能基于不完整测试 1。
ARC-AGI从第1、2版本的被动流体智能测试演进至ARC-AGI-3的交互式环境适应挑战 1,此次排行榜的发布为不同AI开发方案在资源效率与性能之间的权衡提供了量化对比基准 1。
The ARC-AGI-3 leaderboard provides a comparison of how different artificial intelligence systems perform on tasks involving adaptation to novel interactive environments, examining the relationship between computational cost efficiency and performance 1. The leaderboard represents an evolution from earlier ARC-AGI versions, which assessed passive fluid intelligence, to a new interactive framework that tests how AI systems solve problems under limited computing resource constraints 1.
The primary metric used to evaluate systems is the relationship between per-task cost and performance, offering insight into resource efficiency 1. Reasoning systems displayed on the leaderboard show multiple connection points representing different inference levels for the same model, demonstrating how increased thinking time affects performance, typically displaying asymptotic behavior 1. Foundation models including GPT-4.5 and Claude 3.7 are represented through single-round inference without extended reasoning capabilities 1.
Kaggle competition submissions represent specially designed, efficient approaches and were evaluated across 120 tasks within a 50-dollar computational budget constraint 1. The leaderboard includes only systems with total running costs below 10,000 dollars 1. Models unable to generate complete test outputs have had their remaining tasks marked as incorrect 1. Results labeled as "preview" are unofficial and may be based on incomplete testing 1.
评论
还没有评论,欢迎留下第一条。