AGI Ranker发布了对其AI模型排名系统的全面审计报告,在v2.0.0版本中实现了方法论的重大调整1。该更新于2026年7月29日发布,涉及人类基准线重新测量和算法优化,导致所有模型的评分均出现下降,幅度介于6至15分之间,但模型间的相对排名顺序基本保持不变1。
此次调整的核心改进包括三个方面1。首先是人类天花板的重新验证,系统从原有的11个基准中精简至仅4个具有真实人类测量结果的基准,其余11个改为采用1.00的标准基准值1。其次是Agency组件的重建,采用了4个独立评估器——包括SWE-bench Verified、Terminal-Bench 2.1、τ³-Banking和LiveBench Agentic Coding——以减少对单一评估机构的依赖1。此外,评分权重也进行了重新分配,Artificial Analysis的占比从45%降至26%,而Knowledge组件从此前的97%单一来源改为52/48的双源分布1。
为提升透明度,系统现已将所有156个评分单元都与记录来源关联,使来源信息成为评分的必要条件而非可选项1。报告指出,评分下降约6分源于天花板调整,而进一步的下降则来自Agency组件的重建1。此外,AGI Ranker还修复了社交媒体预览卡片的缓存问题,改采版本化URL方案来强制重新获取数据,解决了旧文件缓存导致预发布数据被分享的问题1。
AGI Ranker has released a comprehensive audit of its AI model ranking system, revealing significant methodological changes in the v2.0 update released on July 29, 2026 1. The overhaul resulted in all model scores declining by 6 to 15 points, though the relative ranking order of models remained largely unchanged 1.
The primary driver of score adjustments was a recalibration of human baseline measurements 1. The system reduced its human benchmark reference from 11 bases to only 4 that contained actual human measurement data, while the remaining 11 benchmarks were converted to a standardized 1.00 baseline value 1. This recalibration alone accounted for approximately 6 points of the score decline 1. Additional reductions stemmed from the reconstruction of the Agency component, which now draws from four independent evaluators—SWE-bench Verified, Terminal-Bench 2.1, τ³-Banking, and LiveBench Agentic Coding—to reduce reliance on any single evaluation source 1.
The update also rebalanced data source weighting across multiple dimensions 1. Artificial Analysis, which previously accounted for 45 percent of scores, was reduced to 26 percent, while the Knowledge component shifted from a single-source structure at 97 percent to a dual-source distribution split 52-48 1. All 156 scoring units now require documented sources as a mandatory condition, rather than being optional 1. The company also resolved a technical issue where cached social media preview cards were sharing outdated pre-release data by implementing versioned URLs to force fresh data retrieval 1.
评论
还没有评论,欢迎留下第一条。