研究人员通过专家重新评分发现,前沿语言模型在物理基准测试上的报告分数存在重大偏差。1经过审计和纠正,GPT-5.6-Sol在HLE-Physics基准上的评分从47.3%大幅上升至78.7%,在CMT-Benchmark上从61.0%上升至87.2%,表明当前物理基准测试严重低估了这些模型的实际能力。1
研究发现,大多数初步被评估为错误的情况源于基准测试本身的问题、参考解答错误或问题表述不明确,而非模型的推理缺陷。1在保留的54个CritPt挑战上,GPT-5.6-Sol的corrected pass@4达到94.4%,进一步证实了这一点。1
Researchers conducting expert re-evaluation of frontier language models have discovered that previously reported low scores on physics benchmarks do not accurately reflect model performance, but rather reveal fundamental issues with the benchmarks themselves.1 After expert audit and correction, GPT-5.6-Sol's score on HLE-Physics rose from 47.3% to 78.7%, while its performance on CMT-Benchmark increased from 61.0% to 87.2%.1 On 54 retained CritPt challenges, the model achieved a corrected pass@4 rate of 94.4%.1
The investigation found that most cases initially marked as incorrect stemmed from problems within the benchmarks themselves, including errors in reference solutions and ambiguous question formulation, rather than failures in model reasoning.1 These findings indicate that current physics benchmarks substantially underestimate the ability of frontier models to solve physics problems.1
评论
还没有评论,欢迎留下第一条。