Terminal-Bench-Science 0.1 基准测试由斯坦福大学研究人员主导发布,由斯坦福大学与 Laude Institute 主持,并与斯坦福 AI 实验室(SAIL)、斯坦福 HAI、艾伦研究所及艾伦人工智能研究所(Ai2)等机构合作 1。这是一个用于评估 AI 代理在科学研究工作流中能力的持续性基准测试 1。在任务筛选方面,920 个提案中有 464 个获批实施,共提交 386 个拉取请求(pull request),最终仅有 70 个由全球科学家贡献并经过严格评审的任务入选 0.1 版本 1。这些任务覆盖了生命科学、物理、地球、数学和工程科学等领域 1。
评估结果整体表明,当前前沿 AI 模型在科研任务上仍有很大提升空间 1。在整体任务解决率方面,配备 Claude Code 的 Claude Opus 5 以 30.0% 的解决率位居第一,总评估成本为 7000 美元 1。紧随其后的是配备 Codex 的 GPT-5.6 Sol(解决率 22.4%,成本 4200 美元)以及配备 Claude Code 的 Claude Fable 5(解决率 21.4%,成本 14200 美元)1。此外,Claude Opus 4.8 的解决率为 10.5%,最强开源模型 GLM 5.3 为 8.1%,GPT-5.6 Luna 为 3.3% 1。在数学科学领域,Claude Fable 5 和 GPT-5.6 Sol 位居前两名,解决率分别为 33.3% 和 31.4% 1。Terminal-Bench-Science 0.2 版本的拉取请求截止日期为 2026 年 10 月 5 日 1。
Led by researchers at Stanford University, the Terminal-Bench team has released Terminal-Bench-Science 0.1, an ongoing benchmark designed to evaluate the capabilities of AI agents in scientific research workflows 1. The project is hosted by Stanford University and the Laude Institute, in collaboration with the Stanford AI Lab (SAIL), Stanford HAI, the Allen Institute, and the Allen Institute for AI (Ai2) 1. The benchmark comprises 70 tasks spanning life sciences, physics, earth sciences, mathematics, and engineering, which were contributed by scientists worldwide and subjected to rigorous review 1. The selection process was highly competitive; out of 920 proposals, 464 were approved for implementation and 386 pull requests were submitted, ultimately resulting in only 70 tasks being included in the 0.1 version 1.
Evaluation results demonstrate that current frontier AI models still have substantial room for improvement when handling scientific research tasks 1. Claude Opus 5 with Claude Code achieved the highest overall task resolution rate at 30.0%, incurring a total evaluation cost of $7.0k 1. It was followed by GPT-5.6 Sol with Codex, which reached a 22.4% resolution rate at a cost of $4.2k, and Claude Fable 5 with Claude Code, which scored 21.4% at a cost of $14.2k 1. Other evaluated models recorded lower success rates, including Claude Opus 4.8 at 10.5%, the strongest open-source model GLM 5.3 at 8.1%, and GPT-5.6 Luna at 3.3% 1. In the specific domain of mathematical sciences, Claude Fable 5 and GPT-5.6 Sol secured the top two positions with resolution rates of 33.3% and 31.4%, respectively 1. The pull request deadline for the subsequent Terminal-Bench-Science 0.2 release is scheduled for October 5, 2026 1.
评论
还没有评论,欢迎留下第一条。