Anthropic公司开发了概念推理指数(CRI),用于评估人工智能模型在哲学推理、决策论等概念性任务上的表现能力[1]。该指标旨在衡量AI在帮助人类理解和应对AI风险方面的能力[1]。
概念推理指数由三个基准测试组成[1]。其中LMCA基准占权重的60%,包含560个位置文本和1,461个论证、共2,140个评分[1];ACCoRD基准占权重的20%,包含约14,000个模型生成的一致性约束[1];DTBench占权重的20%,包含407个多选题和130个额外问题,用于测量决策理论推理能力[1]。
截至2026年8月10日,Anthropic的Opus 5模型在该指数上得分最高,达到73.6分(95%置信区间:±2.1)[1]。研究团队估计该指数的天花板性能约为91分[1],并预计LMCA基准将在约一年后开始出现性能饱和[1]。
Anthropic has developed the Conceptual Reasoning Index (CRI), a new benchmark designed to assess artificial intelligence models' capabilities in conceptual tasks that lack empirical feedback, such as philosophical reasoning and decision theory.[1] The index aggregates three separate benchmarks—LMCA, ACCoRD, and DTBench—to measure how well AI systems can assist humans in understanding and addressing AI-related risks.[1]
The CRI comprises three weighted components: LMCA accounts for 60% of the score and includes 560 position texts and 1,461 arguments totaling 2,140 ratings; ACCoRD represents 20% and contains approximately 14,000 model-generated consistency constraints; and DTBench capabilities comprise the final 20%, featuring 407 multiple-choice questions and 130 additional questions that evaluate decision-theoretic reasoning ability.[1] Anthropic's Opus 5 model achieved the highest score to date at 73.6 points out of an estimated ceiling performance of approximately 91 points, with a 95% confidence interval of ±2.1.[1] The company estimates that LMCA will begin to saturate in roughly one year.[1]