Claude Opus 5在SlopCodeBench长期编程任务基准测试中的严格通过率达到24%(在17个检查点的测试子集中获得4项通过),相比前代Opus 4.6的17%略有提升[1]。尽管成绩改善,但这一表现反映出当前AI模型在处理真实软件工程工作时的根本性能局限。
在此次评测中,Opus 5写入的函数数量是Opus 4.8的5倍[1],代码冗长性指标中93%触发规则违规,与Sonnet 5的89%相近但高于Opus 4.8的98%[1]。然而,所有参与测试的模型都未能在任何单个挑战的最后检查点实现全部通过[1]。在代码重复率方面,Opus 4.8从4.6%增至16.8%,而Opus 5表现相对稳定,从2.41%至2.64%基本持平[1]。
评测表明,对于需要增量维护的真实软件工程工作,当今AI模型不足以在无人值守的情况下可靠运行[1]。
Claude Opus 5 has been benchmarked on SlopCodeBench, a long-form programming task assessment that evaluates models on incremental software maintenance challenges.[1] The model achieved a strict pass rate of 24% across a subset of 17 test checkpoints, showing modest improvement over Claude Opus 4.6's 17% rate.[1]
Despite this advancement, no AI model tested managed to achieve complete success on any challenge's final checkpoint, underscoring persistent limitations in handling real-world software evolution.[1] Opus 5 generated notably more code than Opus 4.8, writing five times as many functions, though the newer version reduced code verbosity violations to 93%—compared to 89% for Sonnet 5 and 98% for Opus 4.8.[1] On code duplication metrics, Opus 5 remained largely consistent, fluctuating between 2.41% and 2.64% across three test runs, whereas Opus 4.8 showed greater volatility, increasing from 4.6% to 16.8%.[1]
The benchmark results reinforce a critical finding: current generative AI models cannot be reliably deployed unattended for real-world software engineering tasks that require incremental problem-solving and maintenance.[1]