Anthropic推出了Claude Opus 5大模型1。新模型相比前代Opus 4.8实现了显著的性能提升,同时维持原有价格——输入令牌$5/百万,输出令牌$25/百万1。
Claude Opus 5在多个权威评估基准上实现了业界领先成绩1。在Frontier-Bench v0.1评估中,其性能相比Opus 4.8提升一倍以上1。在ARC-AGI 3评估上,该模型的得分达到次优模型的三倍1。在计算机使用能力方面,Claude Opus 5在OSWorld 2.0基准上以仅Fable 5成本三分之一的价格超越了后者的最佳成绩1。在Zapier AutomationBench上的通过率约为次优模型的1.5倍1。
在代码编程和科学研究领域表现尤为突出1。软件工程任务中的性能翻倍提升1,有机化学任务上较Opus 4.8提高10.2个百分点,蛋白质相关任务上提高7.7个百分点1。在CursorBench 3.2最大努力设置下,该模型性能接近Fable 5峰值成绩,但成本仅为一半1。
新模型已上线所有平台,成为Claude Max的默认模型和Claude Pro的最强选择1。用户可选择Fast模式,其运行速度约为默认速度的2.5倍,但价格为基础价格的两倍1。此外,Claude Opus 5在自动行为审计中被评为目前最对齐的模型,欺骗行为率最低1。
在某些专业领域,该模型仍存在差距1。生物学研究和网络安全方面的性能仍落后于Mythos 51。
Anthropic has released Claude Opus 5, a new large language model that delivers substantial performance improvements over its predecessor Opus 4.8 while maintaining identical pricing 1. The model is now available across all platforms as the default option for Claude Max and the most advanced model for Claude Pro 1.
The new model is priced at $5 per million input tokens and $25 per million output tokens 1. Despite the performance enhancements, Anthropic has kept costs unchanged from the previous generation 1.
Claude Opus 5 demonstrates leading performance across multiple benchmarks 1. On Frontier-Bench v0.1, it surpasses all competing models with more than double the performance of Opus 4.8 1. The model achieves results within 0.5 percent of Fable 5's peak score on CursorBench 3.2's maximum effort setting, while costing half as much 1. On the ARC-AGI 3 evaluation, Claude Opus 5 scores three times higher than the second-best model 1, and on Zapier AutomationBench, its pass rate is approximately 1.5 times that of the next-best performer 1. For the OSWorld 2.0 computer use benchmark, it exceeds Fable 5's best results at one-third of the cost 1.
In specialized domains, Claude Opus 5 shows notable improvements over Opus 4.8: it scores 10.2 percentage points higher on organic chemistry tasks and 7.7 percentage points higher on protein-related tasks 1. A faster operating mode runs approximately 2.5 times faster than the default speed at double the base price 1.
The model also demonstrates strong alignment properties, achieving the lowest deception rate in automated behavioral audits among all models tested 1. However, it remains behind Mythos 5 in biology research and cybersecurity applications 1.
评论
还没有评论,欢迎留下第一条。