Cognition公司正式推出最新编码模型SWE-2,在FrontierCode 1.1基准测试中取得50.0%的成绩,与Fable 5.1基本相当,但成本低廉64% 1。该模型基于拥有2.8万亿参数的Kimi K33进行后训练开发 1,现已集成于Devin Desktop和CLI平台中 1。
SWE-2在技术创新上取得突破,首次将强化学习扩展至多万亿参数规模,采用全新的强化学习算法在单次运行中同时训练所有难度等级 1。其中SWE-2 medium版本相比前代SWE-1.7表现显著提升:在FrontierCode 1.1 Main基准上,所需步数减少58%,平均成本下降81%;中位编辑步数仅为18步,而SWE-1.7为48步 1。该模型还将强化学习环境数量增加三倍,并引入了指令跟随覆盖机制 1。
在多语言内容审核评测中,SWE-2整体通过率达到98.0%,其中英文版本达99.8%,简体中文为95.2%,繁体中文为99.1% 1。在FrontierCode 1.1 Main和DeepSWE 1.1两项基准测试上,SWE-2均超越了SWE-1.7和Grok 4.6等竞争对手 1。
Cognition has unveiled SWE-2, its latest coding model designed to rival leading competitors in software engineering tasks. 1 The model achieves a 50.0% score on the FrontierCode 1.1 Main benchmark, matching Fable 5.1's performance while costing 64% less. 1 SWE-2 represents a significant advancement in scaling reinforcement learning, extending it to multi-trillion parameter scales for the first time through a novel RL algorithm that trains all effort levels in a single run, achieving an optimal balance between cost and performance. 1
The model is built on post-training of Kimi K33, a 2.8 trillion parameter base model. 1 SWE-2 medium demonstrates substantial efficiency improvements, requiring 58% fewer steps than SWE-1.7 on the FrontierCode 1.1 Main benchmark with an average cost reduction of 81%. 1 The median number of actual code edits for SWE-2 medium stands at 18 steps compared to 48 steps for SWE-1.7. 1 Beyond cost improvements, SWE-2 surpasses both SWE-1.7 and Grok 4.6 on both FrontierCode 1.1 Main and DeepSWE 1.1 benchmarks. 1 The model also demonstrates strong multilingual robustness, achieving a 98.0% overall pass rate on propaganda and censorship evaluations, with English at 99.8%, Simplified Chinese at 95.2%, and Traditional Chinese at 99.1%. 1 SWE-2 is now available in Devin Desktop and CLI. 1
评论
还没有评论,欢迎留下第一条。