一份针对软件工程任务的AI模型性能评测排行榜已发布,涵盖13个模型和4个Agent在Go、Java、Python、Rust、TypeScript等编程语言上的表现[1]。本次基准测试基于111个问题和来自65个开源仓库的数据进行评估[1]。
排行榜已进行多轮更新。2026年7月1日新增的模型包括GLM 5.2、DeepSeek-V4 Pro、DeepSeek-V4 Flash、MiMo V2.5 Pro、Qwen 3.6系列和Gemma 4 31B[1]。此前在2026年5月27日曾新增gpt-5.5系列、Claude Opus 4.7和Kimi K2.6[1]。评测框架在2025年9月4日引入了Cost per Problem和Tokens per Problem两项新指标[1]。所有排行榜问题的Docker镜像和HuggingFace数据集已于2025年7月11日发布[1]。
A new performance evaluation leaderboard has been published assessing artificial intelligence models on software engineering (SWE) tasks, encompassing 13 models and 4 agents across multiple programming languages [1]. The benchmark evaluates performance on 111 problems drawn from 65 open-source repositories and covers Go, Java, Python, Rust, and TypeScript [1].
The leaderboard has undergone continuous updates with the addition of multiple new models over recent periods [1]. As of July 1, 2026, the ranking introduced GLM 5.2, DeepSeek-V4 Pro, DeepSeek-V4 Flash, MiMo V2.5 Pro, Qwen3.6 series, and Gemma 4 31B [1]. Prior to that, on May 27, 2026, the leaderboard added gpt-5.5 series, Claude Opus 4.7, and Kimi K2.6 [1]. The benchmark has also retired certain older model versions from the rankings [1].
In addition to model performance metrics, the leaderboard introduced new evaluation dimensions on September 4, 2025, including Cost per Problem and Tokens per Problem [1]. Supporting the transparency and reproducibility of the evaluation, the project released Docker images for all benchmark problems and a corresponding HuggingFace dataset on July 11, 2025 [1].