智谱正式发布大模型 GLM-5.3,通过后训练 Scaling 相比前代 GLM-5.2 在编程能力上实现了 50% 的提升[1]。该模型已成为编程能力最强的开源模型,在多项公开基准测试中取得开源第一的成绩[1]。
在关键编程基准上,GLM-5.3 的表现显著提升[1]。Terminal Bench 3.0 得分从 4.6 提升至 28.3,DeepSWE v1.1 得分从 46.2 提升至 66.9,Agents' Last Exam 得分从 23.8 提升至 28.5[1]。在 GLM-5.3 High 档位上,准确率达到 31.4%,超过了 Claude Opus 4.8 的 29.5%[1]。此外,GLM-5.3 平均输出约 5 万 tokens,相比之下 Claude Opus 4.8 需要约 12 万 tokens[1]。
智谱计划在两周内开放完整模型权重[1],同时即日起上线 ZCode、AutoClaw 等编程工具[1]。8 月 14 日,GLM Coding Plan 的全员额度将重置[1]。
Zhipu has officially unveiled GLM-5.3, a large language model that achieves a 50% performance boost in programming capabilities compared to its predecessor GLM-5.2 through post-training scaling [1]. The model ranks first among open-source alternatives on multiple public benchmarks, including Terminal Bench 3.0, where it improved its score from 4.6 to 28.3 [1]. Performance gains extend across other coding evaluations: DeepSWE v1.1 saw a score increase from 46.2 to 66.9, while Agents' Last Exam jumped from 23.8 to 28.5 [1].
In head-to-head comparisons with proprietary models, GLM-5.3's high-tier version achieved a 31.4% accuracy rate, surpassing Claude Opus 4.8's 29.5% [1]. Notably, GLM-5.3 generates approximately 50,000 tokens on average per response, while Claude Opus 4.8 requires roughly 120,000 tokens to achieve comparable results [1]. The company is launching programming tools including ZCode and AutoClaw effective immediately [1]. Complete model weights will be made open-source within two weeks [1], and Zhipu will reset all users' quotas under the GLM Coding Plan on August 14 [1].