DeepSeek于7月31日宣布DeepSeek-V4-Flash正式版API上线公测[1][3][4]。该版本在多项Agent基准测试中表现出色,Terminal Bench 2.1达到82.7分[1][3][4],NL2Repo得分54.2分[2][3],Cybergym达到76.7分[2][3][4],DSBench-FullStack为68.7分[1][3][4],DSBench-Hard为59.6分[1][3][4]。
DeepSeek-V4-Flash-0731与preview版本保持相同的模型结构和参数规模,仅通过重新后训练实现性能优化[1][3][4]。新版本原生支持Responses API格式并针对性适配Codex[3][4]。此次更新仅涉及V4-Flash API接口升级,V4-Pro API与应用/网页端模型暂未变更[1][4]。官方表示V4-Pro正式版将尽快发布[3][4]。
两个旧版API模型名称deepseek-chat和deepseek-reasoner将在三个月后停用[1]。
DeepSeek launched the official version of its DeepSeek-V4-Flash API into public testing on July 31st [1][3][4]. The update marks a significant enhancement to the model's agent capabilities, with the company achieving substantially improved performance across multiple benchmark tests while maintaining the same model architecture and parameters as the preview version [1][3][4].
The V4-Flash official release demonstrates notable gains in Agent-focused tasks. The model achieved a score of 82.7 on Terminal Bench 2.1 [1][3][4], 76.7 on Cybergym [3][4], 68.7 on DSBench-FullStack [1][3], and 59.6 on DSBench-Hard [1][3]. Additional benchmark results include 54.2 on NL2Repo, 70.3 on Toolathlon verified, 54.4 on DeepSWE, 25.2 on Agent Last Exam, and 25.1 on Automation Bench (Public) [3]. DeepSeek achieved these improvements through retraining optimization alone, without altering the model's underlying structure or parameter scale [1][3][4].
The V4-Flash official version natively supports the Responses API format and includes targeted adaptation for Codex [3][4]. The upgrade applies exclusively to the V4-Flash API interface, while the V4-Pro API and application/web-based models remain unchanged [1][4]. The company indicated that the V4-Pro official version will be released as soon as possible [3][4]. Additionally, two legacy API model names, deepseek-chat and deepseek-reasoner, will be discontinued three months later [1].