DeepSeek正式推出V4-Flash版本,这一轻量级模型在性能上超越了此前的V4-Pro-Preview,并在多项基准测试中与Claude Opus等顶级模型相当[1]。该模型总参数为284B,但激活参数仅13B,展现出显著的高效优势[1]。在定价方面,DeepSeek V4-Flash的成本极为低廉,仅为2块钱每百万token,缓存命中时还可享受一折优惠,整体性价比远超同类产品[1]。
此次更新的核心突破来自两个方面[1]。其一是高质量后训练的应用,通过采用GRPO强化学习等技术手段大幅提升了模型能力,预览版与正式版之间的最大差异便源于此[1]。这种方法的效果已在R1的开发中得到验证——相同模型在GRPO强化学习的作用下,数学能力从15.6%飙升至71%[1]。其二是推出了DeepSeek Harness工具,该工具为语言模型配备了代码运行环境、搜索引擎等外置工具,属于Agent类解决方案[1]。多伦多大学的研究表明,同一模型在不同Agent工具中的表现差异可达数倍,因此工具的引入有望进一步释放模型潜力[1]。
DeepSeek has launched the official version of V4-Flash, a model that outperforms its V4-Pro-Preview predecessor and matches top-tier models like Claude Opus across multiple benchmarks [1]. The lightweight variant features 284 billion total parameters with only 13 billion activated parameters [1].
The release marks two significant technical advances. First, enhanced post-training methodologies—including GRPO reinforcement learning techniques—have substantially boosted model capabilities, mirroring improvements seen in previous iterations where such training elevated mathematical performance from 15.6% to 71% [1]. Second, DeepSeek has introduced the Harness Agent tool, an external framework that equips the model with auxiliary capabilities such as code execution environments and search engine integration [1].
Pricing remains highly competitive at 2 yuan per million tokens, with cache hit optimization offering up to a 10% discount [1]. Research from the University of Toronto indicates that the same model can exhibit performance variations of several times across different agent tools, highlighting the significance of such external infrastructure improvements [1].