深度求索(DeepSeek)于7月31日正式发布DeepSeek-V4-Flash-0731模型的API公测版本[1][2]。该模型拥有304亿参数[1],在Artificial Analysis Intelligence Index评测中获得50分[2],与GPT-5.6 Luna性能相当,但单任务成本低约60%[2]。
在定价方面,DeepSeek-V4-Flash-0731的输入价格为每百万tokens 0.14美元,输出价格为每百万tokens 0.27美元[1],成为当前性价比最优的模型之一[1]。与参数量更大的428亿参数模型MiniMax M3相比,该新模型在Intelligence Index中排名更靠前[1]。此外,DeepSeek为其自有API提供约98%的缓存命中折扣,远超业界90%的标准折扣水平[2]。在性能提升方面,V4-Flash-0731相比4月发布的V4-Flash版本提升10分,比V4 Pro版本高出6分[2]。
DeepSeek released its latest model, DeepSeek-V4-Flash-0731, on July 31st, 2026, featuring 304 billion parameters and enhanced agent capabilities [1]. The model is priced at $0.14 per million tokens for input and $0.27 per million tokens for output, positioning it as one of the most cost-effective options currently available [1][2].
The official API entered public testing on July 31st, with independent benchmark institutions Artificial Analysis and Arena.ai simultaneously releasing updated test results [2]. According to Artificial Analysis, the model achieves an Intelligence Index score of 50, matching the performance of GPT-5.6 Luna while delivering approximately 60% lower per-task costs [2]. The V4-Flash-0731 variant improves 10 points over the April 2026 version of DeepSeek-V4-Flash and 6 points above DeepSeek-V4-Pro [2]. In Frontend Code Arena rankings, DeepSeek-V4-Flash-High secured 1,586 points, placing seventh overall and third among open-category competitors [2].
DeepSeek's per-task cost remains lower than GPT-5.6 Luna even if OpenAI reduces its pricing by 80% [2]. Additionally, DeepSeek offers approximately 98% cache hit discounts on its proprietary API, substantially exceeding the industry standard of 90% [2]. Artificial Analysis ranked the model ahead of MiniMax M3, a 428-billion-parameter competitor [1].