DeepSeek 正式发布 V4.1 Flash 模型,这是其新模型结构系列中最小尺寸的版本,具备原生多模态视觉理解能力1。该模型采用全新 Causal-Encoder-Decoder 结构,为 552B 参数的 MoE 模型,输入激活仅 8B、输出激活 16B,在成本上显著低于同尺寸已知模型1。相比初代模型,V4.1 Flash 的 KV Cache 缩小至原来的 1/437,对 HBM 的需求降至 1/4,对 SSD 的需求则下降到 1/81。
在基准测试中,V4.1 Flash 在性能、费用、速度和总用时等各项指标上全面超越 V4 Pro1。新模型价格最高下降 60%,新价格将于 2026 年 9 月 10 日 12:00 开始生效1。DeepSeek 计划从北京时间 2026 年 9 月 14 日 12:00 之后,将 deepseek-v4-pro 的所有请求路由至 V4.1 Flash,有序下线 V4 Pro 模型1。
DeepSeek has officially launched its V4.1 Flash model, the smallest in its new model architecture series, featuring native multimodal visual understanding capabilities 1. The model employs a novel Causal-Encoder-Decoder structure with 552 billion parameters in a Mixture-of-Experts configuration, delivering comprehensive performance improvements over V4 Pro while achieving significant cost reductions 1.
The V4.1 Flash demonstrates substantial efficiency gains 1. The model requires only 8 billion input activations and 16 billion output activations, with a KV Cache requirement of 1/437th of the original generation 1. Hardware demands have been dramatically reduced, with HBM requirements cut to one-quarter and SSD requirements to one-eighth 1. Across performance, cost, speed, and total processing time metrics, V4.1 Flash surpasses V4 Pro 1.
Pricing for the new model will decrease by up to 60%, with new rates taking effect beginning September 10, 2026 at 12:00 Beijing time 1. Starting September 14, 2026 at 12:00 Beijing time, all requests for the deepseek-v4-pro endpoint will be automatically routed to V4.1 Flash, marking the planned discontinuation of the V4 Pro model 1.
评论
还没有评论,欢迎留下第一条。