在大模型发展的竞争中,不同厂商对原生多模态能力的策略选择出现明显分化。Kimi K3采用参数2.8万亿的大模型方案,整合原生多模态、代码开发和Agent能力,支持100万token的上下文窗口,每个token激活约1040亿参数[1]。该模型在约15万亿混合图文token的联合预训练中按固定比例融合文本与视觉数据[1],曾以1679分登顶Arena Frontend Code榜单[1]。
相比之下,DeepSeek V4-Flash采取了截然不同的技术路径[1]。这款模型的总参数仅为2840亿,每个token激活130亿参数,主要通过后训练优化代码能力,暂未加入视觉模态[1]。DeepSeek创始人梁文锋曾表示:"把AI训练做好,并不需要世界模型,甚至不用多模态"[1],体现了该团队对多模态的不同认识。
与此同时,OpenAI1月企业使用报告显示图片上传在ChatGPT最常用工具中排名第三,反映出用户对多模态能力的实际需求[1]。阿里的Qwen3.8-Max总参数为2.4万亿,每个token激活950亿[1],代表了业界在大模型规模设计上的另一种思路。
Kimi K3 and DeepSeek V4 have adopted fundamentally different technical approaches to integrating multimodal abilities into large language models.[1] Kimi K3, developed by Moon Matters, features 2.8 trillion parameters with approximately 104 billion activated per token, supports a context window of 1 million tokens, and incorporates native multimodal capabilities alongside coding and agent functions.[1] In contrast, DeepSeek V4-Flash operates with 284 billion total parameters and 13 billion activated per token, taking a lighter approach by focusing optimization efforts on coding ability through post-training rather than adding visual modality.[1]
The strategic difference reflects divergent philosophies on model design.[1] DeepSeek founder Liang Wenfeng has stated that "doing AI training well does not require a world model, and doesn't even require multimodality," suggesting skepticism about the necessity of visual capabilities for general-purpose performance.[1] Meanwhile, Kimi K3 has invested substantially in native multimodality, conducting joint pre-training on approximately 15 trillion mixed image-text tokens with fixed-ratio blending of textual and visual data.[1] This capability proved consequential in practice: Kimi K3 achieved first place on the Arena Frontend Code leaderboard with a score of 1,679.[1]
The market response indicates growing user demand for multimodal features.[1] According to OpenAI's January enterprise usage report, image uploads ranked third among the most-used tools in ChatGPT, underscoring the practical value enterprises place on visual input handling.[1] Alibaba's competing Qwen3.8-Max model, operating with 2.4 trillion parameters and 95 billion activated per token, represents a middle-ground approach in the multimodal integration spectrum.[1]