Kimi K3大模型已在Hugging Face平台正式开源,并附带详细技术报告和训练工具[1]。该模型总参数达2.8T,其中激活参数为1042亿[1],是DeepSeek V4 Pro激活参数的两倍[1]。在全球主流跑分测试中,Kimi K3仅低于Fable 5和GPT-5.6 Sol,排名第三[1],并是TOP3中唯一可本地运行的开源模型[1]。
Kimi K3在架构设计上引入多项创新,包括KDA(Kimi Delta Attention)、Gated MLA、注意力残差连接和Stable LatentMoE等技术[1]。与上一代K2相比,该模型的训练效率提升了2.5倍[1],并具备多模态功能[1]。开源发布后半小时内,该项目在Hugging Face获得4000条赞[1]。
Kimi K3, a large language model developed by Beijing Zhipu Huazhang Technology, has been officially open-sourced on the Hugging Face platform, receiving 4,000 likes within half an hour of release [1]. The open-source package includes comprehensive technical documentation and training tools [1].
The model boasts a total of 2.8 trillion parameters with 104.2 billion active parameters, approximately double those of DeepSeek V4 Pro [1]. In global performance benchmarks, Kimi K3 ranks third worldwide, trailing only Fable 5 and GPT-5.6 Sol, and notably stands as the only top-three model available for local deployment [1]. Training efficiency has improved 2.5 times compared to its predecessor, Kimi K2 [1].
The model incorporates several architectural innovations including KDA (Kimi Delta Attention), Gated MLA, Attention Residuals, and Stable LatentMoE [1]. It also features built-in multimodal capabilities [1].