英伟达推出了开放权重的大语言模型Nemotron 3.5 Lightning,该模型采用混合Mixture-of-Experts架构,融合了Mamba-2和MoE层以及Multi-Token Prediction功能[1]。模型包含30B总参数,其中3B为活跃参数,支持1M令牌的超长上下文长度[1]。
该模型已于2026年8月11日在Hugging Face平台发布,正式名称为NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4,采用OpenMDW-1.1许可证[1]。模型预训练数据超过20万亿令牌,数据截止日期为2025年9月,后训练数据截止至2026年5月[1]。Nemotron 3.5 Lightning支持英文、西班牙文、法文、德文、意大利文、日文等多种语言,以及43种编程语言[1],可应用于代理系统、聊天机器人、RAG系统等场景[1]。
Nvidia has released Nemotron 3.5 Lightning, an open-weight large language model featuring a hybrid Mixture-of-Experts architecture that combines Mamba-2 and MoE layers [1]. The model, officially designated NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4, contains 30 billion total parameters with 3 billion active parameters and supports a context length of up to 1 million tokens [1]. The model became available on Hugging Face on August 11, 2026 [1].
The model was trained on over 200 trillion tokens with a data cutoff date of September 2025, and its post-training data extends to May 2026 [1]. Nemotron 3.5 Lightning is designed as a general-purpose reasoning and chat model, with primary optimization for English and programming languages, though it also supports Spanish, French, German, Italian, and Japanese, alongside 43 programming languages [1]. The model is released under the OpenMDW-1.1 license and is intended for applications including agent systems, chatbots, and retrieval-augmented generation (RAG) systems, with commercial use permitted [1].