Nvidia推出了Nemotron 3.5 Lightning开源AI模型,该模型拥有30亿参数,采用混合专家架构,专为长期运行的智能代理工作负载优化[1]。相比同类模型,Nemotron 3.5 Lightning的输出速度快4倍,代理任务完成速度提升30%[1]。
同时发布的NeMo Switchyard开源库具备智能请求路由功能,能够将任务动态分配到最合适的模型,帮助企业在开源、专有和Nvidia模型之间实现高效部署[1]。内部基准测试表明,使用NeMo Switchyard可将成本降低至Opus的约三分之一[1]。
多家企业已验证该方案的实际效果。Boomi使用Switchyard实现了100%的域路由准确率,将59%的流量转向性能快5倍的微调模型,后续延迟下降21%[1]。LangChain在145个多轮Deep Agents任务中通过Switchyard实现74%的成本降低,其中仅7%的请求被路由到前沿模型,精度权衡仅6%[1]。Ramp则利用该工具将成本降低58%,运行时间缩短33%[1]。
Nvidia has unveiled Nemotron 3.5 Lightning, an open-source artificial intelligence model featuring 3 billion parameters designed as a mixture-of-experts architecture tailored for long-running intelligent agent workloads [1]. The model delivers output speed 4 times faster than comparable alternatives and completes agent tasks 30 percent more quickly [1].
Alongside the model release, Nvidia introduced NeMo Switchyard, an open-source library that intelligently routes requests to the most suitable model for each task [1]. This tool enables enterprises to efficiently deploy AI applications by dynamically selecting among open-source, proprietary, and Nvidia-hosted models [1]. Internal benchmarking shows NeMo Switchyard reduces costs to approximately one-third of those associated with Opus [1].
Early adopters have demonstrated substantial benefits from the routing system. Boomi achieved 100 percent accuracy in domain-based routing and redirected 59 percent of traffic to faster fine-tuned models while reducing subsequent latency by 21 percent [1]. LangChain reported a 74 percent cost reduction across 145 multi-turn deep agent tasks, routing only 7 percent of requests to frontier models while accepting a 6 percent precision tradeoff [1]. Ramp lowered costs by 58 percent and reduced runtime by 33 percent through Switchyard deployment [1].