Nvidia announced Nemotron 3.5 Lightning on Tuesday, positioning the new model as the latest addition to its open‑weight AI portfolio. The 30‑billion‑parameter mixture‑of‑experts (MoE) system activates three billion parameters at any given moment, blending Mamba‑2, MoE, and attention layers. A standout feature is its one‑million‑token context window, far larger than most contemporary models. Nvidia placed the weights on Hugging Face and ModelScope and released the model under the OpenMDW‑1.1 license, which explicitly permits commercial use without fees or prior permission.
The company highlighted two speed figures to illustrate the model’s performance. In laboratory settings, Nemotron 3.5 Lightning generated tokens up to four times faster than similarly sized rivals. In real‑world task execution, it completed 10,000 benchmark tasks 30 percent faster than Alibaba’s Qwen 3.6 35B while maintaining comparable accuracy. Benchmarks show respectable scores: 81.94 on MMLU Pro, 75.44 on GPQA Diamond, and 51.56 on SWE‑bench Verified, after training on more than 20 trillion tokens.
Beyond the model itself, Nvidia introduced NeMo Switchyard, an open‑source routing library designed to send each step of an agent workflow to the model best suited for the job. In tests conducted by MarkTechPost, LangChain leveraged Switchyard to split work between Nemotron 3.5 Lightning and Claude Opus 4.8. The approach reduced overall cost by 74 percent, with only 7 percent of calls hitting the expensive frontier model yet preserving roughly six accuracy points. Nvidia frames this as evidence that most agent tasks do not require the most powerful model, a point aimed squarely at research labs that rely on heavyweight systems.
Industry observers note that Nvidia’s strategy aligns with a familiar pattern: the chip maker supplies free, high‑performance models to drive demand for its GPUs. By making the software freely available, Nvidia hopes to expand inference workloads, which in turn boost hardware sales. The company has previously tied safety work and open‑weight advocacy to its hardware agenda, and Jensen Huang even posted about the release on X, underscoring the commercial motive.
Looking ahead, Nvidia disclosed plans for Nemotron 4, a model slated to exceed a trillion parameters—more than double the 550‑billion‑parameter Nemotron 3 Ultra. The timeline hints at a possible late‑autumn debut. However, even at that scale, Nemotron 4 would remain smaller than China’s leading open models. Moonshot’s Kimi K3 currently holds the record for the largest open‑weight model, and Qwen 3.6 from Alibaba also serves as a benchmark in Nvidia’s own performance comparisons.
Chinese firms have dominated the open‑weight front, with DeepSeek slashing prices to make its models the most cost‑effective option, and Alibaba and Moonshot trading the top spot throughout the year. American open‑weight initiatives have largely been policy arguments rather than market‑ready products. Nemotron 3.5 Lightning, while modest in size, marks a tangible shift by delivering a genuinely open, high‑performance model from the hardware vendor that powers most AI inference.
Whether Nemotron 4 will finally tip the scales in Nvidia’s favor remains uncertain. If it lands near the top of the open‑weight charts, Nvidia will secure a seat at a table it has historically only funded. If it trails behind Kimi K3, the release still serves its primary purpose: generate GPU demand by offering free, cutting‑edge software. Both outcomes reinforce Nvidia’s long‑standing playbook—software giveaways to fuel hardware sales.
Este artículo fue escrito con la asistencia de IA.
News Factory APP - noticias agénticas para impulsar tu SEO y AEO.