NVIDIA has just released the first Nemotron 3.5 model: Nemotron 3.5 Lightning, a highly efficient small open weights model with performance similar to gpt-oss-120b at around a quarter of the total parameters
Nemotron 3.5 Lightning is the successor to @nvidia Nemotron 3 Nano 30B A3B, with 31.6B total and 3.6B active parameters. It retains the same hybrid Mamba-Transformer architecture and small size from Nemotron 3 Nano, but makes substantial gains in intelligence and agentic performance.
Key takeaways:
➤ Major intelligence jump: Nemotron 3.5 Lightning scores 24 on the Artificial Analysis Intelligence Index, a +9 point improvement over Nemotron 3 Nano (15). This puts it in line with OpenAI's gpt-oss-120b (24) and only just behind Nemotron 3 Super (26), a model ~4x its size
➤ Optimized for efficiency: Nemotron 3.5 Lightning sits behind the most intelligent small models in its size class such as Qwen3.6 35B A3B (32) and Muse Glimmer (high, 35) - but it is built for a different point on the frontier. In pre-release testing of a DeepInfra endpoint serving the final NVFP4 weights, we measured median output speeds of nearly 670 tokens per second, much faster than those models are served in the market today
➤ Meaningful agentic gains: the largest improvements over Nemotron 3 Nano come on agentic evaluations in GDPval-AA v2 (+334 ELO, moving past gpt-oss-120b and Nemotron 3 Super) and Terminal-Bench v2.1 (24% vs 7%). Combined with its speed and permissive OpenMDW-1.1 license, this positions Lightning as an efficient workhorse model for high-volume agentic deployments
➤ Near-lossless NVFP4 quantization: as with prior Nemotron releases, the model ships in NVFP4 alongside BF16 weights. We measured the NVFP4 variant at 24 on the Intelligence Index and saw minimal degradation compared to the higher-precision weights
Key model details:
➤ 1 million token context window, text-only reasoning model
➤ 31.6B total and 3.6B active parameters
➤ Released under the OpenMDW-1.1 license, open for commercial use without material restrictions
➤ The model weights are available now along with serverless inference from providers including @DeepInfra, @FireworksAI_HQ, @friendliai, @CoreWeave, @gmi_cloud, @nebiusai, and @CrusoeAI