NVIDIA released Nemotron 3.5 Lightning for long-running agent execution.
• 30B total / 3B active parameters, with up to 1M tokens of context. Ready for commercial use.
• Local deployment: NVIDIA ships BF16 and much smaller NVFP4 checkpoints, lists 1× DGX Spark or 1× H100 for single-GPU deployment, and also lists RTX 5090 among supported hardware.
• NVIDIA then adds multi-token prediction, DSpark and DFlash speculative decoding, plus NVFP4 quantization to push generation speed further.
• NVIDIA claims up to 4× output speed; on PinchBench, 86% accuracy and 30% faster completion than Qwen3.6 35B at similar accuracy.