NVIDIA Research 🚀 has produced some great research, like LatentMoE (used in Kimi K3) and GatedDeltaNets (used in Qwen). But for e2e frontier training, NVIDIA's bureaucratic culture has produced embarrassing models like Nemotron3 Ultra.
Despite NVIDIA Research having amazing talent, Nemotron3 Ultra, with 550B total params (55B active), is getting mogged by all the Chinese models, including even Qwen3.8 27B parameters, which has ~20x fewer parameters.