像 Nemotron-3-Super 这样的大型混合 MoE 模型虽然准确,但部署成本高昂。其活跃参数、KV 缓存和 Mamba 状态限制了单个节点在给定每用户 token 速率下所能承载的用户数量。NVIDIA AI 团队发布了 Nemotron-Labs-3-Puzzle-75B-A9B,这是 Nemotron-3-Super 的一个压缩版本。原始模型总参数量为 1207 亿,活跃参数量为 128 亿。压缩后的模型总参数量为 753 亿,活跃参数量为 93 亿。
部署目标在架构搜索开始前就已确定。目标一:在每用户每秒 100 个 token 的条件下,服务器吞吐量提升 2 倍。目标二:在单个 H100 上实现 8 个并发 100 万 token 的请求。Hugging Face 上提供了三个检查点:BF16、FP8 和 NVFP4。
1207 亿/128 亿活跃参数压缩至 753 亿/93 亿活跃参数,同时保留了原有的 88 块混合布局。
在匹配的 NVFP4 和匹配的用户吞吐量下,8 块 B200 的总吞吐量相比 Super 提升了 1.60 倍至 2.14 倍。
单个 H100 上 100 万 token 的并发数从 1 提升至 8,这得益于权重从 70 GB 降至 44.5 GB。
在相同的压缩目标下,迭代式 Puzzle 方法比单步式 Puzzle 方法平均高出 0.57 分。
Arena-Hard-V2(-4.2)和 SWE-Bench(-2.6)是实际付出的代价;RULER 和 AA-LCR 则几乎没有变化。
Nemotron-Labs-3-Puzzle-75B-A9B
Nemotron-3-Super 是一个混合了 Mamba-Transformer 的 MoE 模型。Puzzle-75B-A9B 完全保留了原始模型的块布局。它共有 88 个块:40 个 Mamba 块、40 个 MoE 块和 8 个注意力块。
发生变化的是这些块内部的容量:
数量 | Super | Puzzle-75B-A9B | 比例 总参数 | 1207 亿 | 753 亿 | 62.4% 活跃参数 | 128 亿 | 93 亿 | 73.1% Mamba SSM 状态大小 | 128 | 96 | 75% MoE 路由专家中间大小 | 2688 | 1280-2688 | 平均 59.9% 每 token 激活的路由专家数 | 2 | 2-4 | 平均 18% 活跃路由专家容量(相对值) | 100% | 8.7%-62.3% | 平均 30.9%
路由专家数量、共享专家大小以及 MoE 潜在空间大小保持不变。注意力层未作改动。该研究提出的理由是,Nemotron-3-Super 本身在 KV 缓存方面已经非常高效。Mamba 层被统一剪枝,因为推理框架不支持每层使用不同的 SSM 状态大小。
其结果并非一个均匀缩小的教师模型。上图展示了各层之间的容量分配情况。Puzzle 在选定的中间层和后层保留了容量,而在其他层则大幅削减。
基准测试与性能表现
下表报告了在单个 8×B200 节点上,采用单步解码时的帕累托最优总吞吐量。
场景(输入/输出)UT 下限Super (tok/s)Puzzle-75B-A9B (tok/s)提升倍数50K / 2K>= 1005,1288,2101.60x50K / 2K>= 1253,7846,4121.69x50K / 2K>= 1502,5324,5231.79x8K / 64K>= 10020,93942,6012.03x8K / 64K>= 12513,07427,9182.14x8K / 64K>= 1508,52218,0472.12x
两个模型均在匹配的 NVFP4 权重、FP8 KV 缓存和 FP16 Mamba 状态下提供服务。因此,差距反映的是压缩效果,而非数值格式的变化。以预填充为主的 50K/2K 场景提升最小。以解码为主的 8K/64K 场景提升最大。
在单个 8×H100 节点上,UT = 100 时,提升幅度较小。在 50K/2K 场景下为 1.91 倍,在 8K/64K 场景下为 1.82 倍。这两个模型在此处均使用 FP8 权重、FP8 KV 缓存和 FP32 Mamba 状态。
在单个 H100 上处理 100 万上下文时,瓶颈约束从计算转向内存。Super 的 NVFP4 权重约占 80 GB HBM 预算中的 70 GB。每个 100 万 token 的请求会增加约 4 GB 的 KV 缓存。因此,有效并发数为 1。
Puzzle-75B-A9B 的 NVFP4 权重约占 44.5 GB。注意力布局不变,因此每个请求的 KV 缓存成本不变。在 100 万上下文时的并发数提升至 8。该并发下的聚合解码吞吐量大约是 Super 单请求吞吐量的 4 倍。对 99 万 token 提示词的预填充速度大约快 1.2 倍。
迭代式 Puzzle 的工作原理
Puzzle 是一个分解式神经架构搜索框架,在此以 Puzzletron 形式实现。它定义了一个由多种备选层实现方案组成的离散搜索空间。每种备选方案都会获得一个质量评分。然后,一个混合整数规划会在部署约束条件下,为每一层选择一种备选方案。
三种剪枝技术构成了该搜索空间:
中间通道剪枝:每个路由专家内部的通道根据其对专家输出的贡献进行排序。一个 MoE 层内的所有专家都被剪枝到统一大小,以保证内核兼容性。
Top-k 缩减:每个 token 被路由到的专家数量因层而异,最多不超过父模型的 k=22。
Mamba SSM 剪枝:SSM 状态大小从 128 通道降至 96 通道。
对 SSM 结果进行了测量。将 128 通道降至 96 通道,在解码阶段可使 SSM 内核加速 1.2 倍至 1.3 倍。这一效果在批次大小为 8 到 512 之间均成立。通道根据其对 Mamba 层输出的预估贡献进行排序。该预估基于 6700 万 token 的验证数据取平均值。附录 A 显示,在激进剪枝条件下,此方法优于随机通道选择。
原始公式假设替换质量的影响大致是可叠加的。每个候选模块在未经修改的父模型内部进行评分。这忽略了替换之间的高阶交互作用。
迭代式 Puzzle 将有限压缩与短程知识蒸馏恢复交替进行。它构建了一个序列 M0, M1, … MR,而不是直接跳转到目标模型。评分是针对当前压缩后的模型重新计算的,而非原始父模型。
共使用了三个阶段:
MoE 权重降至教师模型容量的 75%,Mamba SSM 状态降至 75%。使用 240 亿 token 进行恢复训练。
MoE 权重降至教师模型容量的 60%。使用 432 亿 token 进行恢复训练。
激活的路由专家预算降至 50%,并进行异构分配。使用 528 亿 token 进行恢复训练。
上表将本方法与针对同一目标的单步 Puzzle 基线进行了比较。三步流程在十个基准测试上的平均得分为 69.05,而基线为 68.48。在 MMLU-Pro、GPQA、HLE、AA-LCR、LiveCodeBench、SciCode 和 RULER-256K 上均有提升。IFBench-Instruction 下降了 0.2 分,IFBench-Prompt 下降了 0.5 分。
恢复:蒸馏、强化学习与冗长度控制
知识蒸馏使用了来自 Nemotron-3-Nano 的 30% 预训练数据和 70% SFT 数据。在 Puzzle 阶段,KD 使用了 32K 的序列长度。恢复阶段则在 128K 长度下进行训练,并扩展至 512K。预算上限为 1000 亿 token,全局批次大小为 1600 万 token,在 Megatron-LM 框架中执行。
RL 后训练采用了 Nemotron-3-Super RL 流水线的第二阶段,重点聚焦软件工程。阶段 2.1 进行了单步工具使用对比。阶段 2.2 转向端到端沙箱 RL,智能体在此阶段最多可运行 200 轮。两个阶段均使用 KL 惩罚系数为 0。团队对学习率进行了扫描,然后对得到的权重进行了平均。
上图 4 展示了每个阶段的贡献。短上下文知识蒸馏将大多数类别的性能恢复至 Nemotron-3-Super 的 97% 以上。长上下文知识蒸馏则专门提升了长输入和长生成基准测试的性能。研究团队表示,RL 在这些实验中的影响较小。
冗长性是一个不易察觉的细节。在最后一次 Puzzle 迭代之后,模型生成的 token 数量达到了 Super 模型的 132%。经过完整的恢复流水线后,这一比例降至 99%。
部署:量化与多 token 预测
生成了两种训练后量化方案:FP8 W8A8 面向 Hopper 架构,NVFP4 W4A4 面向 Blackwell 架构。
组件 | BF16 基线 | FP8 检查点 | NVFP4 检查点 稀疏和共享 MoE GEMM | BF16 | FP8 | NVFP4 Mamba GEMM | BF16 | FP8 | FP8 Mamba SSM 缓存 | FP32 | FP32 | FP16+SRK KV 缓存 | FP8 | FP8 | FP8 路由器 | FP32 | FP32 | FP32 注意力 QKV/输出、MoE 潜在投影、LM 头 | BF16 | BF16 | BF16
两种方案均在 256 个训练后 SFT 样本上进行了校准。NVFP4 使用了最大校准,而非 Super 模型所使用的 AutoQuantize 敏感性搜索。由此产生的检查点量化程度略高,但性能表现相似。
NVFP4 在 Hopper 架构上并非原生支持。它仍被用于 100 万上下文窗口的 H100 目标场景,因为该场景受限于 HBM 容量。
Puzzle-75B-A9B 继承了 Super 模型的共享 MTP 头。参数在 MTP 步骤之间共享,因此一个头在推理时递归应用。直接迁移 Super 模型训练好的头,得到了相似的接受长度。
研究团队随后发现了一个训练-推理不匹配的问题。教师强制 MTP 训练输入的是完整的移位隐藏状态序列。而自回归草稿生成输入的则是目标模型和 MTP 生成隐藏状态的混合。在更深的草稿位置,接受率会下降。
对迁移后的头部进行持续训练解决了这个问题。在 draft 长度为 7 的 SPEED-Bench 上,平均接受长度从 3.45 提升到了 4.34。这大约提升了 25% 到 30%,主要集中在较靠后的 draft 位置。与 Super 不同,NVFP4 检查点几乎没有退化:4.31 对比 4.34。
压缩在何处有益,在何处有害
基准测试 (BF16) | Super | Puzzle-75B-A9B | Delta MMLU-Pro | 83.8 | 82.4 | -1.4 AIME 25 (无工具) | 92.2 | 89.7 | -2.5 GPQA (无工具) | 80.5 | 78.6 | -1.9 LiveCodeBench | 82.1 | 81.1 | -1.0 SciCode (子任务) | 42.3 | 40.6 | -1.7 SWE-Bench (OpenHands) | 59.5 | 56.9 | -2.6 Arena-Hard-V2 | 72.8 | 68.6 | -4.2 AA-LCR | 56.8 | 56.9 | +0.1 RULER 1M | 93.9 | 92.2 | -1.7 MMLU-ProX | 79.5 | 77.5 | -2.0
该研究论文自身的总结是,指令遵循和智能体类评测的损失最大。Arena-Hard-V2 是最严重的情况,下降了 4.2 分。RULER 在 256K、512K 和 1M 长度下,分数波动大致在 1 到 2 分以内。
有三项 BF16 结果没有出现退化。AA-LCR 提升了 0.1 分,Scale AI Multi-Challenge 持平于 56.6 分,TauBench Telecom 提升了 0.4 分。
NVFP4 在压缩基础上带来的额外代价很小。在 RULER 1M 上,NVFP4 检查点得分为 93.2,高于 BF16 的 92.2。HLE 是 NVFP4 代价最明显的例子,分数从 16.5 降至 15.7。FP8 的结果见附录 E,与 BF16 的结果非常接近。SWE-Bench 未报告 FP8 检查点的结果。
单 GPU 上的超长上下文 RAG:一个处理 1M 上下文的文档分析服务,从 1 个并发请求提升到了 8 个。在该并发度下的聚合解码吞吐量大约提升了 4 倍。
交互式编码助手:在 8K/64K 范围内,当用户吞吐量 >= 100 tok/s 时,单个节点可提供 2.03 倍的 token 量。考虑到冗长程度调整后,每分钟完成的请求数量提升了 2.16 倍。
预填充密集型文档处理管线:在 50K/2K 范围内,仅获得 1.60 倍的提升。当提示词处理占据计算主导地位时,压缩的帮助较小。
智能体 SWE 循环:请对照你的任务组合,检查 SWE-Bench 上 2.6 分的差距。强化学习恢复针对的是这项能力,并且仅部分恢复了它。
在匹配的 NVFP4 和匹配的用户吞吐量下,总吞吐量相比 Super 提升了 1.60 倍到 2.14 倍
单张 H100 上的 1M token 并发数从 1 个请求提升到了 8 个
在 draft 长度为 7 的 SPEED-Bench 上,MTP 接受长度从 3.45 提升到了 4.34
长上下文准确率在 RULER 的 256K、512K 和 1M 长度下,波动保持在 1 到 2 分以内
生成冗余在 99% 的 Super 处终止,因此 token 增益在请求级别得以保留。
发布了三个检查点:BF16、FP8 和 NVFP4。
Arena-Hard-V2 下降 4.2 分,SWE-Bench 下降 2.6 分。
RL 恢复带来的影响很小,论文中已直接说明。
Mamba 剪枝是均匀的,因为框架无法逐层改变 SSM 状态大小。
潜在维度剪枝已被放弃:NVFP4 MoE 内核需要潜在维度为 512 的倍数。
在这篇 v2 预印本中,文本和表格在多个吞吐量倍数上存在不一致。
分离式预填充增益仅为 5% 到 7%,并且增加了服务复杂性。
Large hybrid MoE models like Nemotron-3-Super are accurate but expensive to serve. Their active parameters, KV cache, and Mamba state cap how many users a node can hold at a given per-user token rate. NVIDIA AI team has released Nemotron-Labs-3-Puzzle-75B-A9B, a compressed variant of Nemotron-3-Super. The parent model has 120.7B total and 12.8B active parameters. The compressed model has 75.3B total and 9.3B active parameters.
The deployment target was fixed before the architecture search began. Target one was 2x server throughput at 100 tokens per second per user. Target two was 8 concurrent 1M-token requests on a single H100. Three checkpoints on Hugging Face: BF16, FP8, and NVFP4.
120.7B/12.8B active compresses to 75.3B/9.3B active, with the 88-block hybrid layout preserved.
8xB200 total throughput rises 1.60x to 2.14x over Super at matched NVFP4 and matched user throughput.
Single-H100 1M-token concurrency goes 1 to 8, driven by a 70 GB to 44.5 GB weight drop.
Iterative Puzzle beats single-step Puzzle by 0.57 average points at the same compression target.
Arena-Hard-V2 (-4.2) and SWE-Bench (-2.6) are the real costs; RULER and AA-LCR barely move.
Nemotron-Labs-3-Puzzle-75B-A9B
Nemotron-3-Super is a hybrid Mamba-Transformer MoE model. Puzzle-75B-A9B preserves the parent’s block layout exactly. It has 88 blocks: 40 Mamba, 40 MoE, and 8 attention blocks.
What changed is capacity inside those blocks:
QuantitySuperPuzzle-75B-A9BRatioTotal parameters120.7B75.3B62.4%Active parameters12.8B9.3B73.1%Mamba SSM state size1289675%MoE routed expert intermediate size26881280-2688Mean 59.9%Activated routed experts per token224-18Mean 50%Active routed expert capacity (relative)100%8.7%-62.3%Mean 30.9%
The number of routed experts, the shared expert size, and the MoE latent size are unchanged. Attention layers were left untouched. The proposed research’s stated reason is that Nemotron-3-Super is already very KV-cache efficient. Mamba layers were pruned uniformly, because inference frameworks do not support a different SSM state size per layer.
The result is not a uniformly scaled-down teacher. The above figure shows the allocation across depth. Puzzle preserved capacity in selected middle and late layers, and cut hard elsewhere.
Benchmark and Performance
The below table reports Pareto-optimal total throughput on a single 8xB200 node, with single-step decoding.
Scenario (in/out)UT floorSuper (tok/s)Puzzle-75B-A9B (tok/s)Boost50K / 2K>= 1005,1288,2101.60x50K / 2K>= 1253,7846,4121.69x50K / 2K>= 1502,5324,5231.79x8K / 64K>= 10020,93942,6012.03x8K / 64K>= 12513,07427,9182.14x8K / 64K>= 1508,52218,0472.12x
Both models were served at matched NVFP4 weights, FP8 KV cache, and FP16 Mamba state. The gap therefore reflects compression, not a change in numeric format. The prefill-heavy 50K/2K regime gains least. The decode-heavy 8K/64K regime gains most.
On a single 8xH100 node at UT = 100, the gains are smaller. They are 1.91x on 50K/2K and 1.82x on 8K/64K. Both models there use FP8 weights, FP8 KV cache, and FP32 Mamba state.
On a single H100 at 1M context, the binding constraint flips from compute to memory. Super’s NVFP4 weights occupy about 70 GB of the 80 GB HBM budget. Each 1M-token request adds about 4 GB of KV cache. Effective concurrency is therefore 1.
Puzzle-75B-A9B’s NVFP4 weights occupy about 44.5 GB. Attention layout is unchanged, so per-request KV cost is unchanged. Concurrency at 1M rises to 8. Aggregate decode throughput at that concurrency is roughly 4x Super’s single-request throughput. Prefill of a 990K-token prompt is about 1.2x faster.
How Iterative Puzzle Works
Puzzle is a decomposed neural architecture search framework, implemented here as Puzzletron. It defines a discrete search space of alternative layer implementations. Each alternative gets a quality score. A mixed-integer program then selects one alternative per layer under a deployment constraint.
Three pruning techniques form the search space:
Intermediate channel pruning: Channels inside each routed expert are ranked by contribution to the expert’s output. All experts within one MoE layer are pruned to a uniform size, for kernel compatibility.
Top-k reduction: The number of experts a token is routed to varies per layer, up to the parent’s k=22.
Mamba SSM pruning: The SSM state size drops from 128 to 96 channels.
The SSM result is measured. Dropping 128 channels to 96 speeds the SSM kernel 1.2x to 1.3x during decode. This holds at batch sizes between 8 and 512. Channels were ranked by estimated contribution to the Mamba layer output. The estimate averaged over 67M tokens of validation data. Appendix A shows this beats random channel selection under aggressive pruning.
The original formulation assumes replacement quality impacts are approximately additive. Each candidate block is scored inside the unmodified parent. That ignores higher-order interactions between replacements.
Iterative Puzzle alternates bounded compression with short knowledge distillation recovery. It builds a sequence M0, M1, … MR instead of jumping to the target. Scores are recomputed against the current compressed model, not the original parent.
Three stages were used:
MoE weights to 75% of teacher capacity, Mamba SSM state to 75%. Healed for 24B tokens.
MoE weights to 60% of teacher capacity. Healed for 43.2B tokens.
Activated routed-expert budget to 50%, allocated heterogeneously. Healed for 52.8B tokens.
The above table compares this against a single-step Puzzle baseline at the same target. The three-step procedure averages 69.05 across ten benchmarks, against 68.48. Gains appear on MMLU-Pro, GPQA, HLE, AA-LCR, LiveCodeBench, SciCode, and RULER-256K. IFBench-Instruction fell 0.2 points and IFBench-Prompt fell 0.5.
Recovery: Distillation, RL, and Verbosity
Knowledge distillation ran on 30% pretraining data and 70% SFT data from Nemotron-3-Nano. During the Puzzle phase, KD used a 32K sequence length. Recovery then trained at 128K, and scaled to 512K. The budget was up to 100B tokens, with a 16M-token global batch, in Megatron-LM.
RL post-training adopted Stage 2 of the Nemotron-3-Super RL pipeline, focused on software engineering. Phase 2.1 did single-step tool-use comparison. Phase 2.2 moved to end-to-end sandbox RL, where agents run up to 200 turns. Both phases used a KL penalty of 0. The team swept learning rates, then averaged the resulting weights.
The above Figure 4 shows what each stage contributed. Short-context KD recovers most categories to over 97% of Nemotron-3-Super. Long-context KD then lifts long-input and long-generation benchmarks specifically. The research team states that RL’s impact in these experiments was small.
Verbosity is the quiet detail. After the last Puzzle iteration, the model generated 132% of Super’s token count. That fell to 99% after the full recovery pipeline.
Deployment: Quantization and Multi-Token Prediction
Two post-training quantization recipes were produced: FP8 W8A8 targets Hopper and NVFP4 W4A4 targets Blackwell.
ComponentBF16 baselineFP8 checkpointNVFP4 checkpointSparse and shared MoE GEMMsBF16FP8NVFP4Mamba GEMMsBF16FP8FP8Mamba SSM cacheFP32FP32FP16+SRKV cacheFP8FP8FP8RouterFP32FP32FP32Attention QKV/output, MoE latent projections, LM headBF16BF16BF16
Both recipes calibrated on 256 post-training SFT samples. NVFP4 used max calibration, not the AutoQuantize sensitivity search used for Super. The resulting checkpoint is slightly more aggressively quantized, and performed similarly.
NVFP4 is not natively supported on Hopper. It is still used for the 1M-context H100 target, because HBM capacity binds there.
Puzzle-75B-A9B inherits a shared MTP head from Super. Parameters are shared across MTP steps, so one head applies recursively at inference. Transferring Super’s trained head directly gave similar acceptance lengths.
The research team then identifies a training-inference mismatch. Teacher-forced MTP training feeds the full shifted hidden-state sequence. Autoregressive drafting instead feeds a mixture of target-model and MTP-generated hidden states. Acceptance rates fall at deeper draft positions.
Continued training on the transferred head addresses this. On SPEED-Bench at draft length 7, average acceptance length rose from 3.45 to 4.34. That is roughly 25% to 30%, concentrated at later draft positions. Unlike Super, the NVFP4 checkpoint barely degrades: 4.31 against 4.34.
Where Compression Helps and Where It Hurts
Benchmark (BF16)SuperPuzzle-75B-A9BDeltaMMLU-Pro83.882.4-1.4AIME25 (no tools)92.289.7-2.5GPQA (no tools)80.578.6-1.9LiveCodeBench82.181.1-1.0SciCode (subtask)42.340.6-1.7SWE-Bench (OpenHands)59.556.9-2.6Arena-Hard-V272.868.6-4.2AA-LCR56.856.9+0.1RULER 1M93.992.2-1.7MMLU-ProX79.577.5-2.0
The research paper’s own summary is that instruction-following and agentic evaluations lose most. Arena-Hard-V2 is the worst case, at -4.2 points. RULER stays within roughly 1 to 2 points at 256K, 512K, and 1M.
Three BF16 results do not regress. AA-LCR gains 0.1, Scale AI Multi-Challenge ties at 56.6, and TauBench Telecom gains 0.4.
NVFP4 costs little on top of compression. On RULER 1M the NVFP4 checkpoint scores 93.2, above BF16’s 92.2. HLE is the clearest NVFP4 cost, dropping from 16.5 to 15.7. FP8 results sit in Appendix E, and track BF16 closely. SWE-Bench is not reported for the FP8 checkpoint.
Ultra-long-context RAG on one GPU: A document analysis service at 1M context goes from 1 concurrent request to 8. Aggregate decode throughput at that concurrency is roughly 4x.
Interactive coding assistants: At UT >= 100 tok/s in the 8K/64K regime, one node serves 2.03x the tokens. Adjusted for verbosity, that is 2.16x the completed requests per minute.
Prefill-heavy document pipelines: The 50K/2K regime gains only 1.60x. Compression helps less when prompt processing dominates compute.
Agentic SWE loops: Check the 2.6-point SWE-Bench gap against your task mix. RL recovery targeted this capability, and only partly restored it.
1.60x to 2.14x total throughput over Super at matched NVFP4 and matched user throughput
1M-token concurrency on a single H100 rises from 1 request to 8
MTP acceptance length improves from 3.45 to 4.34 on SPEED-Bench at draft length 7
Long-context accuracy holds within 1 to 2 points on RULER at 256K, 512K, and 1M
Generation verbosity ends at 99% of Super, so token gains survive at request level
Three checkpoints published: BF16, FP8, and NVFP4
Arena-Hard-V2 drops 4.2 points and SWE-Bench drops 2.6 points
RL recovery had a small measured impact, which the paper states directly
Mamba pruning is uniform, because frameworks cannot vary SSM state size per layer
Latent-dimension pruning was dropped: NVFP4 MoE kernels need a latent dim that is a multiple of 512
Prose and tables disagree on several throughput multiples in this v2 preprint
Disaggregated prefill gains are only 5% to 7%, and add serving complexity