突破 DeepSeek-V4-Pro 服务能力的极限
张天宇、高雨松、张云
1. 引言
DeepSeek-V4-Pro 是一个 1.6 万亿参数的混合专家(MoE)模型,同时发布了 FP8 和 FP4 两种权重的版本。这种规模的模型天然受益于 NVIDIA Blackwell GPU 等加速器,后者提供更大的 HBM 容量、更高的计算吞吐量以及原生的 FP4 Tensor Core 支持。然而,H20 GPU 尽管缺乏这些优势,仍被广泛部署。
硬件限制并不会降低服务要求。长上下文预填充仍须控制首 token 时间(TTFT)。交互式解码必须满足每个服务层级的时间每输出 token(TPOT)目标。持续流量必须在聚合吞吐量与 KV 缓存容量之间取得平衡。短输入、长上下文、延迟敏感型请求以及高并发会以不同方式给系统施加压力;没有任何一种通用配置能够同时很好地应对所有这些情况。
一个模型需要多种服务配置。工作负载特征、服务级别目标(SLO)以及实测的硬件行为共同决定部署拓扑和执行路径:
- 将服务配置与工作负载相匹配。在本文评估的配置中,预填充根据实测的上下文长度范围在 PP2 和 PP4 之间进行选择,而解码则使用针对不同延迟、吞吐量和 KV 容量目标优化的配置。
- 优化预填充路径。我们优化了 Attention-CP8 → MoE-TP8 以及上下文并行通信,然后针对长上下文和短上下文工作负载产生的真实路由形态进行调优。
- 优化解码路径。我们优化了 DSpark 投机解码路径,针对不同的解码 SLO 细化执行、专家路由以及通信与计算的重叠。
突破延迟极限。在批大小为 1 时,单节点 H20-141GB 参考配置达到 271 输出 token/秒,而 B300 上报告的数值为 383.7 token/秒。尽管硬件差距巨大,针对特定工作负载的系统优化将实测解码性能比缩小至 1.42 倍。详细的基准测试设置以及基于日志的吞吐量提取方法见附录 B.3。
覆盖服务(serving)的完整性能包络。延迟结果只代表系统的一个侧面。在更广泛的配置(profile)族中,优化后的预填充(prefill)达到每节点每秒 8.45k 输入 token,并在 43.7 秒内处理完一个 1M token 的提示词。在面向吞吐量的解码(decode)方面,DP16-EP16 效率参考配置达到每节点每秒 4.67k 输出 token,对应平均 TPOT 为 27.4 毫秒。这些结果有意取自不同的配置,每个配置都针对上下文长度、延迟、吞吐量和容量约束的不同组合进行了选择和优化。
这里的贡献是一套方法论,而非单一基准。场景化服务(scenario-specific serving)允许每个工作负载在现有硬件上评估的各配置中,向更优的实测运行点移动。我们希望这里呈现的部署选择、优化方法和实测结果,能为在算力、内存、带宽或互联约束下服务前沿模型的团队提供实用参考。
2. 从硬件约束到服务配置(Serving Profiles)
2.1 硬件约束与服务角色
图 1. 硬件差距:H20 对比 B300。
Blackwell 提供原始性能;H20 提供可规模化的部署能力。B300 具备原生 FP4 Tensor Core、高得多的 FP8 吞吐量,以及显著更大的 HBM。H20 无法匹敌其计算能力,但它仍可大规模获取,并提供高内存带宽和 900 GB/s 的 NVLink。本研究中的每个节点包含八块通过 NVLink 连接的 GPU。预填充不保留跨请求的长期状态,因此其硬件选择主要受 TTFT、计算和通信效率的制约。解码必须在整个生成过程中保留每个活跃请求的 KV cache,这使得 HBM 容量直接限制上下文长度和并发度。在本研究所涉及的部署中,这促使我们使用 H20-141GB 用于解码,使用 H20-96GB——其容量足以满足我们的预填充工作负载——用于预填充。
图 2. 按服务角色的硬件分配。
2.2 容量选择
服务容量最终来源于共享的 HBM 预算:模型权重和每请求的 KV 状态会竞争同一块内存。我们将“全 token 容量”定义为:在模型权重和运行时缓冲区分配完毕后,每个 rank 所能容纳的全注意力 KV token 的最大数量。它是一个内存上限,而非可接受批处理规模的直接保证。
使用 Humming MXFP4AFP8 缩减权重占用
首先缩减权重占用。Humming MXFP4AFP8 使用 MXFP4 专家权重配合在线 FP8 激活,以减少 H20 GPU(缺乏原生 FP4 Tensor Core)上的权重占用和内存流量。SGLang 集成已可在 sglang#23754 中获取。我们将在后续专门文章中介绍 Humming/SGLang 集成。模型级精度结果和公开参考测量数据见附录 D.2。
通过 Online C128 扩展 KV 容量
为 KV 缓存留出增长空间。Offline C128 基线为每个压缩页保留逐索引状态。Online C128 则维护一个紧凑的聚合状态,将更多 HBM 释放给 KV 缓存池。它会引入额外的状态维护和推测验证工作,但我们在测试中未观察到 TPOT 回退。
综合容量提升
图 3. 使用 Humming MXFP4AFP8 和 Online C128 的容量扩展。
容量提升在权重和 KV 状态两个维度上叠加生效。通过缩减权重占用,Humming MXFP4AFP8 将全 token 容量扩展到 Baseline FP8 + Offline C128 配置的 1.71 倍(DP32-EP32)和 4.47 倍(PP2-TP8)。Online C128 随后缩减 C128 辅助状态占用,在 Humming 基础上再提供 2.268 倍的提升。两项技术结合后,容量分别提升至基线的 3.88 倍(DP32-EP32)和 10.14 倍(PP2-TP8)。附录 D.1 提供了完整数据。
2.3 场景特定的服务配置
Prefill 配置
图 4. Prefill 配置:相同执行路径,不同流水线深度。
正确的流水线深度取决于需要流水化的工作量大小。PP2-CP8-TP8 和 PP4-CP8-TP8 共享相同的 Attention-CP8 → MoE-TP8 执行路径。在拓扑层面,它们的主要区别在于流水线深度:PP2 将模型分布在两个阶段,而 PP4 使用四个阶段。
短上下文更倾向于较低的流水线开销;长上下文则能暴露更多的并行度。短输入产生的块(chunk)较少,会使更深的流水线填充不足,从而让填充(fill)、排空(drain)和跨阶段传输的成本更加突出。长上下文能提供足够的块,使四个阶段保持忙碌;由于每个阶段的层数更少,额外的节点便转化为更多的预填充并行度。在我们的部署中,这些特性促使我们对较短上下文采用 PP2-CP8-TP8,而对长上下文工作负载采用 PP4-CP8-TP8。
低延迟解码配置
图 5. 低延迟解码:TP8 参考配置与 PP2-TP8 服务配置。
低延迟始于最短的执行路径。单节点 TP8 与 PP2-TP8 共享相同的 Attention-TP8 → MoE-TP8 执行路径;区别在于模型是否跨节点划分。单节点 TP8 将所有层放置在单个 H20-141GB 节点上,避免了跨阶段通信和同步。PP2-TP8 则将模型划分到两个流水线阶段上。
最快的拓扑并不总是最实用的。单节点 TP8 执行路径更短,但模型权重和服务状态共享单个节点的 HBM,留给 KV 缓存的空间有限。它无法同时支持长上下文和更大的批处理规模。PP2-TP8 虽然增加了额外的流水线开销,但将模型权重分布到两个节点上,为 KV 状态释放了更多 HBM。针对我们的延迟和容量目标,我们采用单节点 TP8 作为批大小为 1 的延迟参考,并采用 PP2-TP8 作为低延迟服务配置。
高吞吐解码配置
图 6. 高吞吐解码:DP16-EP16 参考配置与 DP32-EP32 容量配置。
高吞吐解码同时扩展数据并行和专家并行。两种配置都使用 Attention-DP → MoE-EP 执行路径。DP16-EP16 是最小的部署单元;DP32-EP32 在相同拓扑内同时扩展 DP 和 EP。
扩展(Scale-out)优先考虑请求容量而非单 GPU 吞吐量。更大的 EP 组将专家权重分布到更多 GPU 上,释放 HBM 给 KV cache,从而容纳更多并发请求。与此同时,MoE 流量中留在节点内的比例变小,跨节点传输的比例变大,这可能会降低单 GPU 效率。在本文评估的配置中,我们以 DP16-EP16 作为最小部署单元和效率基准,以 DP32-EP32 来扩展请求容量。
3. 预填充(Prefill):平衡计算与通信
预填充性能是一个系统性问题。专家负载不均衡、上下文并行通信以及生产环境的路由形态共同决定了 TTFT;仅仅优化单个内核是不够的。
3.1 为什么选择 MoE-TP 而非 MoE-EP
图 7. 用 MoE-TP 替代 MoE-EP。
流量更少也可能耗时更长。MoE-EP 只交换被路由的 token,但真实预填充流量表现出显著的专家倾斜。拥有热门专家的 rank 需要执行更多计算,成为掉队者(straggler);所有其他 rank 在合并(combine)步骤都要等待最慢的路径。更低的通信量并不等于更低的 TTFT。
在最小化流量之前先平衡计算。对于本文评估的 H20 预填充工作负载,PP2 和 PP4 均使用 MoE-TP。全序列的 all-gather 和 reduce-scatter 会引入更多通信,但这些流量走的是高带宽 NVLink,成本稳定且可预测。所有 TP rank 对相同的被路由 token 执行张量并行计算,从而避免专家倾斜演变为 rank 级别的长尾问题。对于该工作负载,可预测的通信比不可预测的不均衡更划算。该实现可在 sglang#24947 中获取。
3.2 加速与融合预填充集合通信
图 8. 对称内存集合通信与预填充融合。
构建可复用的集合通信快速路径。MoE-TP 用可预测的集合通信流量取代了不可预测的专家负载不均衡,使通信效率成为下一个瓶颈。我们让对称内存在 TP 和 CP 之间可复用,使 AllReduce、AllGather 和 ReduceScatter 能够共享已注册缓冲区的快速路径,并适用相应的 Hopper 加速。上游配套工作涵盖内存池所有权、通信器注册、MoE-TP 集合通信缓冲区,以及 CP Attention 和 KV 缓存缓冲区路径。
然后缩短 Prefill 关键路径。仅靠更快的集合通信并不能消除通信与计算之间的边界。针对 32K 单块场景,我们构建了一条融合路径,将基于拷贝引擎驱动的 AllGather 与融合 FP8 量化和共享专家 GEMM 重叠执行,然后在第二个 Triton kernel 中合并 TopK 归约、共享专家加法和 ReduceScatter。这七个算子被重组为三个执行组,在匹配的 PP4 A/B 测试中,TTFT 降低了约 3.5%。
3.3 针对真实路由形态调优 Humming
图 9. 针对真实路由形态调优 Humming。
通用调优无法覆盖真正重要的形态。Prefill 路由将 token 不均匀地分布到 384 个专家上,因此有效的 M 维度聚集在一小组离散值上。W13 和 W2 也作用于不同的形态,因此单一的通用启发式策略无法同时优化两条路径。
基于生产路由进行调优。我们从真实路由直方图中提取高频形态,为 W13 和 W2 分别构建精确形态配置,并在 kernel、流水线阶段和匹配 A/B 层面进行验证。优化目标不是合成范围的 M,而是我们实际服务的路由分布。在 32K 的匹配 PP4 A/B 测试中,选定的 MoE kernel 延迟降低约 21%,转化为端到端 TTFT 降低 11.35%。
4. Decode:优化投机采样与 MoE 执行
在我们的实现中,解码优化是特定于配置文件的。PP2-TP8 需要跨推测流水线阶段进行协调,而 DP32-EP32 则专注于在高并发下优化精炼步骤和专家路由。Humming 融合与重叠优化了这些服务拓扑之下的共享 MoE 热路径。
4.1 低延迟 PP2-TP8:跨流水线阶段扩展 DSpark
图 10. 跨 PP2 阶段协调 DSpark。
流水线并行分割了推测循环。在 PP2-TP8 中,目标执行跨越两个流水线阶段,而 DSpark 草稿模型仅驻留在最后一个阶段。阶段 0 将目标隐藏状态发送到阶段 1,阶段 1 执行验证、接受 token,并为下一轮生成候选。
让两个阶段如同一个整体般推进。每一轮推测都跨越流水线边界。我们在一个统一的执行协议下协调两个阶段及所需的中间传输,防止阶段进入不同的轮次,同时避免冗余同步。PP 特定的 DSpark 集成正在被合入 sglang#32281。
4.2 高吞吐 DP32-EP32:消除高并发瓶颈
图 11. DP32-EP32 瓶颈消除。
本小节中匹配的 A/B 结果使用 DP32-EP32,在 4K 上下文下每个 DP rank 有 32 个并发请求。
为精炼步骤选择正确的执行形态。精炼步骤应用全词表投影来对 DSpark 的候选集重新评分。在高并发下,逐行的点积归约会为每个活动行反复读取词表权重,在每个解码步骤中造成持续的尾部延迟。我们将活动行合并为一次转置 GEMM,减少了冗余的内存流量并缩短了精炼路径。每 GPU 吞吐量提升了 22.8%。
根据实测路由放置专家。DSpark 流量也表现出显著的专家倾斜。我们从代表性请求中记录路由亲和性,并使用它来配置专家并行负载均衡(EPLB)和冗余专家,防止少量热门专家反复延长关键路径。每 GPU 吞吐量提升了 13.5%。
4.3 Humming 解码热路径:融合与重叠
图 12. Humming 解码热路径优化。
这些优化位于服务拓扑层之下,可被基于 Humming 的解码配置复用。下方匹配结果采用 DP32-EP32 配置,4K 上下文,每个 DP rank 并发 32 个请求。
移除额外的量化过程。我们将 SwiGLU 激活与量化融合,使融合后的内核直接生成 W2 所需的数据和缩放因子。这消除了对中间缓冲区的重复访问,并移除了独立的量化过程,使 W2 能够更早启动。在匹配的 DSpark A/B 测试中,单卡吞吐量提升了 44.0%。
将通信与 W2 重叠。我们将此前工作(sglang#9660)中的单批次重叠(SBO)机制适配为 Humming-Aware SBO。基于逐 tile 的信号,DeepEP 可以在某个 W2 输出 tile 完成时立即启动对应的 combine 发送,无需等待整个 GEMM 完成。在相同运行点下,此前的非投机匹配 A/B 测试中,SBO 相比 FP8 传输层级恢复了 4.12% 的吞吐量。
5. 评估:系统收益与配置权衡
5.1 预填充:累积收益与上下文长度权衡
图 13. 预填充吞吐量累积增益。
PP2 强化了短上下文配置。PP2 在全部九个输入长度上均有提升,几何平均吞吐量增益为 36.5%,峰值总输入吞吐量达 16,900 tokens/s。其更浅的流水线减少了短请求的填充与排空开销,使 PP2 能够以更少的资源维持更低的 TTFT。
PP4 将增益延续至长上下文。PP4 在同样的九个测试点上实现了 31.8% 的几何平均吞吐量增益。随着上下文长度增长,更深的流水线有足够的工作量来摊薄其固定成本:总输入吞吐量在 512K 时达到 25,860 tokens/s,在 1M 时仍保持 23,970 tokens/s。
图 14. PP2 与 PP4 之间的 TTFT 权衡。
上下文长度会改变 PP2/PP4 之间的权衡关系。相对于 PP4,PP2 在 4K 上下文下将 TTFT 降低了 16.7%,在 32K 下降低了 19.5%。在 8K、16K 和 64K 下,两种配置的差异保持在 2% 以内。从 128K 开始,PP4 建立起决定性优势,在 128K、256K、512K 和 1M 下,相对于 PP2 的 TTFT 分别降低了 26.2%、33.3%、42.1% 和 44.8%。因此,我们将路由边界视为一种基于实测上下文长度范围得出的运行策略,而非一个通用的交叉点。
附录 A.1–A.2 提供了完整的 TTFT 和总输入吞吐量结果。
5.2 低延迟解码:性能与容量权衡
图 15. 优化后 DSpark 的峰值 TPOT 提升。
优化后的 DSpark 重新设定了延迟基线。在图 15 所示的四种输入长度下,优化后的 DSpark 在批大小为 1 时将峰值 TPOT 降低了 74.8%–78.0%。在每对测量所共有的最大批大小下,降幅仍保持在 52.2%–60.0%。这一提升在 8K 到 1M 的范围内均能保持,而非仅限于短上下文或单请求执行场景。
图 16. 批大小为 1 时的解码吞吐量:H20-141GB 与 B300 参考对比。
实测服务性能远高于仅凭峰值算力比值所预期的水平。在图 16 所示的四种输入长度下,采用 PP2-TP8 配置的优化 DSpark 在批大小为 1 时达到 150–174 tokens/s。单节点 TP8 参考配置达到 183–271 tokens/s。就实际执行路径所使用的精度而言,B300 的峰值 Tensor Core 算力约为 H20-141GB 的 45.6 倍(B300 FP4 对比 H20 FP8),内存带宽为其 1.67 倍。然而,实测的最高生成速率分别为 B300 上的 383.7 tokens/s 和 H20-141GB 上的 271 tokens/s——比值为 1.42 倍。即便面对如此强大的硬件参考,针对工作负载的优化仍使 H20-141GB 参考配置在实测服务性能上大幅接近对手。
就我们的生产目标而言,容量更倾向于 PP2-TP8。单节点 TP8 速度更快,但在 1M 上下文下,其 KV-cache 容量仅够 batch size 为 1 时使用,无法容纳更大的 batch 或更多并发请求。通过将模型权重分布到两个流水线阶段,PP2-TP8 在 1M、512K 和 256K 上下文下分别支持 batch size 4、8 和 16。配合 Online C128,其全 token 容量可达 11.04M tokens/rank。对于与我们类似的上下文长度和并发目标,我们建议保留单节点 TP8 作为延迟基准,并使用 PP2-TP8 作为低延迟服务配置。附录 B 和附录 D.1 提供了完整的性能和容量数据。
5.3 高吞吐解码:前沿增益与配置权衡
图 17. 吞吐量–交互性帕累托前沿。
图 17 展示了吞吐量–交互性前沿如何随系统演进。横轴是交互性,单位为 tokens/s/user;纵轴是吞吐量,单位为 tokens/s/GPU。在这些 DP/EP 配置中,每个 DP rank 对应一个 GPU;交互性等于每 GPU 吞吐量除以每个 DP rank 的并发请求数。越靠近右上方的点,代表用户可见生成速度与 GPU 效率的组合越优。四条曲线代表系统的累积演进,而非第 4 节中任何单一优化的孤立收益。
MTP 指多 token 预测;(3, 1, 4) 配置使用三步投机采样、top-k 为 1、四个草稿 token。
系统优化推动了整个前沿的移动。在 4K 上下文、每个 DP rank 32 个并发请求下,每 GPU 吞吐量从 319.92 tokens/s/GPU 提升至 703.15 tokens/s/GPU,提升 2.20 倍。在 1M 上下文、每个 DP rank 1 个请求下,吞吐量从 27.05 tokens/s/GPU 提升至 66.82 tokens/s/GPU。前三个系统里程碑在 1M 上下文下每个 DP rank 只能处理一个请求;最终系统支持四个请求,并达到 177.48 tokens/s/GPU。扩展后的运行范围既来自更快的执行速度,也来自更大的容量。
图 18. 每 GPU 吞吐量:DP16-EP16 与 DP32-EP32 对比。
较小的部署单元可在选定的高并发运行点上保持效率。在我们此前基于 H20 服务 DeepSeek-V3/R1 的工作中,我们发现较小的 EP 部署单元可以在每个节点内保留更大比例的 MoE 流量。DeepSeek-V4-Pro 在图 18 绘制的运行点上表现出同样的优势:当每个 DP rank 的并发请求数为 16 和 32 时,DP16-EP16 的每 GPU 吞吐量比 DP32-EP32 高出约 3.6%–20%。完整扫描结果并非在每个并发级别上都呈单调变化,因此我们将 DP16-EP16 作为效率参考,而非 DP32-EP32 的通用替代方案。
图 19. 每个 DP rank 的长上下文请求容量。
容量改变了首选的高吞吐量配置。DP16-EP16 在每 GPU 上更高效,但 DP32-EP32 将专家权重分布到更多 rank 上,为 KV cache 释放了更多 HBM。在 256K、512K 和 1M 上下文长度下,每个 DP rank 的最大并发请求数分别从 8、4、2 增加到 16、8、4——一致地实现了 2 倍扩展。对于与我们类似的长上下文并发目标的部署,这种额外容量使 DP32-EP32 成为面向容量的高吞吐量配置,而 DP16-EP16 仍可作为效率参考。附录 C 和附录 D.1 提供了完整数据。
6. 结论
一个模型并不需要一种折中的配置。我们为 H20 上的 DeepSeek-V4-Pro 构建了面向场景的服务栈。Prefill 根据上下文长度在 PP2 和 PP4 之间切换。Decode 使用 PP2-TP8 实现低延迟,使用 DP32-EP32 实现高吞吐量。通过协同设计容量、部署拓扑和执行路径,H20 在计算资源有限且缺乏原生 FP4 Tensor Cores 的情况下,仍能支撑 1M token 上下文并满足多种服务 SLO。
可迁移的成果是一套以场景驱动的方法论。服务配置不应仅从硬件规格或孤立的基准测试中选取。我们建议从工作负载、SLO、上下文长度和并发度出发,然后通过性能剖析来识别瓶颈资源,并将其转化为具体的拓扑和执行路径决策。我们希望这套方法论能帮助 AI 基础设施团队在多样化的资源约束下——无论瓶颈是算力、内存容量、内存带宽还是互联——构建实用的前沿模型服务系统,并将这些经验分享给更广泛的开源生态。
致谢
我们感谢 SGLang 团队和社区在 SGLang 框架上的杰出工作。我们也感谢以下团队和合作者的支持与贡献:
- 蚂蚁集团 SCT 团队:徐永飞、张倩瑜、顾泽凯、黄志林、王法康、傅建豪、杜卓轩、詹霞、黄纯、刘琦、陈曦、毛宇涵、程培成、高翰林、姚静华
- 蚂蚁集团 Venus 团队:林金镇
- SGLang 社区:张鹏
附录 A. 预填充结果
A.1 Humming PP2 预填充:基线 vs. 最终配置
| 输入长度 | 基线 TTFT(毫秒) | 基线总输入吞吐量(tokens/秒) | 最终 TTFT(毫秒) | 最终总输入吞吐量(tokens/秒) |
|---|---|---|---|---|
| 4K | 775.8 | 5,280 | 573.3 | 7,140 |
| 8K | 1202.1 | 6,810 | 907.6 | 9,030 |
| 16K | 2059.8 | 7,950 | 1649.5 | 9,930 |
| 32K | 4137.5 | 7,920 | 2470.3 | 13,260 |
| 64K | 6195.7 | 10,580 | 4063.8 | 16,130 |
| 128K | 10744.4 | 12,200 | 7975.9 | 16,430 |
| 256K | 20542.2 | 12,760 | 15507.2 | 16,900 |
| 512K | 44544.6 | 11,770 | 34982.6 | 14,990 |
| 1M | 100304.2 | 10,450 | 79214.2 | 13,240 |
A.2 Humming PP4 预填充:基线 vs. 最终配置
| 输入长度 | 基线 TTFT(毫秒) | 基线总输入吞吐量(tokens/秒) | 最终 TTFT(毫秒) | 最终总输入吞吐量(tokens/秒) |
|---|---|---|---|---|
| 4K | 924.6 | 4,430 | 687.9 | 5,950 |
| 8K | 1174.5 | 6,970 | 890.3 | 9,200 |
| 16K | 2202.0 | 7,440 | 1635.4 | 10,020 |
| 32K | 4185.6 | 7,830 | 3068.4 | 10,680 |
| 64K | 5252.4 | 12,480 | 3982.6 | 16,460 |
| 128K | 7793.4 | 16,820 | 5882.5 | 22,280 |
| 256K | 13210.7 | 19,840 | 10348.9 | 25,330 |
| 512K | 26350.1 | 19,900 | 20273.1 | 25,860 |
| 1M | 55532.3 | 18,880 | 43742.5 | 23,970 |
附录 B. 低延迟解码结果
B.1 不同输入长度和批处理大小下的峰值 TPOT
B.1.1 无投机解码 PP2-TP8
| 输入长度 / 批处理大小(峰值 TPOT,毫秒) | 1 | 2 | 4 | 8 | 16 |
|---|---|---|---|---|---|
| 8K | 26.39 | 30.86 | 31.31 | 31.79 | 31.74 |
| 32K | 25.72 | 26.58 | 27.81 | 31.06 | 37.97 |
| 64K | 25.75 | 26.62 | 28.13 | 29.19 | 38.75 |
| 128K | 25.94 | 26.94 | 28.38 | 29.75 | 38.51 |
| 256K | 26.08 | 27.21 | 28.84 | 32.43 | 38.83 |
| 512K | 26.25 | 27.51 | 29.16 | 33.70 | - |
| 1M | 26.42 | 27.81 | 29.52 | - | - |
B.1.2 优化版 DSpark PP2-TP8
| 输入长度 / 批大小(峰值 TPOT,毫秒) | 1 | 2 | 4 | 8 | 16 | 32 |
|---|---|---|---|---|---|---|
| 4K | 5.91 | 6.76 | 7.97 | 10.00 | 14.55 | 19.23 |
| 8K | 5.80 | 6.87 | 8.85 | 10.48 | 15.18 | 19.60 |
| 32K | 6.14 | 7.04 | 8.39 | 10.83 | 14.86 | 20.46 |
| 64K | 6.15 | 7.13 | 8.73 | 10.39 | 15.49 | 21.65 |
| 128K | 6.77 | 7.02 | 8.91 | 11.59 | 16.17 | 24.78 |
| 256K | 5.76 | 6.98 | 8.61 | 11.98 | 17.72 | - |
| 512K | 6.35 | 7.95 | 9.87 | 14.30 | - | - |
| 1M | 6.65 | 8.92 | 12.43 | - | - | - |
B.2 批大小为 1 时的输出吞吐量
| 输入长度 | 无投机解码 PP2-TP8(tokens/s) | 优化版 DSpark PP2-TP8(tokens/s) | 单节点 TP8(tokens/s) |
|---|---|---|---|
| 4K | - | 169 | 213 |
| 8K | 38 | 172 | 260 |
| 16K | - | - | 244 |
| 32K | 39 | 163 | 269 |
| 64K | 39 | 163 | 246 |
| 128K | 39 | 148 | 267 |
| 256K | 38 | 174 | 271 |
| 512K | 38 | 157 | 254 |
| 1M | 38 | 150 | 183 |
B.3 基准测试设置
硬件:一个 8× H20-141GB 解码节点。
解码服务器
--tp-size 8 \
--mem-fraction-static 0.91 \
--max-running-requests 1 \
--cuda-graph-max-bs 1 \
--cuda-graph-bs 1 \
--moe-runner-backend humming \
--moe-a2a-backend none \
--speculative-algorithm DSPARK \
--speculative-num-draft-tokens 7 \
--speculative-dspark-block-size 7 \
--speculative-moe-runner-backend triton \
--speculative-moe-a2a-backend none
客户端基准测试
python3 -m sglang.bench_serving \
--backend sglang-oai-chat \
--host 127.0.0.1 \
--port 8000 \
--model <MODEL_PATH> \
--dataset-name random \
--dataset-path <DATASET_PATH> \
--random-input-len 262144 \
--random-output-len 4096 \
--random-range-ratio 1.0 \
--num-prompts 10 \
--max-concurrency 1 \
--warmup-requests 0 \
--seed 1
输出吞吐量取自服务器 TP0 解码批处理日志行;我们丢弃最高和最低 20% 的样本,对剩余部分取平均值。B300 的数据遵循所链接来源中报告的设置。
附录 C. 高吞吐量解码结果
C.1 DP32-EP32 搭配 FP8 + MTP(3, 1, 4)
| 输入长度 / 每个 DP 秩的并发请求数(tokens/s/GPU) | 1 | 2 | 4 | 8 | 16 | 32 |
|---|---|---|---|---|---|---|
| 4K | 30.49 | 58.58 | 102.89 | 174.75 | 253.15 | 319.92 |
| 8K | 30.34 | 58.29 | 102.38 | 174.67 | 251.62 | 318.32 |
| 16K | 29.70 | 56.55 | 99.47 | 170.01 | 242.22 | 302.43 |
| 32K | 29.58 | 56.35 | 98.28 | 164.26 | 234.13 | - |
| 64K | 29.07 | 55.73 | 96.43 | 161.60 | - | - |
| 128K | 28.39 | 54.06 | 92.89 | 153.55 | - | - |
| 256K | 28.35 | 53.02 | 90.89 | - | - | - |
| 512K | 27.51 | 51.49 | - | - | - | - |
| 1M | 27.05 | - | - | - | - | - |
C.2 DP32-EP32 搭配 FP8 + 优化版 MTP(3, 1, 4)
| 输入长度 / 每个 DP 秩的并发请求数(tokens/s/GPU) | 1 | 2 | 4 | 8 | 16 | 32 |
|---|---|---|---|---|---|---|
| 4K | 36.84 | 69.86 | 131.96 | 232.94 | 389.94 | 514.77 |
| 8K | 32.58 | 69.51 | 131.53 | 222.06 | 348.80 | 416.82 |
| 16K | 31.89 | 67.44 | 127.79 | 216.14 | 341.85 | 395.99 |
| 32K | 31.49 | 67.21 | 124.49 | 208.83 | 337.97 | - |
| 64K | 30.95 | 66.47 | 123.68 | 205.44 | - | - |
| 128K | 30.22 | 64.47 | 119.14 | - | - | - |
| 256K | 30.18 | 63.23 | - | - | - | - |
| 512K | 29.28 | 61.40 | - | - | - | - |
| 1M | 28.79 | - | - | - | - | - |
C.3 DP32-EP32 搭配 FP8 + DSpark
| 输入长度 / 每个 DP 秩的并发请求数(tokens/s/GPU) | 1 | 2 | 4 | 8 | 16 | 32 |
|---|---|---|---|---|---|---|
| 4K | 53.1 | 94.8 | 181.2 | 338.1 | 495.8 | 591.8 |
| 8K | 44.5 | 88.4 | 170.1 | 317.3 | 495.5 | - |
| 16K | 43.6 | 88.3 | 165.3 | 308.8 | 455.5 | - |
| 32K | 43.0 | 87.3 | 161.0 | 298.4 | - | - |
| 64K | 42.3 | 86.3 | 158.0 | - | - | - |
| 128K | 41.3 | 83.8 | - | - | - | - |
| 256K | 41.2 | - | - | - | - | - |
| 512K | 40.0 | - | - | - | - | - |
| 1M | 39.3 | - | - | - | - | - |
C.4 DP32-EP32 搭配 Humming MXFP4AFP8 + 在线 C128 + DSpark
| 输入长度 / 每个 DP 秩的并发请求数(tokens/s/GPU) | 1 | 2 | 4 | 8 | 16 | 32 |
|---|---|---|---|---|---|---|
| 4K | 75.32 | 127.10 | 235.85 | 417.53 | 564.08 | 703.15 |
| 8K | 75.60 | 128.29 | 238.01 | 417.34 | 560.68 | 709.64 |
| 16K | 74.00 | 124.47 | 231.25 | 406.21 | 539.72 | 674.19 |
| 32K | 73.07 | 122.27 | 225.28 | 392.47 | 521.70 | 601.67 |
| 64K | 71.81 | 120.92 | 221.05 | 386.11 | 516.54 | 599.63 |
| 128K | 70.12 | 117.29 | 212.93 | 366.88 | 487.69 | - |
| 256K | 70.03 | 115.03 | 208.35 | 345.21 | 457.62 | - |
| 512K | 67.95 | 111.71 | 191.99 | 302.80 | - | - |
| 1M | 66.82 | 105.82 | 177.48 | - | - | - |
C.5 DP16-EP16 搭配 Humming MXFP4AFP8 + 在线 C128 + DSpark
| 输入长度 / 每个 DP 秩的并发请求数(tokens/s/GPU) | 1 | 2 | 4 | 8 | 16 | 32 |
|---|---|---|---|---|---|---|
| 4K | 76.80 | 129.62 | 236.83 | 397.42 | 584.37 | 759.73 |
| 8K | 76.69 | 130.55 | 237.53 | 398.60 | 582.03 | 762.09 |
| 16K | 76.16 | 127.88 | 233.10 | 388.54 | 571.22 | 745.23 |
| 32K | 74.07 | 124.77 | 226.24 | 378.79 | 559.05 | 722.51 |
| 64K | 74.57 | 124.69 | 223.84 | 373.13 | 541.46 | 695.35 |
| 128K | 72.36 | 120.34 | 219.66 | 365.13 | 518.98 | - |
| 256K | 71.38 | 119.19 | 211.64 | 340.72 | - | - |
| 512K | 69.54 | 115.14 | 198.81 | - | - | - |
| 1M | 67.39 | 106.50 | - | - | - | - |
附录 D. 容量结果
D.1 解码容量扩展
| 解码配置 | 配置 | 全 token 容量(tokens/秩) | 与上一阶段对比 | 与 FP8 基线对比 |
|---|---|---|---|---|
| DP32-EP32 | 基线 FP8 + 离线 C128 | 1,475,328 | - | 1.00× |
| Humming MXFP4AFP8 + 离线 C128 | 2,526,720 | 1.71× | 1.71× | |
| Humming MXFP4AFP8 + 在线 C128 | 5,731,328 | 2.268× | 3.88× | |
| PP2-TP8 | 基线 FP8 + 离线 C128 | 1,089,024 | - | 1.00× |
| Humming MXFP4AFP8 + 离线 C128 | 4,869,888 | 4.47× | 4.47× | |
| Humming MXFP4AFP8 + 在线 C128 | 11,044,906 | 2.268× | 10.14× |
D.2 Humming 精度验证
我们在 GSM8K1000 上评估了 DP16-EP16 Humming MXFP4AFP8 + 在线 C128 + DSpark 配置。该配置实现了 95.5% 的精确匹配准确率,有一个无效响应和零系统错误,通过了我们 95.0% 的验收阈值。
作为公开参考,上游 SGLang Humming 集成在 DeepSeek-V4-Flash 的 200 示例 GSM8K 评估中报告了以下结果:
| 后端 | GSM8K 准确率 |
|---|---|
| Marlin MXFP4A16 | 96.5%–97.0% |
| FlashInfer MXFP4 | 96.5%–97.0% |
| Humming MXFP4A16 | 96.5%–97.5% |
| Humming MXFP4AFP8 | 97.0% |
在这项公开对比中,Humming MXFP4AFP8 没有出现明显的精度下降。由于它使用的是 DeepSeek-V4-Flash 而非 DeepSeek-V4-Pro,我们将其视为外部参考,而不是对我们服务配置的匹配精度损失测量。
Pushing the Limits of Serving DeepSeek-V4-Pro
Tianyu Zhang, Yusong Gao, Yun Zhang
1. Introduction
DeepSeek-V4-Pro is a 1.6-trillion-parameter Mixture-of-Experts (MoE) model released with both FP8 and FP4 weights. Models at this scale naturally benefit from accelerators such as NVIDIA Blackwell GPUs, which offer more HBM, higher compute throughput, and native FP4 Tensor Cores. Yet H20 GPUs remain widely deployed, despite lacking those advantages.
Hardware constraints do not relax serving requirements. Long-context prefill must still control time to first token (TTFT). Interactive decode must satisfy the time-per-output-token (TPOT) target of each service tier. Sustained traffic must balance aggregate throughput against KV-cache capacity. Short inputs, long contexts, latency-sensitive requests, and high concurrency stress the system in different ways; no universal configuration can serve all of them well.
One model needs multiple serving profiles. Workload characteristics, service-level objectives (SLOs), and measured hardware behavior jointly inform the deployment topology and execution path:
- Match serving profiles to the workload. In the configurations evaluated here, prefill selects between PP2 and PP4 based on the measured context-length range, while decode uses profiles optimized for different latency, throughput, and KV-capacity targets.
- Optimize the prefill path. We optimize
Attention-CP8 → MoE-TP8and context-parallel communication, then tune for the real routing shapes produced by long- and short-context workloads. - Optimize the decode path. We optimize the DSpark speculative-decoding path, refine execution, expert routing, and communication–computation overlap for distinct decode SLOs.
Push the latency frontier. At batch size 1, the single-node H20-141GB reference reaches 271 output tokens/s, compared with the 383.7 tokens/s reported on B300. Despite the substantial hardware gap, workload-specific system optimization narrows the observed decode performance ratio to 1.42×. Detailed benchmark settings and the log-based throughput extraction methodology are provided in Appendix B.3.
Cover the serving envelope. The latency result represents only one edge of the system. Across the broader profile family, optimized prefill reaches 8.45k input tokens/s per node and processes a 1M-token prompt in 43.7 seconds. For throughput-oriented decode, the DP16-EP16 efficiency reference reaches 4.67k output tokens/s per node, corresponding to an average TPOT of 27.4 ms. These results intentionally come from different profiles, each selected and optimized for a different combination of context length, latency, throughput, and capacity constraints.
The contribution is a methodology, not a single benchmark. Scenario-specific serving allows each workload to move toward a better measured operating point among the profiles evaluated on the available hardware. We hope the deployment choices, optimization methods, and measurements presented here provide a practical reference for teams serving frontier models under compute, memory, bandwidth, or interconnect constraints.
2. From Hardware Constraints to Serving Profiles
2.1 Hardware Constraints and Serving Roles
Figure 1. Hardware Gap: H20 vs. B300.
Blackwell offers raw performance; H20 offers deployable scale. B300 provides native FP4 Tensor Cores, much higher FP8 throughput, and substantially more HBM. H20 cannot match its compute capability, but it remains available at scale and provides high memory bandwidth and 900 GB/s NVLink. Each node in this study contains eight GPUs connected by NVLink. Prefill does not retain long-lived per-request state, so its hardware choice is governed primarily by TTFT, compute, and communication efficiency. Decode must retain the KV cache of every active request throughout generation, making HBM capacity a direct limit on context length and concurrency. For the deployment studied here, this led us to use H20-141GB for decode and H20-96GB—whose capacity was sufficient for our prefill workloads—for prefill.
Figure 2. Hardware Assignment by Serving Role.
2.2 Capacity Choices
Serving capacity ultimately comes from a shared HBM budget: model weights and per-request KV state compete for the same memory. We define full-token capacity as the maximum number of full-attention KV tokens that each rank can hold after model weights and runtime buffers have been allocated. It is a memory ceiling rather than a direct guarantee of admissible batch size.
Reducing Weight Footprint with Humming MXFP4AFP8
Reduce the weight footprint first. Humming MXFP4AFP8 uses MXFP4 expert weights with online FP8 activations to reduce weight footprint and memory traffic on H20 GPUs, which lack native FP4 Tensor Cores. The SGLang integration is available in sglang#23754. We will cover the Humming/SGLang integration in a dedicated follow-up post. Model-level accuracy results and public reference measurements are provided in Appendix D.2.
Expanding KV Capacity with Online C128
Give the KV cache room to grow. The Offline C128 baseline retains per-index state for each compressed page. Online C128 instead maintains a compact aggregate state, releasing more HBM to the KV-cache pool. It introduces additional state maintenance and speculative-verification work, but we observed no TPOT regression in our tests.
Combined Capacity Gains
Figure 3. Capacity Scaling with Humming MXFP4AFP8 and Online C128.
Capacity gains compound across weights and KV state. By reducing the weight footprint, Humming MXFP4AFP8 expands full-token capacity to 1.71× the Baseline FP8 + Offline C128 configuration for DP32-EP32 and 4.47× for PP2-TP8. Online C128 then reduces the C128 auxiliary-state footprint, providing another 2.268× increase on top of Humming. Combined, the two techniques raise capacity to 3.88× the baseline for DP32-EP32 and 10.14× for PP2-TP8. Appendix D.1 provides the complete data.
2.3 Scenario-Specific Serving Profiles
Prefill Profiles
Figure 4. Prefill Profiles: Same Execution Path, Different Pipeline Depth.
The right pipeline depth depends on how much work there is to pipeline. PP2-CP8-TP8 and PP4-CP8-TP8 share the same Attention-CP8 → MoE-TP8 execution path. At the topology level, their primary difference is pipeline depth: PP2 distributes the model across two stages, while PP4 uses four.
Short contexts favor lower pipeline overhead; long contexts expose more parallelism. Short inputs produce fewer chunks, leaving a deeper pipeline underfilled and making fill, drain, and cross-stage transfer costs more prominent. Long contexts provide enough chunks to keep four stages busy; with fewer layers per stage, the additional nodes translate into more prefill parallelism. In our deployment, these characteristics led us to use PP2-CP8-TP8 for shorter contexts and PP4-CP8-TP8 for long-context workloads.
Low-Latency Decode Profiles
Figure 5. Low-Latency Decode: TP8 Reference and PP2-TP8 Serving Profile.
Low latency starts with the shortest execution path. Single-node TP8 and PP2-TP8 share the same Attention-TP8 → MoE-TP8 execution path; the difference is whether the model is partitioned across nodes. Single-node TP8 places all layers on one H20-141GB node and avoids cross-stage communication and synchronization. PP2-TP8 partitions the model across two pipeline stages.
The fastest topology is not always the most serviceable one. Single-node TP8 has the shorter execution path, but model weights and serving state share the HBM of one node, leaving limited room for the KV cache. It cannot simultaneously support long contexts and larger batch sizes. PP2-TP8 pays additional pipeline overhead but distributes the model weights across two nodes, releasing more HBM for KV state. For our latency and capacity targets, we use single-node TP8 as the batch-size-1 latency reference and PP2-TP8 as the low-latency serving profile.
High-Throughput Decode Profiles
Figure 6. High-Throughput Decode: DP16-EP16 Reference and DP32-EP32 Capacity Profile.
High-throughput decode scales data and expert parallelism together. Both profiles use the Attention-DP → MoE-EP execution path. DP16-EP16 is the smallest deployment unit; DP32-EP32 expands both DP and EP within the same topology.
Scale-out prioritizes request capacity over per-GPU throughput. A larger EP group distributes expert weights across more GPUs, releasing HBM for the KV cache and admitting more concurrent requests. At the same time, a smaller fraction of MoE traffic remains within each node, while a larger fraction crosses nodes, which can reduce per-GPU efficiency. In the profiles evaluated here, we use DP16-EP16 as the smallest deployment unit and efficiency reference, and DP32-EP32 to expand request capacity.
3. Prefill: Balancing Compute and Communication
Prefill performance is a system problem. Expert imbalance, context-parallel communication, and production routing shapes jointly determine TTFT; optimizing an isolated kernel is not enough.
3.1 Why MoE-TP Instead of MoE-EP
Figure 7. Replacing MoE-EP with MoE-TP.
Less traffic can still take longer. MoE-EP exchanges only routed tokens, but real prefill traffic exhibits significant expert skew. Ranks that own hot experts perform more computation and become stragglers; all other ranks wait for the slowest path at the combine step. Lower communication volume does not translate into lower TTFT.
Balance compute before minimizing traffic. For the H20 prefill workloads evaluated here, both PP2 and PP4 use MoE-TP. Full-sequence all-gather and reduce-scatter introduce more communication, but the traffic remains on high-bandwidth NVLink and has stable, predictable cost. All TP ranks execute tensor-parallel computation over the same routed tokens, preventing expert skew from becoming a rank-level long tail. For this workload, predictable communication is cheaper than unpredictable imbalance. The implementation is available in sglang#24947.
3.2 Accelerating and Fusing Prefill Collectives
Figure 8. Symmetric-Memory Collectives and Prefill Fusion.
Build a reusable collective fast path. MoE-TP replaces unpredictable expert imbalance with predictable collective traffic, making communication efficiency the next bottleneck. We made symmetric memory reusable across TP and CP, allowing AllReduce, AllGather, and ReduceScatter to share registered-buffer fast paths and applicable Hopper acceleration. The supporting upstream work spans memory-pool ownership, communicator registration, MoE-TP collective buffers, and the CP Attention and KV-cache buffer paths.
Then shorten the Prefill critical path. Faster collectives alone do not remove the boundaries between communication and computation. For the 32K single-chunk case, we built a fused path that overlaps a copy-engine-driven AllGather with fused FP8 quantization and shared-expert GEMM, then combines TopK reduction, shared-expert addition, and ReduceScatter in a second Triton kernel. This reorganizes seven operators into three execution groups and reduces TTFT by approximately 3.5% in a matched PP4 A/B.
3.3 Tuning Humming for Real Routing Shapes
Figure 9. Tuning Humming for Real Routing Shapes.
Generic tuning misses the shapes that matter. Prefill routing distributes tokens unevenly across 384 experts, so the effective M dimension clusters into a small set of discrete values. W13 and W2 also operate on different shapes, so a single generic heuristic cannot optimize both paths.
Tune from production routing. We extract high-frequency shapes from real routing histograms, build separate exact-shape configurations for W13 and W2, and validate them at the kernel, pipeline-stage, and matched A/B levels. The optimization target is not a synthetic range of M, but the routing distribution we actually serve. In a matched PP4 A/B at 32K, selected MoE kernel latency falls by approximately 21%, translating into an 11.35% end-to-end TTFT reduction.
4. Decode: Optimizing Speculation and MoE Execution
Decode optimization is profile-specific in our implementation. PP2-TP8 requires coordination across speculative pipeline stages, while DP32-EP32 focuses on optimizing the refinement step and expert routing at high concurrency. Humming fusion and overlap improve the shared MoE hot path beneath these serving topologies.
4.1 Low-Latency PP2-TP8: Extending DSpark Across Pipeline Stages
Figure 10. Coordinating DSpark Across PP2 Stages.
Pipeline parallelism splits the speculative loop. In PP2-TP8, target execution spans two pipeline stages, while the DSpark drafter resides only on the final stage. Stage 0 sends target hidden states to Stage 1, which performs verification, accepts tokens, and generates candidates for the next round.
Make two stages advance as one. Every speculative round crosses the pipeline boundary. We coordinate both stages and the required intermediate transfers under one execution protocol, preventing the stages from entering different rounds while avoiding redundant synchronization. The PP-specific DSpark integration is being upstreamed in sglang#32281.
4.2 High-Throughput DP32-EP32: Removing High-Concurrency Bottlenecks
Figure 11. DP32-EP32 Bottleneck Removal.
The matched A/B results in this subsection use DP32-EP32 at 4K with 32 concurrent requests per DP rank.
Choose the right execution shape for refinement. The refinement step applies a full-vocabulary projection to rescore DSpark's candidate set. At high concurrency, the row-wise dot-reduce repeatedly reads the vocabulary weights for every active row, creating a persistent tail in each decode step. We combine active rows into one transposed GEMM, reducing redundant memory traffic and shortening the refinement path. Per-GPU throughput improves by 22.8%.
Place experts from measured routing. DSpark traffic also exhibits significant expert skew. We record routing affinity from representative requests and use it to configure expert-parallel load balancing (EPLB) and redundant experts, preventing a small number of hot experts from repeatedly extending the critical path. Per-GPU throughput improves by 13.5%.
4.3 Humming Decode Hot Path: Fusion and Overlap
Figure 12. Humming Decode Hot-Path Optimizations.
These optimizations sit below the serving topology and can be reused by Humming-based decode profiles. The matched results below use DP32-EP32 at 4K with 32 concurrent requests per DP rank.
Remove the extra quantization pass. We fuse the SwiGLU activation with quantization so that the fused kernel directly produces the data and scale required by W2. This eliminates repeated access to an intermediate buffer and removes the standalone quantization pass, allowing W2 to start earlier. In the matched DSpark A/B, per-GPU throughput improves by 44.0%.
Overlap communication with W2. We adapt the Single-Batch Overlap (SBO) mechanism from our previous work (sglang#9660) into Humming-Aware SBO. Per-tile signals allow DeepEP to begin the corresponding combine send as soon as a W2 output tile completes, without waiting for the entire GEMM. In an earlier matched non-spec A/B at the same operating point, SBO recovers 4.12% throughput relative to the FP8-transport tier.
5. Evaluation: System Gains and Profile Trade-offs
5.1 Prefill: Cumulative Gains and Context-Length Trade-offs
Figure 13. Cumulative Prefill Throughput Gains.
PP2 strengthens the short-context profile. PP2 improves at all nine input lengths, with a geometric-mean throughput gain of 36.5% and a peak total input throughput of 16,900 tokens/s. Its shallower pipeline reduces fill-and-drain overhead for short requests, allowing PP2 to maintain lower TTFT with fewer resources.
PP4 carries the gains into long context. PP4 delivers a geometric-mean throughput gain of 31.8% across the same nine points. As context length grows, the deeper pipeline has enough work to amortize its fixed cost: total input throughput reaches 25,860 tokens/s at 512K and remains 23,970 tokens/s at 1M.
Figure 14. TTFT Trade-off Between PP2 and PP4.
Context length shifts the PP2/PP4 trade-off. Relative to PP4, PP2 lowers TTFT by 16.7% at 4K and 19.5% at 32K. The two profiles remain within 2% at 8K, 16K, and 64K. PP4 establishes a decisive advantage from 128K onward, reducing TTFT relative to PP2 by 26.2%, 33.3%, 42.1%, and 44.8% at 128K, 256K, 512K, and 1M, respectively. We therefore treat the routing boundary as an operating policy derived from the measured context-length range rather than a universal crossover point.
Appendix A.1–A.2 provide the complete TTFT and total-input-throughput results.
5.2 Low-Latency Decode: Performance and Capacity Trade-offs
Figure 15. Peak TPOT Gains from Optimized DSpark.
Optimized DSpark resets the latency baseline. Across the four input lengths shown in Figure 15, Optimized DSpark reduces peak TPOT by 74.8%–78.0% at batch size 1. At the largest batch size shared by each pair of measurements, the reduction remains 52.2%–60.0%. The gain holds from 8K through 1M rather than being confined to short contexts or single-request execution.
Figure 16. Batch-Size-1 Decode Throughput: H20-141GB and B300 Reference.
Observed serving performance is much closer than peak-compute ratios alone suggest. Across the four input lengths shown in Figure 16, Optimized DSpark on PP2-TP8 reaches 150–174 tokens/s at batch size 1. The single-node TP8 reference reaches 183–271 tokens/s. For the precisions used by the actual execution paths, B300 has approximately 45.6× the peak Tensor Core compute of H20-141GB (B300 FP4 versus H20 FP8) and 1.67× its memory bandwidth. Yet the highest observed generation rates are 383.7 tokens/s on B300 and 271 tokens/s on H20-141GB, respectively—a ratio of 1.42×. Even against this much stronger hardware reference, workload-specific optimization brings the H20-141GB reference substantially closer in observed serving performance.
Capacity favors PP2-TP8 for our production targets. Single-node TP8 is faster, but at a 1M context it has enough KV-cache capacity only for batch size 1. It cannot admit a larger batch or more concurrent requests. By distributing model weights across two pipeline stages, PP2-TP8 supports batch sizes 4, 8, and 16 at 1M, 512K, and 256K, respectively. With Online C128, its full-token capacity reaches 11.04M tokens/rank. For context-length and concurrency targets similar to ours, we recommend retaining single-node TP8 as the latency reference and using PP2-TP8 as the low-latency serving profile. Appendix B and Appendix D.1 provide the complete performance and capacity data.
5.3 High-Throughput Decode: Frontier Gains and Profile Trade-offs
Figure 17. Throughput–Interactivity Pareto Frontiers.
Figure 17 shows how the throughput–interactivity frontier evolves with the system. The horizontal axis is interactivity in tokens/s/user, and the vertical axis is throughput in tokens/s/GPU. In these DP/EP profiles, each DP rank maps to one GPU; interactivity is per-GPU throughput divided by the number of concurrent requests per DP rank. Points farther toward the upper right provide a better combination of user-visible generation speed and GPU efficiency. The four curves represent cumulative system evolution rather than the isolated gain of any optimization in Section 4.
MTP denotes multi-token prediction; the (3, 1, 4) configuration uses three speculative steps, top-k 1, and four draft tokens.
System optimization moves the entire frontier. At 4K with 32 concurrent requests per DP rank, per-GPU throughput rises from 319.92 tokens/s/GPU to 703.15 tokens/s/GPU, a 2.20× increase. At 1M with one request per DP rank, it rises from 27.05 tokens/s/GPU to 66.82 tokens/s/GPU. The first three system milestones can each process only one request per DP rank at 1M; the final system supports four and reaches 177.48 tokens/s/GPU. The expanded operating envelope comes from both faster execution and greater capacity.
Figure 18. Per-GPU Throughput: DP16-EP16 vs. DP32-EP32.
Smaller deployment units preserve efficiency at selected high-concurrency operating points. In our earlier work on serving DeepSeek-V3/R1 on H20, we found that a smaller EP deployment unit can keep a larger fraction of MoE traffic within each node. DeepSeek-V4-Pro shows the same advantage at the operating points plotted in Figure 18: with 16 and 32 concurrent requests per DP rank, DP16-EP16 delivers approximately 3.6%–20% higher per-GPU throughput than DP32-EP32. The full sweep is not monotonic across every concurrency level, so we use DP16-EP16 as an efficiency reference rather than a universal replacement for DP32-EP32.
Figure 19. Long-Context Request Capacity per DP Rank.
Capacity shifts the preferred high-throughput profile. DP16-EP16 is more efficient per GPU, but DP32-EP32 distributes expert weights across more ranks and releases additional HBM for the KV cache. At 256K, 512K, and 1M, the maximum concurrent requests per DP rank increase from 8, 4, and 2 to 16, 8, and 4, respectively—a consistent 2× expansion. For deployments with long-context concurrency targets similar to ours, this additional capacity favors DP32-EP32 as the capacity-oriented high-throughput profile, while DP16-EP16 remains useful as an efficiency reference. Appendix C and Appendix D.1 provide the complete data.
6. Conclusion
One model does not require one compromise profile. We built a scenario-specific serving stack for DeepSeek-V4-Pro on H20. Prefill switches between PP2 and PP4 according to context length. Decode uses PP2-TP8 for low latency and DP32-EP32 for high throughput. By co-designing capacity, deployment topology, and execution path, H20 can sustain 1M-token contexts and meet multiple serving SLOs despite limited compute and the absence of native FP4 Tensor Cores.
The transferable result is a scenario-driven methodology. Serving profiles should not be selected from hardware specifications or isolated benchmarks alone. We recommend starting from the workload, SLO, context length, and concurrency, then using profiling to identify the binding resource and translate it into concrete topology and execution-path decisions. We hope this methodology helps AI infrastructure teams build practical frontier-model serving systems under diverse resource constraints—whether the bottleneck is compute, memory capacity, memory bandwidth, or interconnect—and share those lessons with the broader open-source ecosystem.
Acknowledgements
We would like to thank the SGLang Team and Community for their outstanding work on the SGLang framework. We also thank the following teams and collaborators for their support and contributions:
- Ant Group SCT Team: Yongfei Xu, Qianyu Zhang, Zekai Gu, ZhiLin Huang, Fakang Wang, Jianhao Fu, Zhuoxuan Du, Xia Zhan, Chun Huang, Qi Liu, Xi Chen, Yuhan Mao, Peipeng Cheng, Hanlin Gao, Jinghua Yao
- Ant Group Venus Team: Jinzhen Lin
- SGLang Community: Peng Zhang
Appendix A. Prefill Results
A.1 Humming PP2 Prefill: Baseline vs. Final Profile
| Input Length | Baseline TTFT (ms) | Baseline Total Input Throughput (tokens/s) | Final TTFT (ms) | Final Total Input Throughput (tokens/s) |
|---|---|---|---|---|
| 4K | 775.8 | 5,280 | 573.3 | 7,140 |
| 8K | 1202.1 | 6,810 | 907.6 | 9,030 |
| 16K | 2059.8 | 7,950 | 1649.5 | 9,930 |
| 32K | 4137.5 | 7,920 | 2470.3 | 13,260 |
| 64K | 6195.7 | 10,580 | 4063.8 | 16,130 |
| 128K | 10744.4 | 12,200 | 7975.9 | 16,430 |
| 256K | 20542.2 | 12,760 | 15507.2 | 16,900 |
| 512K | 44544.6 | 11,770 | 34982.6 | 14,990 |
| 1M | 100304.2 | 10,450 | 79214.2 | 13,240 |
A.2 Humming PP4 Prefill: Baseline vs. Final Profile
| Input Length | Baseline TTFT (ms) | Baseline Total Input Throughput (tokens/s) | Final TTFT (ms) | Final Total Input Throughput (tokens/s) |
|---|---|---|---|---|
| 4K | 924.6 | 4,430 | 687.9 | 5,950 |
| 8K | 1174.5 | 6,970 | 890.3 | 9,200 |
| 16K | 2202.0 | 7,440 | 1635.4 | 10,020 |
| 32K | 4185.6 | 7,830 | 3068.4 | 10,680 |
| 64K | 5252.4 | 12,480 | 3982.6 | 16,460 |
| 128K | 7793.4 | 16,820 | 5882.5 | 22,280 |
| 256K | 13210.7 | 19,840 | 10348.9 | 25,330 |
| 512K | 26350.1 | 19,900 | 20273.1 | 25,860 |
| 1M | 55532.3 | 18,880 | 43742.5 | 23,970 |
Appendix B. Low-Latency Decode Results
B.1 Peak TPOT Across Input Lengths and Batch Sizes
B.1.1 No-Spec PP2-TP8
| Input Length / Batch Size (Peak TPOT, ms) | 1 | 2 | 4 | 8 | 16 |
|---|---|---|---|---|---|
| 8K | 26.39 | 30.86 | 31.31 | 31.79 | 31.74 |
| 32K | 25.72 | 26.58 | 27.81 | 31.06 | 37.97 |
| 64K | 25.75 | 26.62 | 28.13 | 29.19 | 38.75 |
| 128K | 25.94 | 26.94 | 28.38 | 29.75 | 38.51 |
| 256K | 26.08 | 27.21 | 28.84 | 32.43 | 38.83 |
| 512K | 26.25 | 27.51 | 29.16 | 33.70 | - |
| 1M | 26.42 | 27.81 | 29.52 | - | - |
B.1.2 Optimized DSpark PP2-TP8
| Input Length / Batch Size (Peak TPOT, ms) | 1 | 2 | 4 | 8 | 16 | 32 |
|---|---|---|---|---|---|---|
| 4K | 5.91 | 6.76 | 7.97 | 10.00 | 14.55 | 19.23 |
| 8K | 5.80 | 6.87 | 8.85 | 10.48 | 15.18 | 19.60 |
| 32K | 6.14 | 7.04 | 8.39 | 10.83 | 14.86 | 20.46 |
| 64K | 6.15 | 7.13 | 8.73 | 10.39 | 15.49 | 21.65 |
| 128K | 6.77 | 7.02 | 8.91 | 11.59 | 16.17 | 24.78 |
| 256K | 5.76 | 6.98 | 8.61 | 11.98 | 17.72 | - |
| 512K | 6.35 | 7.95 | 9.87 | 14.30 | - | - |
| 1M | 6.65 | 8.92 | 12.43 | - | - | - |
B.2 Batch-Size-1 Output Throughput
| Input Length | No-Spec PP2-TP8 (tokens/s) | Optimized DSpark PP2-TP8 (tokens/s) | Single-Node TP8 (tokens/s) |
|---|---|---|---|
| 4K | - | 169 | 213 |
| 8K | 38 | 172 | 260 |
| 16K | - | - | 244 |
| 32K | 39 | 163 | 269 |
| 64K | 39 | 163 | 246 |
| 128K | 39 | 148 | 267 |
| 256K | 38 | 174 | 271 |
| 512K | 38 | 157 | 254 |
| 1M | 38 | 150 | 183 |
B.3 Benchmark Settings
Hardware: one 8× H20-141GB decode node.
Decode server
--tp-size 8 \
--mem-fraction-static 0.91 \
--max-running-requests 1 \
--cuda-graph-max-bs 1 \
--cuda-graph-bs 1 \
--moe-runner-backend humming \
--moe-a2a-backend none \
--speculative-algorithm DSPARK \
--speculative-num-draft-tokens 7 \
--speculative-dspark-block-size 7 \
--speculative-moe-runner-backend triton \
--speculative-moe-a2a-backend none
Client benchmark
python3 -m sglang.bench_serving \
--backend sglang-oai-chat \
--host 127.0.0.1 \
--port 8000 \
--model <MODEL_PATH> \
--dataset-name random \
--dataset-path <DATASET_PATH> \
--random-input-len 262144 \
--random-output-len 4096 \
--random-range-ratio 1.0 \
--num-prompts 10 \
--max-concurrency 1 \
--warmup-requests 0 \
--seed 1
Output throughput is extracted from the server's TP0 Decode batch log lines; we discard the highest and lowest 20% of samples and average the remainder. The B300 number follows the setup reported in the linked source.
Appendix C. High-Throughput Decode Results
C.1 DP32-EP32 with FP8 + MTP (3, 1, 4)
| Input Length / Concurrent Requests per DP Rank (tokens/s/GPU) | 1 | 2 | 4 | 8 | 16 | 32 |
|---|---|---|---|---|---|---|
| 4K | 30.49 | 58.58 | 102.89 | 174.75 | 253.15 | 319.92 |
| 8K | 30.34 | 58.29 | 102.38 | 174.67 | 251.62 | 318.32 |
| 16K | 29.70 | 56.55 | 99.47 | 170.01 | 242.22 | 302.43 |
| 32K | 29.58 | 56.35 | 98.28 | 164.26 | 234.13 | - |
| 64K | 29.07 | 55.73 | 96.43 | 161.60 | - | - |
| 128K | 28.39 | 54.06 | 92.89 | 153.55 | - | - |
| 256K | 28.35 | 53.02 | 90.89 | - | - | - |
| 512K | 27.51 | 51.49 | - | - | - | - |
| 1M | 27.05 | - | - | - | - | - |
C.2 DP32-EP32 with FP8 + Optimized MTP (3, 1, 4)
| Input Length / Concurrent Requests per DP Rank (tokens/s/GPU) | 1 | 2 | 4 | 8 | 16 | 32 |
|---|---|---|---|---|---|---|
| 4K | 36.84 | 69.86 | 131.96 | 232.94 | 389.94 | 514.77 |
| 8K | 32.58 | 69.51 | 131.53 | 222.06 | 348.80 | 416.82 |
| 16K | 31.89 | 67.44 | 127.79 | 216.14 | 341.85 | 395.99 |
| 32K | 31.49 | 67.21 | 124.49 | 208.83 | 337.97 | - |
| 64K | 30.95 | 66.47 | 123.68 | 205.44 | - | - |
| 128K | 30.22 | 64.47 | 119.14 | - | - | - |
| 256K | 30.18 | 63.23 | - | - | - | - |
| 512K | 29.28 | 61.40 | - | - | - | - |
| 1M | 28.79 | - | - | - | - | - |
C.3 DP32-EP32 with FP8 + DSpark
| Input Length / Concurrent Requests per DP Rank (tokens/s/GPU) | 1 | 2 | 4 | 8 | 16 | 32 |
|---|---|---|---|---|---|---|
| 4K | 53.1 | 94.8 | 181.2 | 338.1 | 495.8 | 591.8 |
| 8K | 44.5 | 88.4 | 170.1 | 317.3 | 495.5 | - |
| 16K | 43.6 | 88.3 | 165.3 | 308.8 | 455.5 | - |
| 32K | 43.0 | 87.3 | 161.0 | 298.4 | - | - |
| 64K | 42.3 | 86.3 | 158.0 | - | - | - |
| 128K | 41.3 | 83.8 | - | - | - | - |
| 256K | 41.2 | - | - | - | - | - |
| 512K | 40.0 | - | - | - | - | - |
| 1M | 39.3 | - | - | - | - | - |
C.4 DP32-EP32 with Humming MXFP4AFP8 + Online C128 + DSpark
| Input Length / Concurrent Requests per DP Rank (tokens/s/GPU) | 1 | 2 | 4 | 8 | 16 | 32 |
|---|---|---|---|---|---|---|
| 4K | 75.32 | 127.10 | 235.85 | 417.53 | 564.08 | 703.15 |
| 8K | 75.60 | 128.29 | 238.01 | 417.34 | 560.68 | 709.64 |
| 16K | 74.00 | 124.47 | 231.25 | 406.21 | 539.72 | 674.19 |
| 32K | 73.07 | 122.27 | 225.28 | 392.47 | 521.70 | 601.67 |
| 64K | 71.81 | 120.92 | 221.05 | 386.11 | 516.54 | 599.63 |
| 128K | 70.12 | 117.29 | 212.93 | 366.88 | 487.69 | - |
| 256K | 70.03 | 115.03 | 208.35 | 345.21 | 457.62 | - |
| 512K | 67.95 | 111.71 | 191.99 | 302.80 | - | - |
| 1M | 66.82 | 105.82 | 177.48 | - | - | - |
C.5 DP16-EP16 with Humming MXFP4AFP8 + Online C128 + DSpark
| Input Length / Concurrent Requests per DP Rank (tokens/s/GPU) | 1 | 2 | 4 | 8 | 16 | 32 |
|---|---|---|---|---|---|---|
| 4K | 76.80 | 129.62 | 236.83 | 397.42 | 584.37 | 759.73 |
| 8K | 76.69 | 130.55 | 237.53 | 398.60 | 582.03 | 762.09 |
| 16K | 76.16 | 127.88 | 233.10 | 388.54 | 571.22 | 745.23 |
| 32K | 74.07 | 124.77 | 226.24 | 378.79 | 559.05 | 722.51 |
| 64K | 74.57 | 124.69 | 223.84 | 373.13 | 541.46 | 695.35 |
| 128K | 72.36 | 120.34 | 219.66 | 365.13 | 518.98 | - |
| 256K | 71.38 | 119.19 | 211.64 | 340.72 | - | - |
| 512K | 69.54 | 115.14 | 198.81 | - | - | - |
| 1M | 67.39 | 106.50 | - | - | - | - |
Appendix D. Capacity Results
D.1 Decode Capacity Scaling
| Decode Profile | Configuration | Full-Token Capacity (tokens/rank) | Vs. Previous Stage | Vs. FP8 Baseline |
|---|---|---|---|---|
| DP32-EP32 | Baseline FP8 + Offline C128 | 1,475,328 | - | 1.00× |
| Humming MXFP4AFP8 + Offline C128 | 2,526,720 | 1.71× | 1.71× | |
| Humming MXFP4AFP8 + Online C128 | 5,731,328 | 2.268× | 3.88× | |
| PP2-TP8 | Baseline FP8 + Offline C128 | 1,089,024 | - | 1.00× |
| Humming MXFP4AFP8 + Offline C128 | 4,869,888 | 4.47× | 4.47× | |
| Humming MXFP4AFP8 + Online C128 | 11,044,906 | 2.268× | 10.14× |
D.2 Humming Accuracy Validation
We evaluated the DP16-EP16 Humming MXFP4AFP8 + Online C128 + DSpark profile on GSM8K1000. The profile achieves 95.5% exact-match accuracy, with one invalid response and zero system errors, passing our 95.0% acceptance threshold.
As a public reference, the upstream SGLang Humming integration reports the following results on a 200-example GSM8K evaluation of DeepSeek-V4-Flash:
| Backend | GSM8K Accuracy |
|---|---|
| Marlin MXFP4A16 | 96.5%–97.0% |
| FlashInfer MXFP4 | 96.5%–97.0% |
| Humming MXFP4A16 | 96.5%–97.5% |
| Humming MXFP4AFP8 | 97.0% |
No accuracy degradation is visible for Humming MXFP4AFP8 in this public comparison. Because it uses DeepSeek-V4-Flash rather than DeepSeek-V4-Pro, we treat it as an external reference rather than a matched accuracy-loss measurement for our serving profile.