SemiAnalysis · @SemiAnalysis_ · X·2026-09-01 01:02·1天前
AI 导读

NVIDIA LPU 支持 3 种分离式推理: 1. Rubin Prefill + LPU Decode,实现最快交互性 2. Rubin Prefill + Rubin Decode Attention + LPU Decode FFN,面向曲线中段 3. Rubin Prefill + Rubin Decode Verification + LPU Drafter,面向曲线中左段 在低交互性场景下,纯 Rubin 仍占优势。期待 Rubin + LPU 在 AgentX 等开源智能体基准上的性能曲线表现。

SemiAnalysis@SemiAnalysis_
41AI 编辑部评分,满分 100
2026-09-01 01:02· 1天前
AI 导读

NVIDIA LPU 支持 3 种分离式推理: 1. Rubin Prefill + LPU Decode,实现最快交互性 2. Rubin Prefill + Rubin Decode Attention + LPU Decode FFN,面向曲线中段 3. Rubin Prefill + Rubin Decode Verification + LPU Drafter,面向曲线中左段 在低交互性场景下,纯 Rubin 仍占优势。期待 Rubin + LPU 在 AgentX 等开源智能体基准上的性能曲线表现。

NVIDIA LPU supports 3 types of disaggregated inferencing:

  1. Rubin Prefill + LPU Decode for the fastest interactivity
  2. Rubin Prefill + Rubin Decode Attention + LPU Decode FFN for the middle of the curve
  3. Rubin Prefill + Rubin Decode Verification + LPU Drafter for the middle-left of the curve

For low interactivity, raw Rubin still takes the win. Looking forward to seeing Rubin + LPU performance curves on open-source agentic benchmarks like AgentX.