NVIDIA LPU supports 3 types of disaggregated inferencing:
- Rubin Prefill + LPU Decode for the fastest interactivity
- Rubin Prefill + Rubin Decode Attention + LPU Decode FFN for the middle of the curve
- Rubin Prefill + Rubin Decode Verification + LPU Drafter for the middle-left of the curve
For low interactivity, raw Rubin still takes the win. Looking forward to seeing Rubin + LPU performance curves on open-source agentic benchmarks like AgentX.