# 英伟达复活 Rubin CPX，2027 年一季度量产

- 来源：郭明錤｜Ming-Chi Kuo (@mingchikuo)
- 发布时间：2026-08-31 22:04
- AIHOT 分数：51
- AIHOT 链接：https://aihot.virxact.com/items/cmthcqzyw0a72rodm1uaict8g
- 原文链接：https://x.com/mingchikuo/status/2094426162193993964

## AI 摘要

英伟达已重启此前被传砍掉的 Rubin CPX 项目，预计 1Q27 量产。新版 CPX 预填充性能更强，GPU 改用 168GB HBM4（Rubin 为 288GB），采用独立 MGX ETL 机架，需与 Vera Rubin NVL72 以 1:1 配对，通过以太网 RDMA 传输 KV 缓存，主打长上下文预填充的每美元最佳性能。

## 正文

Just as the market had come to believe that Rubin CPX had been dropped from Nvidia’s product roadmap, my latest industry checks indicate that Nvidia has revived the program, with production expected to begin in 1Q27. Compared with the previous design, the revived Rubin CPX delivers stronger prefill performance and features major changes to both its GPU specifications and rack architecture, underscoring the high priority Nvidia places on prefill solutions.

Key changes:

1. Rubin CPX GPU specs
CPX delivers near-Rubin compute performance and matches Rubin’s maximum power rating of 2,300 W per GPU. CPX moves to 168 GB of HBM4, vs. 288 GB on Rubin and 128 GB of GDDR7 on the previous CPX design.

2. Rack design
The new CPX uses a standalone MGX ETL rack rather than sharing a rack with Rubin, as in the previous design. Customers can opt for 64, 128, 192, or 256 CPX GPUs depending on their needs. Within a CPX rack, each group of 64 CPX GPUs forms a rack module comprising eight compute trays (eight CPX GPUs per tray) and one switch tray.

3. Scale-up and scale-out
NVLink is used only for scale-up among the eight CPX GPUs within each tray, with 1–1.5 TB/s of NVLink bandwidth per CPX (vs. 3.6 TB/s per Rubin). Inter-tray scale-out within each rack module runs over Spectrum-6 Ethernet using all-copper L1 links. Across rack modules, scale-out is handled by each module’s Spectrum-6 switch over OSFP optical links.

4. How it works
CPX must be paired with Vera Rubin NVL72, and Nvidia recommends a 1:1 ratio of CPX to Rubin GPUs. CPX handles prefill and builds the KV cache, which is then transferred to Rubin over Ethernet RDMA for decode.

5. Product positioning: Best performance per dollar for long-context prefill
Over 50% of today’s AI inference workload comes from processing input context and building the corresponding KV cache. CPX therefore offers a more flexible, lower-cost way to handle prefill. Each eight-CPX tray has approximately 1.34 TB of HBM4, sufficient for most long-context prefill workloads and associated KV cache requirements.
