Just as the market had come to believe that Rubin CPX had been dropped from Nvidia’s product roadmap, my latest industry checks indicate that Nvidia has revived the program, with production expected to begin in 1Q27. Compared with the previous design, the revived Rubin CPX delivers stronger prefill performance and features major changes to both its GPU specifications and rack architecture, underscoring the high priority Nvidia places on prefill solutions.
Key changes:
1. Rubin CPX GPU specs CPX delivers near-Rubin compute performance and matches Rubin’s maximum power rating of 2,300 W per GPU. CPX moves to 168 GB of HBM4, vs. 288 GB on Rubin and 128 GB of GDDR7 on the previous CPX design.
2. Rack design The new CPX uses a standalone MGX ETL rack rather than sharing a rack with Rubin, as in the previous design. Customers can opt for 64, 128, 192, or 256 CPX GPUs depending on their needs. Within a CPX rack, each group of 64 CPX GPUs forms a rack module comprising eight compute trays (eight CPX GPUs per tray) and one switch tray.
3. Scale-up and scale-out NVLink is used only for scale-up among the eight CPX GPUs within each tray, with 1–1.5 TB/s of NVLink bandwidth per CPX (vs. 3.6 TB/s per Rubin). Inter-tray scale-out within each rack module runs over Spectrum-6 Ethernet using all-copper L1 links. Across rack modules, scale-out is handled by each module’s Spectrum-6 switch over OSFP optical links.
4. How it works CPX must be paired with Vera Rubin NVL72, and Nvidia recommends a 1:1 ratio of CPX to Rubin GPUs. CPX handles prefill and builds the KV cache, which is then transferred to Rubin over Ethernet RDMA for decode.
5. Product positioning: Best performance per dollar for long-context prefill Over 50% of today’s AI inference workload comes from processing input context and building the corresponding KV cache. CPX therefore offers a more flexible, lower-cost way to handle prefill. Each eight-CPX tray has approximately 1.34 TB of HBM4, sufficient for most long-context prefill workloads and associated KV cache requirements.