Nvidia Corp has reportedly brought its Rubin CPX program back onto its roadmap after the market had largely assumed the product was dropped. Industry checks conducted by Ming-Chi Kuo of TF International Securities show production for the revised CPX is set to begin in the first quarter of 2027.
According to Kuo, the updated Rubin CPX departs from prior CPX iterations in several technical areas. Each CPX GPU in the new design offers compute performance comparable to the Rubin GPU and carries a peak power rating of 2,300 watts. Memory is configured as 168 GB of HBM4 per CPX unit, while the Rubin GPU supports 288 GB of HBM4 and the earlier CPX design used 128 GB of GDDR7.
The physical and system architecture has also been revised. Rather than sharing rack space with Rubin, the new CPX units will be deployed in a separate MGX ETL rack. Customers will be able to choose systems populated with 64, 128, 192, or 256 CPX GPUs. The design organizes every 64 GPUs into a module composed of eight compute trays - each tray containing eight GPUs - plus a single switch tray per module.
For intra-tray scaling, NVLink will interconnect eight CPX GPUs within a tray, delivering between 1 and 1.5 TB per second of NVLink bandwidth per CPX. By comparison, each Rubin GPU offers 3.6 TB per second of NVLink bandwidth. Tray-to-tray communications will rely on Spectrum-6 Ethernet using copper links, while cross-module connections will use OSFP optical links.
Kuo emphasized an operational dependency between the CPX and Rubin systems: Nvidia requires CPX to be paired with the Vera Rubin NVL72 at a 1:1 ratio. In that workflow, CPX units perform prefill operations and construct the KV cache. The KV cache is then transferred to Rubin systems via Ethernet RDMA for decode processing.
On workload composition, Kuo noted that more than half of current AI inference workloads consist of processing input context and building KV caches. In this design context, each eight-CPX tray contains roughly 1.34 TB of HBM4 memory.
Context and implications
The revived Rubin CPX program reflects a shift in Nvidia's system-level approach, with distinct rack segmentation and a paired deployment model alongside Rubin NVL72. The revised memory and interconnect choices indicate trade-offs between capacity, bandwidth and rack-level topology that will shape how customers deploy these systems for inference workloads.
More detailed product performance and customer adoption timelines beyond the stated 1Q27 production start were not provided in Kuo's checks.