NVIDIA Revives Rubin CPX Chip for AI Inference, Mass Production Set for 2027

NVIDIA restarts Rubin CPX chip plan for AI inference with 168GB HBM4. Mass production expected in Q1 2027 for long- context processing tasks.

NVIDIA Rubin CPX
NVIDIA Rubin CPX

has restarted the Rubin CPX chip plan, a move that changes how data centers will handle the heavy lifting of AI inference. This shift matters because the chip targets the prefill stage, which currently consumes more than half of all AI inference workloads. Buyers looking to optimize long-context processing now have a dedicated hardware path rather than relying solely on general-purpose accelerators.

New chip targets long-context inference with 168GB HBM4 memory

The Rubin CPX serves as a specialized complement to the standard Rubin architecture, specifically designed for million-token prefill and long-context inference tasks. According to analyst Guo Mingchi, the new version features major adjustments in specifications, rack design, and interconnect architecture. The product positioning shifts toward greater flexibility and cost-effectiveness compared to previous iterations.

Rubin CPX Specifications

  • Memory: 168GB HBM4
  • Compute Performance: 30-50 Petaflops
  • Precision: NVFP4
  • Power Consumption: 2300W
  • NVLink Bandwidth: 1-1.5TB/s

Technical specifications for the restarted chip include 168GB of HBM4 memory, an increase from the earlier 128GB GDDR7 plan. Each GPU delivers 30 to 50 Petaflops of compute performance using NVFP4 precision. Power consumption reaches up to 2300W per chip, while NVLink bandwidth provides 1 to 1.5TB/s for scale-up operations within a single tray.

Rack configurations support 64, 128, 192, or 256 units per rack, with each 64-unit module consisting of eight compute trays and one switch tray. Interconnects use NVLink for scale-up, Spectrum-6 Ethernet for scale-out, and OSFP fiber for cross-rack connections. We looked at Vera Rubin NVL72 earlier while tracking NVIDIA's broader AI hardware roadmap.

Mass production is expected in the first quarter of 2027, according to the latest industry survey from Tianfeng International Securities. The chip will handle input processing and KV Cache creation, offering a cost-effective alternative to standard Rubin for these specific tasks. Exact final specifications may change before production begins, as the details come from an analyst report rather than official NVIDIA confirmation.

Discussion

0 comments

Log in to join the thread with a thoughtful take, question, or correction.

Add to the discussion