Nvidia reportedly restarts "Rubin CPX" AI accelerator project and switches to 168GB HBM4 memory

📅 2026-09-01

Abstract:

According to supply chain analyst Ming-Chi Kuo, NVIDIA’s previously shelved “Rubin CPX” AI accelerator project has now been restarted and the memory architecture has been redesigned. This product was originally planned to be equipped with 128GB GDDR7 video memory, but now it is equipped with 168GB HBM4 high-bandwidth memory.

Nvidia announced at the end of last year that it planned to develop a dedicated accelerator based on the Rubin GPU family, mainly for large-scale agent artificial intelligence workloads. After the project was announced, it was reported that it was temporarily shelved, but Nvidia seems to have completed re-planning. Rubin CPX will be deployed in independent racks, each rack can be configured with 64 to 256 independent CPX GPUs.

Different from conventional Rubin GPUs responsible for decoding tasks, CPX GPUs are mainly used for the "pre-filling" stage in the inference process of large language models and are responsible for processing input context and KV cache. NVIDIA is trying to separate the pre-filling and decoding tasks, with the CPX GPU taking care of the pre-filling work, and then the regular Rubin GPU taking care of the decoding, thereby improving the overall inference efficiency.

A rack tray equipped with 8 CPX GPUs can handle up to 1.34TB of long context pre-population data and related KV cache, all relying on HBM4 memory. HBM4 can provide higher memory bandwidth, helping to speed up the processing of long context data. Since the previous design did not require integrated HBM4, Nvidia must also redesign the chip packaging, and it is expected to use TSMC's CoWoS-S or CoWoS-L advanced packaging technology.

In terms of computing performance, Rubin CPX adopts a monolithic chip design and can provide 30 gigaflops of NVFP4 computing performance, or 30 PFLOPS. The chip also integrates 4 NVENC video encoding engines and 4 NVDEC video decoding engines, allowing it to complete more efficient multimedia-related workloads without relying on external processors.

As the parameter scale of large language models approaches the trillions level, handing over pre-filling tasks and decoding tasks to different racks is expected to significantly increase the speed of token generation and delivery. Nvidia also recommends maintaining a 1:1 ratio of CPX GPUs to regular Rubin GPUs in the decoding rack.

Related tags

Related articles

Comments

0/500
Captcha (click to refresh)
No comments yet