Nvidia’s Rubin Ultra GPU is moving forward with HBM4E memory, not the HBM4 used in its predecessor, and it is bringing a meaningful capacity upgrade along for the ride. The chip is targeting up to 768GB of HBM4E using 12-Hi stacks, a substantial leap from the 288GB of HBM4 that the current Rubin GPUs support.
The Kyber rack-scale platform, which includes the NVL144 and NVL576 systems, remains active and on schedule alongside Rubin Ultra for a second-half 2027 release.
What changed, and why it matters
The original Rubin Ultra design called for 1TB of memory using 16-Hi HBM4E stacks. That target has since been revised downward to 768GB via 12-Hi stacks, a direct response to anticipated supply constraints across the HBM4E production chain.
The company sources HBM from SK Hynix, Samsung, and Micron, and production of next-generation HBM4E is expected to be tight. Nvidia is also evaluating lower-capacity 8-Hi HBM4E configurations as a fallback, giving itself room to maneuver if procurement gets difficult.
Nvidia pushes back on delay reports
Reports had circulated suggesting Rubin Ultra could slip into 2028, which would have been a meaningful setback in Nvidia’s annual cadence of GPU releases. Nvidia has publicly denied those claims, reaffirming the second-half 2027 target.
The company’s roadmap has Vera Rubin GPUs arriving in 2026, followed by Rubin Ultra in 2027. The Rubin Ultra architecture also incorporates a multi-chiplet design intended to increase GPU density per rack. The Kyber NVL576, as the name implies, scales to 576 GPUs in a single rack configuration.
Supply chain flexibility as competitive strategy
SK Hynix is widely regarded as the leading supplier of HBM at current node generations, and its production capacity for HBM4E will be a gating factor for the entire industry in 2026 and 2027. Samsung and Micron are ramping their own HBM4E programs, but both have faced qualification delays with major customers. Nvidia’s multi-supplier approach reduces single-source dependency, though it does not eliminate the category risk entirely.
For context on why memory architecture matters this much: in large language model training, the GPU’s memory bandwidth and capacity determine how large a model can fit on a single chip, which in turn determines how efficiently the system can train without constantly offloading data between chips. More memory per GPU means fewer inter-chip data transfers, which means faster training at lower energy cost. The jump from 288GB to 768GB is not a spec sheet vanity metric.
Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy.

1 day ago
88








English (US) ·