Nvidia's upcoming Vera Rubin platform, set to arrive later this year, will take the stage as the AI world shifts towards an era dominated not by frontier training runs but by the demands of agentic AI inference at massive scale. The hunger for generated tokens in agentic workflows and the demands of delivering them quickly, efficiently, and at low unit cost now dominate the discussion.
We’ve already gone in depth on new performance data around the Vera CPU and how it helps to accelerate agentic AI workloads, but that’s not all Nvidia is sharing today. It’s also detailing some new features of the Rubin architecture and how those features are meant to increase inference efficiency from the GPU level to rack-scale and data-center-scale implementations of this accelerator platform.
(Image credit: Nvidia)
The full Vera Rubin NVL72 rack-scale system is built up from 36 Vera CPUs and 72 Rubin GPUs, but our focus today is on the GPU proper. Rubin joins two compute dies onto a single package using the Nvidia High Bandwidth Interface. The resulting chip offers 224 Streaming Multiprocessors (SMs) containing a total of 896 Tensor Cores alongside 288GB of HBM4 memory providing 22 TB/s of memory bandwidth.
Latest Videos From Watch full video here:
As an inference-focused accelerator, Nvidia touts Rubin’s 50 sparse PFLOPS of NVFP4 inference throughput as its headline performance figure, although that’s only one of a dizzying array of data types this chip can handle. Here are some key rates to keep in mind for this chip so far:
Swipe to scroll horizontally Nvidia Rubin GPU Row 0 - Cell 1 NVFP4 Inference 50 PFLOPS (with sparsity) NVFP4 Training 35 PFLOPS FP8/FP6 Training 17.5 PFLOPS INT8 250 TOPS FP16/BF16 4 PFLOPS TF32 2 PFLOPS FP32 130 TFLOPS FP64 33 TFLOPS
Let’s dive into some of Rubin’s refinements for inference workloads to understand how Nvidia aims to keep all of those resources fully utilized.
The Rubin Tensor Memory Accelerator efficiently manages growing MoE models
First up, Nvidia highlights efficiency improvements in the Tensor Memory Accelerator (TMA) that help feed the Tensor Cores with data. The TMA is a dedicated engine built to handle memory address calculations and perform direct loads of array data into a GPU's shared local memory.
... continue reading