Skip to content
Tech News
← Back to articles

Nvidia is challenging 20 years of datacenter CPU design with its new Vera chip

read original more articles

The big picture: Nvidia lifted the embargo on its Vera CPU deep dive this week, and the disclosure amounts to a direct challenge to two decades of x86 datacenter design philosophy. This is the most detail the company has shared on the chip since it first appeared on the Rubin roadmap, and it confirms something I have suspected for a while: Nvidia is not treating the CPU as an attach story anymore. It is treating it as a battleground with a lot of potential dollars at play.

Vera is built around the Olympus core, the first custom CPU core Nvidia has ever brought to the datacenter and the first custom core the company has designed anywhere since the Denver and Carmel efforts of the Tegra era nearly a decade ago.

Ryan Shrout is a longtime technology analyst and industry veteran who has spent over two decades covering PC hardware, graphics, and semiconductors. He previously led technical marketing at Intel and was the founding editor of PC Perspective. He is currently President and GM at Signal65. You can follow him on X @ryanshrout.

Grace used licensed off-the-shelf Arm Neoverse V2 cores. Olympus is an Nvidia design from the ground up, a wide, high-IPC core with a 10-wide decode front end that reorders aggressively and prefetches based on patterns like graph structures in memory.

88 of those cores sit on a monolithic compute die, running 176 threads through a partitioned scheme Nvidia calls Spatial Multithreading, a deliberate departure from the opportunistic resource sharing of traditional SMT.

That monolithic choice matters. Nvidia still uses chiplets for the memory controllers and I/O, but the compute die is one piece of silicon connected by a second-generation scalable coherency fabric. The company measures bisection bandwidth across the die at roughly 3.4 terabytes per second.

The memory subsystem is LPDDR5X hardened for the datacenter with ECC and full telemetry, delivering up to 1.2 TB/s of bandwidth, roughly 3x the memory bandwidth per core and about 5x the bandwidth per watt of conventional DDR-based server designs (all based on Nvidia claims).

The headline claims stack up as roughly 2x faster performance from the Olympus core, 3x the core-to-core bandwidth of chiplet-based competition, and 40% lower memory latency under load through the LPDDR5X subsystem.

Vera ships in two forms: a dense liquid-cooled rack packing 256 CPUs and more than 22,000 cores, and a conventional air-cooled 2U with two sockets. Dell has committed to multiple PowerEdge systems built on it. Nvidia sizes the opportunity as a $200 billion expansion of the CPU market, which explains a lot of the recent market dynamics.

The argument behind the architecture

... continue reading