Skip to content
Tech News
← Back to articles

Nvidia says Groq racks will be online this year following $20 billion purchase

read original more articles
Why This Matters

Nvidia's launch of the Groq 3 LPX rack signifies a major step in advancing low-latency AI inference, crucial for delivering faster, more responsive AI services. This development underscores Nvidia's strategic push into specialized AI hardware, which can enable premium cloud services and open new revenue streams. The move also highlights the competitive landscape in AI chip technology, with other players like AMD and Cerebras also focusing on low-latency solutions.

Key Takeaways

The Nvidia Groq 3 LPU chip during the Nvidia GTC conference in San Jose, California, March 18, 2026.

Nvidia announced on Monday that its Groq 3 LPX rack is in full production, marking the commercialization of technology from the company's largest acquisition on record.

The Groq rack will be deployed alongside Vera central processors and Rubin graphics processors at neocloud Nebius , and will be online later this year, Nvidia senior director Dion Harris told reporters.

Nvidia's race to manufacture Groq's chip and make it available to customers highlights the growing importance of low-latency inference that's needed to make AI agents feel responsive without long lags for users, especially for coding. Cloud companies can charge more for these kind of tokens, Nvidia says.

"For folks who are serving tokens, it unlocks the ability to offer premium tiers of service for those users and those customers who actually demand the most latency-sensitive" service agreements, Harris said on the call.

In December, Nvidia bought assets from chip startup Groq for $20 billion, the company's largest purchase..

The Groq architecture includes 500 megabytes of speedy SRAM on the chip's die itself to reduce memory-related bottlenecks. Groq chips are manufactured by Samsung, while Taiwan Semiconductor Manufacturing makes Nvidia's GPUs.

Nvidia packages 256 individual Groq 3 chips into its LPX racks. Nvidia said that its Groq 3 LPX rack can deliver 3,400 tokens per second, citing a benchmark from Artificial Analysis.

It's a competitive space. Smaller GPU maker Advanced Micro Devices announced earlier this year it would integrate its rack-scale systems with chips from Cerebras , which recently went public, focusing on low-latency inference. OpenAI's newly announced Ultrafast mode currently promises 750 tokens per second, and is "powered by Cerebras."

Low-latency chips don't replace the GPU, the workhorse of AI chips, which can do training as well as inference and are flexible enough to adapt to new technologies and models. Low-latency chips like Groq mainly focus on a part of serving models called the "decode" phase.

... continue reading