Skip to content
Tech News
← Back to articles

Frore claims its LiquidJet can drop Nvidia Rubin GPU temperatures by 10°C — can also boost performance by 15% as hyperscalers eye using delidded GPUs in production environments

read original more articles
Why This Matters

Frore's LiquidJet cooling technology offers a significant advancement in thermal management for high-performance AI GPUs, potentially increasing their efficiency and output while reducing heat-related risks. This innovation is particularly crucial as data centers and hyperscalers seek to optimize performance and energy efficiency in AI workloads, leading to cost savings and enhanced hardware longevity.

Key Takeaways

It is not a secret that proper cooling ensures longevity and enables hardware to demonstrate its full potential. But when it comes to data center AI hardware, proper cooling also means higher sustained performance, which directly translates into money earned by the owner. Frore Systems, a maker of cooling solutions that are made using semiconductor-grade tools, seems to have a perfect idea of how to reduce the temperature of next-generation AI accelerators and increase their performance by 15%.

Frore Systems last week published a white paper which suggests that improvements to the entire cooling stack — from the GPU packaging and thermal interface materials (TIMs) to coldplates and coolant temperatures — can increase token generation per watt by more than 30%. Meanwhile, one of the company's boldest projections based on an analytical thermal model* is that its LiquidJet coldplate technology alone can lower Nvidia Rubin GPU junction temperatures by up to 12°C, which translates into a 10% to 25% improvement in tokens/Watt, while a 10°C reduction could increase token generation by around 15%.

Indeed, modern AI accelerators, such as the upcoming Nvidia Rubin, can dissipate up to 2,400 W, and their die temperatures can easily hit 95°C or more. But while 95°C is not necessarily a problem for silicon longevity, leakage current certainly is. Leakage current rises exponentially with temperature, approximately doubling for every 10°C increase in maximum junction temperature, which is when transistor switching itself also becomes less efficient. As a consequence, hotter GPUs require higher voltages to sustain clocks, which eventually forces Dynamic Voltage and Frequency Scaling (DVFS) to reduce clocks to remain within thermal limits, which in turn will reduce performance and token generation.

Latest Videos From Tom's Hardware Watch full video here:

(Image credit: Frore Systems)

This all leads to a simple conclusion: the better the cooling, the higher the performance and token output. Which is generally right. However, cooling is not as simple, as it depends on multiple factors that can be optimized. Furthermore, for AI data centers, cooling itself is no longer a way to preserve CPUs and accelerators from overheating, but really is a way to maximize their performance and token money generation.

Nvidia designs its platforms around Tj(max) temperature; it is one of the fundamental design constraints for the GPU, package, and cooling solution. This works like this:

Nvidia specifies a maximum allowable junction temperature (Tj,max limit). This is the temperature the silicon must not exceed during normal operation. The exact value is not always public, but Frore uses 95°C for Rubin in its analysis.

The GPU continuously monitors its junction temperature using tens or hundreds of on-die thermal sensors, yet power management monitors the hottest region.

DVFS attempts to maximize performance while staying below the thermal and power limits, so if the GPU has thermal headroom, it can sustain higher clocks or lower voltage. If the junction temperature rises, the firmware gradually adjusts voltage and frequency. If necessary, it throttles to prevent exceeding Tj(max).

... continue reading