Skip to content
Tech News
← Back to articles

Gimlet's Series B

read original more articles
Why This Matters

Gimlet Labs' recent Series B funding highlights the company's focus on advancing AI inference efficiency at scale, addressing critical industry challenges related to power consumption and latency. This investment underscores the importance of optimizing AI infrastructure to support exponential growth in token generation and AI workloads, ensuring sustainable and faster AI services for consumers and businesses alike.

Key Takeaways

Announcing Gimlet Labs’ Series B

Faster, More Efficient AI Inference at Scale

Today, we are announcing our $300M Series B raise, led by Andreessen Horowitz and joined by Sapphire Ventures, Menlo Ventures, 645 Ventures, Arm, Eclipse, Emergence, Factory, Hudson River Trading, M12, OnePrime Capital, Prosperity7, QuantumLight, Samsung Ventures, Tiger Global Management, Triatomic, Wing Ventures, and XTX Markets.

Since March, we’ve added billions in contracted revenue, gigawatts of datacenter pipeline, and are quickly scaling to hundreds of megawatts in managed capacity.

The Need for Throughput and Speed

Since our Series A just over 5 months ago, the demand for tokens has continued to increase at a breakneck pace. Monthly token generation has increased 6X in 12 months, with some projections stating another 20X increase by 2030.

To support this demand, the industry has embarked on perhaps the most ambitious investment project in modern history, already close to $1T per year and expected to reach a cumulative $7T by 2030. Power is quickly becoming the most critical resource bottleneck. AI data centers consumed approximately 18 GW of capacity in 2025, with that number expected to triple by 2030. Supporting this growth requires new power generation, grid capacity, interconnections, and behind-the-meter energy production. We’ll eventually hit the limits of the resources we can deliver at scale. Maximizing throughput per kW enables the industry to continue scaling while reducing the impact on physical resources.

In addition to raw throughput, we’re also seeing an increasing need for very fast inference tiers and low-latency tokens, driven by:

Agentic workloads. Agents execute multi-step inference loops, often requiring dozens of sequential model calls and tool interactions. Latency compounds across these calls.

Agents execute multi-step inference loops, often requiring dozens of sequential model calls and tool interactions. Latency compounds across these calls. Larger models. Models expanded from a few 100B parameters just a few years ago, reaching 3-10T parameters today.

... continue reading