Skip to content
Tech News
← Back to articles

OpenAI says its Jalapeño chip can power faster AI responses than the competition

read original more articles
Why This Matters

OpenAI's new Jalapeño AI chip significantly enhances inference efficiency by delivering faster responses and lower latency than competing solutions, which can lead to more responsive AI applications and improved user experiences. Its higher energy efficiency also supports sustainable scaling of AI services, addressing growing demand while reducing operational costs. This advancement underscores the ongoing innovation in AI hardware, potentially setting new industry standards for performance and efficiency.

Key Takeaways

is a news writer who covers the streaming wars, consumer tech, crypto, social media, and much more. Previously, she was a writer and editor at MUO.

Posts from this author will be added to your daily email digest and your homepage feed.

OpenAI says its new AI chip, Jalapeño, completes tasks more efficiently and returns responses faster than other AI systems, according to a blog post published on Tuesday. During a briefing with reporters, OpenAI hardware vice president Richard Ho said Jalapeño offers the “best of both worlds” with lower latency and higher throughput, as AI systems typically “have to make a trade-off between the two.”

First introduced in June, Jalapeño is an Application-Specific Integrated Circuit (ASIC) made in partnership with Broadcom. It’s designed for AI inference — the process of running a trained AI model to complete a task or deploy an agent.

This chart measures Jalapeño’s time between tokens (TBT) — or the time it takes to deliver a response. Image: OpenAI

To measure Jalapeño’s performance, OpenAI used InferenceX, a benchmarking platform that shows how well AI systems handle inference. The test compared Jalapeño’s performance against the best results recorded at the time, which were with Nvidia’s GB200 or GB300 superchips. OpenAI says Jalapeño delivered 1.5 to 1.9 times more AI work per watt across GPT-OSS 120B, DeepSeek R1, and Kimi K2.5 1T than the comparison systems, while offering 1.7 to 3.6 times lower end-to-end latency across the three models. That means the chip can provide users with “faster responses, more responsive agents, and more reliable access as the demand grows,” according to Ho.

Jalapeño can perform more work while using less energy, according to OpenAI. Image: OpenAI

OpenAI plans to deploy Jalapeño in “small volumes” by the end of this year, but will begin to “ramp the volume up” into 2027, Ho added. The company doesn’t say how many chips it plans to deploy next year, however.

Even with these performance improvements, Ho said OpenAI doesn’t expect to replace its entire chip lineup with Jalapeño, saying its overall compute strategy includes “very good partners,” like Nvidia. OpenAI will continue developing the second and third generations of the new chip.