Skip to content
Tech News
← Back to articles

Google reveals faster and cheaper Gemini 3.6 Flash, says 3.5 Pro is still in testing

read original more articles
Why This Matters

Google's release of Gemini 3.6 Flash marks a significant step in advancing AI model efficiency and performance, offering faster, more cost-effective solutions for developers and businesses. The improvements focus on better coding capabilities, multimodal features, and reduced token usage, which can lead to substantial cost savings and enhanced AI applications. Despite delays in the Gemini 3.5 Pro, these updates demonstrate Google's ongoing commitment to refining its AI offerings for a competitive edge in the industry.

Key Takeaways

Google announced a significant evolution of its AI models at I/O in May with the release of Gemini 3.5 Flash, and it’s not slowing down. The company has revealed three new AI models today, including its first version of Gemini geared toward cybersecurity. However, none of the new models is the delayed Gemini 3.5 Pro, which was supposed to launch in June.

Gemini 3.5 Flash, which was the star of the show at I/O, has already been deprecated. In its place, developers and users will find Gemini 3.6 Flash. Google makes the usual claims about this model—it’s marginally more capable and better at coding, and it has great multimodal features.

Google says the changes to 3.6 Flash were made in response to user feedback on the 3.5 release. In general, Gemini 3.5 Flash didn’t appear to live up to Google’s promises around code generation. Perhaps that’s simply a consequence of Google’s intense focus on efficiency as businesses have started to fret over the cost of AI tokens.

Credit: Google Credit: Google

In the DeepSWE test for coding, 3.6 Flash jumps to 49 percent versus 37 percent for 3.5 Flash. The new model now supports computer use as a standard feature in the Gemini API, too. The OSWorld test for computer use shows a modest boost to 83 percent from 3.5’s 78.4 percent score. Efficiency was a big focus for Gemini 3.5 Flash, and Google says that effort has been amped up with 3.6. Even with small benchmark gains, Gemini 3.6 Flash uses about 17 percent fewer tokens.

In agentic workflows (like the one below), Gemini 3.6 Flash should complete tasks more accurately, in fewer steps, and with fewer tokens. That could save developers (and Google) a lot of money. The new model has a lower API cost, at $1.50/1M input tokens and $7.50/1M output tokens. It was $1.50 and $9, respectively, for 3.5 Flash.

Token efficiency in Gemini 3.6 Flash. Token efficiency in Gemini 3.6 Flash.

Google is not done with the 3.5 branch yet, though. It has also released Gemini 3.5 Flash Lite and 3.5 Flash Cyber. The new Flash Lite is Google’s most efficient modern AI, hitting an impressive 350 tokens per second. The company claims this model is ideal for scaling agentic systems without breaking the bank. Based on benchmark numbers, the new Flash Lite is almost on par with frontier models from about a year ago, but it’s cheap. Pricing is set at $0.30/1M input tokens and $2.50/1M output tokens, though that is slightly higher than the previous 3.1 Flash Lite ($0.25 and $1.50).