Skip to content
Tech News
← Back to articles

Gemini now analyzes your videos more accurately, and at a lower cost

read original more articles
Why This Matters

Google's Gemini now features agentic video understanding, enabling more accurate and cost-effective analysis of long videos by intelligently selecting frames, audio, and transcripts. This advancement allows for precise tracking of movements, object counting, and complex question answering, enhancing AI capabilities in video processing. The new features are set to improve tools for content creators and enterprise users, with broader availability planned soon.

Key Takeaways

Edgar Cervantes / Android Authority

TL;DR Google has introduced agentic video understanding in Gemini.

Gemini can now easily analyze long videos with sub-second accuracy.

It can also track physical movement and count distinct objects in videos.

Google has been steadily improving Gemini to make it more useful. We’ve already spotted the company working on new features for Gemini on mobile, while also releasing new Gemini models such as Gemini 3.7 Flash. However, Google is now giving Gemini the ability to analyze videos in a more cost-effective and efficient manner.

Google announced today that it’s bringing agentic video understanding to the latest Gemini models: Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite. This new capability allows users to upload videos to Gemini and have the AI analyze them. It also reduces token usage by up to 88% and offers up to 7% better accuracy.

Users have been able to upload videos to Gemini for a while now, but the AI could only perform what Google calls “static” processing on them. This meant it would split the video into individual frames, resulting in slower performance and higher costs.

With agentic video understanding, Gemini can now decide what to watch and at what speed. It can also choose between frames, audio, and the transcript to analyze videos more efficiently. This means Gemini can pinpoint split-second changes, answer complex questions across multi-hour videos, inspect videos for visual artifacts, and even count and track physical movements and objects.

Right now, agentic video understanding is only available via the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform. However, the company also said the feature will roll out to the Gemini app soon. Google will also soon start using agentic video understanding to power YouTube’s “Ask YouTube” feature, which could make it easier for creators to analyze their videos with Gemini.

Follow