Introducing agentic video understanding with Gemini
Google has launched a new agentic video understanding feature for the Gemini 3.7 Flash, Gemini 3.6 Flash, and Gemini 3.5 Flash-Lite models. Unlike traditional static video processing that ingests media at a fixed frame rate, this capability pairs core model reasoning with native video tools. It allows the models to dynamically scan visual frames, audio, and transcripts, taking an active role in determining what to watch and at what speed.
The new feature improves analysis accuracy by up to 7 percent while cutting token consumption by up to 88 percent and reducing costs by up to 66 percent. These efficiency gains are especially beneficial for long-form video content, ranging from 10-minute guides to multi-hour recordings, where static methods previously forced a compromise between high costs and lost details. It unlocks specific video processing capabilities, including sub-second moment retrieval, precise anomaly detection, counting actions and objects, and long-form needle-in-a-haystack searches.
Developers can access the feature immediately via the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform by setting the processing configuration to agentic, using standard token pricing with no extra feature fee. Google is also rolling out the capability to billions of users through the Gemini app and plans to power the Ask YouTube feature on YouTube watch pages in the coming months.