← ALL RELEASES

GOOGLE · 01 Sep 2026

Introducing agentic video understanding with Gemini

Google has launched a new agentic video understanding feature for Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite. This capability is available immediately for video uploads and YouTube videos through the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform. It uses standard API token pricing with no additional feature fee.

Unlike traditional static processing that ingests media at a fixed frame rate, agentic video understanding combines core reasoning with native video tools to dynamically search, scan, and inspect target video segments across visual frames, audio, and transcripts. The model actively determines what to watch, at what speed, and through which modality, fetching only the required moments and signals.

This approach reduces token consumption by up to 88 percent and lowers analysis costs by up to 66 percent while boosting accuracy by up to 7 percent. These efficiency gains are especially beneficial for long-form content ranging from 10-minute guides to multi-hour recordings. Specific capabilities include sub-second moment retrieval, long-form needle-in-a-haystack searches, anomaly detection, and precise counting of actions and objects. Among the tested models, Gemini 3.7 Flash with agentic understanding delivers the best overall quality and cost efficiency.

The feature will soon roll out to all users in the Gemini app across Flash and Flash-Lite models, and it will power the Ask YouTube feature on the YouTube watch page in the coming months.

Read the original ↗