Google has launched agentic video understanding across Gemini 3.7 Flash, 3.6 Flash and 3.5 Flash-Lite, so the models can search, scan and inspect long videos instead of splitting them into frames at a fixed rate.
The capability is live today for video uploads and YouTube videos through the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform, Google said on 1 September 2026.
How Gemini video understanding works
Static processing still ingests a clip at a fixed frames-per-second rate, defaulting to 1 FPS. Agentic video understanding pairs Gemini’s core reasoning with native video tools, so the model decides what to watch, at what speed, and whether to use frames, audio or the transcript, fetching only the moments it needs.
Across standard video analysis benchmarks, Google says the agentic path cuts token consumption by up to 88% and analysis costs by up to 66%, while improving accuracy by up to 7%. Gemini API docs on ai.google.dev describe the same mode as up to 88% more token-efficient and about 7% higher quality on long-form content.

Android Authority reported the token and accuracy figures and noted that Gemini previously had to split videos into individual frames.
What it can do with long clips
Google lists four jobs for the agentic loop: sub-second moment retrieval, including split-second state changes missed at 1 FPS; long-form search that can answer complex queries across multi-hour videos; anomaly detection by resampling interesting windows at a higher frame rate; and counting repeated physical movements and distinct objects over time.
Google says those efficiency gains are especially pronounced on long-form video, from 10-minute how-to guides to 90-minute lectures and multi-hour recordings, where static processing forces a choice between high token costs or dropping detail. Google says Gemini 3.7 Flash with agentic understanding offered the best quality, and the best quality-to-cost mix, of the models it tested.
tbreak had already flagged backend groundwork for Gemini 3.7 Flash before this week’s launch.
Where Gemini video understanding is available
The API path is open now. Google Cloud’s video-understanding docs, last updated 1 September 2026, list the same three Flash models and label agentic video understanding as a Generative AI Preview. There is no extra feature fee. Google says it uses standard Gemini API token pricing; the launch post does not publish per-token rates.
The feature will roll out to all users in the Gemini app across Flash and Flash-Lite models soon. In the coming months it will also power YouTube’s Ask YouTube feature on the video watch page, so answers can be grounded in the visuals.
Google did not name countries. The API is available now through AI Studio and Enterprise. Timing for the Gemini app, and for Ask YouTube, is unconfirmed in the UAE.
tbreak previously covered Ask YouTube as a US YouTube Premium search experiment. This week’s post points to the watch page, not that earlier search trial.
Developers turn the mode on by setting processing to “agentic” in the API configuration. The same studio path already carries other Gemini video tools, including the recent Google Flow update with Gemini Omni 1.1.
The Gemini API in AI Studio and Enterprise can run agentic video understanding today. The Gemini app is next. Ask YouTube on the watch page is listed for the coming months. Google has not confirmed a UAE consumer date.


















