Home/ai-models/Google launches agentic video analysis for Gemini models, cutting token use and cost
A detailed pencil sketch of a film reel partially unspooling, with glowing data points hovering over specific frames, illustrating selective video analysis; no text, no logos.
AI ModelsPublished 1 September 20262 min read

Google launches agentic video analysis for Gemini models, cutting token use and cost

Google announced that its Gemini 3.7 Flash, 3.6 Flash and 3.5 Flash‑Lite models now support an agentic video understanding capability.

The new feature lets developers set the API configuration to “agentic” in Google AI Studio or the Gemini Enterprise Agent Platform to process video uploads and YouTube videos.

Agentic video understanding differs from the prior static approach that consumes video at a fixed frame‑per‑second rate, typically one frame per second.

How Agentic Video Understanding Works

Instead of ingesting every frame, the model takes an active, goal‑directed role, deciding which segments to watch, at what speed, and through which modality—visual frames, audio, or transcript.

During this loop the model invokes an internal tool that loads only the relevant portion of the video file, thereby avoiding unnecessary processing.

This dynamic scanning reduces the development effort that previously required manual coding to locate important moments.

Performance Gains and Use Cases

Benchmarks on standard video analysis tasks show token consumption falling by as much as 88 % and analysis costs dropping by up to 66 % compared with static processing.

Accuracy improves by up to 7 % when the agentic mode is enabled, with Gemini 3.7 Flash delivering the strongest quality‑cost balance.

The efficiency gains are most pronounced on long‑form content, ranging from ten‑minute how‑to guides to ninety‑minute lectures and multi‑hour recordings.

For such extended videos, static processing forces a trade‑off between high token costs and loss of critical details, a dilemma that agentic video understanding resolves.

By focusing on the moments that matter, Gemini 3.7 Flash with agentic understanding sits on the accuracy‑to‑cost Pareto frontier among tested video models.

Rohan Doshi, Senior Product Manager at Google DeepMind, and Mario Lučić, Research Director at Google DeepMind, highlighted that the capability unlocks sub‑second moment retrieval, more precise anomaly detection, and accurate counting.

Sub‑second moment retrieval enables developers to pinpoint split‑second state changes that static frame rates often miss.

Improved anomaly detection can surface unexpected events in surveillance or quality‑control footage with greater reliability.

Precise counting supports applications such as inventory tracking or audience measurement without exhaustive frame‑by‑frame analysis.

Because the model can also query audio and transcript streams, it can extract spoken cues that visual frames alone would overlook.

The agentic approach therefore expands the range of video‑centric AI solutions that can be built cost‑effectively on Google’s platform.

Developers can activate the feature today, allowing immediate experimentation without waiting for separate tooling.

Google’s rollout positions its Gemini family as a competitive option for enterprises seeking scalable video insights.

Why This Matters: the agentic video feature reduces token usage and cost while improving accuracy, making large‑scale video analysis more accessible.

#ai-models#ai#digest#auto

This digest was compiled from:

Share this digest

Share on XWhatsAppLinkedInTelegram

People Also Ask

Share your thoughts

Reactions, corrections, or insights — all welcome.

0/2000