Google has introduced agent-based video analysis across its Gemini Flash model line, allowing models to selectively scan video files instead of processing continuous frame rates. Available on Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite, the feature slashes token consumption and API costs by up to 88 percent.

Rather than analyzing fixed one-frame-per-second inputs, the model autonomously decides which video segments, frame rates, and audio tracks to inspect based on the query. This targeted approach enables precise tracking of short state changes, sub-second edits, and repetitive movements across multi-hour video files without generating massive token overhead.

The feature is available at standard API token rates in Google AI Studio and Gemini Enterprise. Google plans to bring agentic video analysis to the consumer Gemini app and use it to power YouTube’s interactive query tools.

Why it matters

  • Agentic video sampling cuts API token consumption by up to 88% on long video inputs.

  • Sub-second scanning enables fine-grained automated video editing and anomaly tracking at scale.

  • Developers can enable the capability in the Gemini API without paying additional feature fees.

Source: the-decoder.com