
Google's Gemini cuts video-analysis token usage by up to 88%
Google's Gemini Flash models now analyze video selectively instead of scanning every frame, The Decoder reports, cutting token usage by up to 88% on benchmark tests.
- The change applies to Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite, and is live now through the Gemini API and Google AI Studio at standard token rates
- Instead of processing video at a fixed frame rate, the models now run a loop that selectively pulls frames, audio, or transcripts only from the sections that matter, deciding on their own what to retrieve and how fast to move
- On benchmarks including 1H-VideoQA and LongVideoBench, token usage dropped up to 88% and costs fell 66%, with accuracy holding steady or improving slightly despite processing less data
- Google says Gemini 3.7 Flash posted the highest accuracy at the lowest cost per query on its 1H-VideoQA evaluation, ahead of GPT 5.6 Terra, Claude Opus 5.0, and Grok 4.6 in its own comparison chart
The technique builds on "agentic vision," a capability Google shipped for Gemini 3 Flash in January that let models write their own Python code to manipulate images inside a think-act-observe loop. Applying the same self-directed approach to video is what lets the model skip frames instead of grinding through footage linearly, the same shift that made the earlier image feature notable. Google describes the target range as anything from a 10-minute tutorial to a multi-hour recording, treated by the same underlying loop.
The efficiency argument matters more for video than almost any other input type, because a fixed frame rate multiplies cost with runtime. A 10-minute tutorial and a 90-minute lecture cost roughly proportional amounts to process the old way, no matter how much of either recording is actually relevant to the question being asked. Letting the model decide what to skip is the same kind of resource discipline Intokened has tracked elsewhere as AI usage scales: agents are already out-consuming humans on platforms like OpenRouter, and Google's own Flash Cyber tool already showed the company optimizing narrower, task-specific models rather than defaulting to its biggest one for every job.
None of Google's benchmark numbers have been independently verified, and the company is comparing its own model against competitors using its own chart rather than a neutral third-party test. Still, the direction is clear enough: as AI video analysis moves from a novelty into something run at scale, on hours of footage rather than short clips, the cost of processing every frame the dumb way stops being a rounding error.
Nothing here should be taken as financial advice — just information to consider.

Comments (0)
No comments yet — be the first!
The market talks all day. We write when it says something
Short, and it tells you why it came
Related news
Most readTop 7
Silicon Valley Workers Are Wearing Noise-Cancelling Masks to Dictate AI Prompts
270AI





