Google DeepMind announced agentic video understanding for Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite, citing up to 88% token reduction and 7% accuracy gain.
Google DeepMind Adds Agentic Video Understanding to Gemini Models
The combination of sharply lower token costs and higher accuracy makes video analysis materially cheaper and more capable for developers and enterprise users building on Gemini. The rollout to YouTube's watch page extends the capability to a large consumer surface.
The full picture
Google DeepMind launched agentic video understanding for Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite. The feature lets models dynamically search, scan, and inspect video segments across frames, audio, and transcripts, rather than ingesting at a fixed frame rate. According to Google DeepMind, this reduces token consumption by up to 88% and costs by up to 66%, while improving accuracy by up to 7% on standard benchmarks. Gemini 3.7 Flash with the feature sits at the accuracy-to-cost Pareto frontier among the models tested. Newly enabled capabilities include sub-second moment retrieval, long-form needle-in-a-haystack search, anomaly detection, and accurate action and object counting. The feature will power YouTube's 'Ask YouTube' on the video watch page and will roll out to Gemini app users across Flash and Flash-Lite models.
How it developed
Sources
Related
Want this in your inbox?
I send a short email each morning with the stories that moved. If you would rather just read here, that works too.
Subscribe free