The Information Machine
Following·since 1 Sep 2026·New·2 sources

Google DeepMind Adds Agentic Video Understanding to Gemini Models

The gist

The combination of sharply lower token costs and higher accuracy makes video analysis materially cheaper and more capable for developers and enterprise users building on Gemini. The rollout to YouTube's watch page extends the capability to a large consumer surface.

The full picture

Google DeepMind launched agentic video understanding for Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite. The feature lets models dynamically search, scan, and inspect video segments across frames, audio, and transcripts, rather than ingesting at a fixed frame rate. According to Google DeepMind, this reduces token consumption by up to 88% and costs by up to 66%, while improving accuracy by up to 7% on standard benchmarks. Gemini 3.7 Flash with the feature sits at the accuracy-to-cost Pareto frontier among the models tested. Newly enabled capabilities include sub-second moment retrieval, long-form needle-in-a-haystack search, anomaly detection, and accurate action and object counting. The feature will power YouTube's 'Ask YouTube' on the video watch page and will roll out to Gemini app users across Flash and Flash-Lite models.

How it developed
1 September 2026

Google DeepMind announced agentic video understanding for Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite, citing up to 88% token reduction and 7% accuracy gain.

Sources
The daily email

Want this in your inbox?

I send a short email each morning with the stories that moved. If you would rather just read here, that works too.

Subscribe free