Reddit
Gemini 3.7 Flash cuts video-understanding token consumption by up to 88% by deciding which segments to watch
Google launched agentic video understanding across Gemini 3.7 Flash, 3.6 Flash and 3.5 Flash-Lite. Instead of scanning a video start to finish at a fixed rate, the model runs an internal agentic loop that decides what to watch, at what speed, and through which channel (frames, audio or transcript), fetching only the segments needed to answer the query. Google reports up to 88% lower token consumption, up to 7% higher accuracy and roughly 66% lower cost, placing it on the accuracy-to-cost pareto frontier for tested video models. It is live now via the Gemini API in AI Studio and the Gemini Enterprise Agent Platform, with the consumer app to follow.
↳ Follow the thread