Google DeepMind launches Gemini agent-based video understanding capabilities.
Google DeepMind announced the introduction of agent-based video understanding capabilities in its Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite models. By dynamically scanning video clips, token consumption can be reduced by up to 88%, costs by up to 66%, and accuracy improved by up to 7%. This feature is now available through the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform.
Key Facts
On September 1, 2026, Google DeepMind announced in a blog post the introduction of agent-based video understanding in Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite models. This feature allows models to dynamically scan video segments instead of processing the entire video at a fixed frame rate, significantly reducing token consumption and cost while improving accuracy. According to official data, this feature can reduce token consumption by up to 88%, cost by up to 66%, and accuracy by up to 7%. This feature is now available through the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform, supporting video uploads and YouTube videos.
Background and Impact
Proxy-based video understanding, similar to proxy-based vision, combines code execution with the image understanding capabilities of the Gemini model. It leverages Gemini's native video tools to achieve sub-second time-lapse retrieval, more accurate anomaly detection, and precise counting. This feature is particularly suitable for long video processing, such as 10-minute tutorials, 90-minute lectures, or multi-hour recordings, solving the problems of high token costs or loss of detail in static processing. Google DeepMind stated that this feature will be gradually rolled out to Gemini applications and plans to support YouTube's "Ask YouTube" feature in the coming months.
Restrictions and Sources
The performance data mentioned above (such as an 88% reduction in token consumption, a 66% reduction in cost, and a 7% improvement in accuracy) are from the official Google DeepMind blog and have not been independently verified. This feature is currently only available through the Gemini API; please refer to the official documentation for Google AI Studio and the Gemini Enterprise Agent Platform for specific availability, pricing, and regional support. Detailed feedback from early partners is not yet publicly available. Source:Google DeepMind Blog。