How does agentic video understanding improve AI performance?
The new agentic video understanding feature allows Gemini models to dynamically scan video segments rather than relying on static frame processing. By actively determining which parts of a video to inspect, the technology reduces token consumption by up to 88% and lowers costs by up to 66%. Additionally, the update improves overall video analysis accuracy by up to 7% across supported models.
This capability is now available for developers via the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform. By enabling an agentic loop, the models can perform complex tasks such as sub-second moment retrieval, anomaly detection, and precise object counting more effectively than previous methods. The feature is particularly beneficial for processing long-form content, such as multi-hour recordings or lengthy lectures.
When will this feature be available to all users?
While currently accessible to developers, Google plans to expand the feature to broader audiences in the coming months. It will roll out to the Gemini app across Flash and Flash-Lite models, and will also be integrated into YouTube's 'Ask YouTube' feature. This integration aims to provide users with higher-quality, visually grounded answers directly on video watch pages.