Google Launches Agentic Video Understanding for Gemini

Turan Adalat
2 Min Read

How does agentic video understanding improve AI performance?

The new agentic video understanding feature allows Gemini models to dynamically scan video segments rather than relying on static frame processing. By actively determining which parts of a video to inspect, the technology reduces token consumption by up to 88% and lowers costs by up to 66%. Additionally, the update improves overall video analysis accuracy by up to 7% across supported models.

This capability is now available for developers via the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform. By enabling an agentic loop, the models can perform complex tasks such as sub-second moment retrieval, anomaly detection, and precise object counting more effectively than previous methods. The feature is particularly beneficial for processing long-form content, such as multi-hour recordings or lengthy lectures.

When will this feature be available to all users?

While currently accessible to developers, Google plans to expand the feature to broader audiences in the coming months. It will roll out to the Gemini app across Flash and Flash-Lite models, and will also be integrated into YouTube's 'Ask YouTube' feature. This integration aims to provide users with higher-quality, visually grounded answers directly on video watch pages.

Share This Article
Turan Adalat Novruzova has been working professionally in the media industry since 2010. Over the years, she has worked as a reporter, editor, and presenter at AzTV, ARB TV, ANS, and Baku TV. Her areas of expertise include current affairs, society, politics, and social issues. She has produced reports on a wide range of topics and has also voiced news texts for the Axar.az news website. She is currently a presenter of the morning news program on ATV. In addition, she provides training as a speech and diction specialist.