메뉴
BL
The Decoder • 24일 전

구글 제미나이 에이전트 기반 영상 분석, 토큰 최대 88% 절감

IMP
7/10
핵심 요약

구글이 제미나이 Flash 모델에 에이전트 기반 영상 분석 기능을 도입했습니다. 프레임별 고정 스캔 대신 모델이 스스로 필요한 구간과 방식(프레임, 오디오, 트랜스크립트)을 결정해 토큰 사용량을 최대 88%, 비용을 66% 줄이면서 정확도도 높였습니다. 현재 Gemini API에서 이용 가능하며, 향후 제미나이 앱과 유튜브 'Ask YouTube'에도 적용될 예정입니다.

번역된 본문

구글 제미나이의 새로운 에이전트 기반 영상 분석, 토큰 사용량 최대 88% 절감

핵심 요점:

  • 구글은 제미나이 Flash 모델에 고정된 패턴으로 프레임 단위 스캔하는 대신 영상 속에서 능동적으로 탐색하는 에이전트 기반 영상 분석 기능을 탑재하고 있습니다.
  • 모델이 어떤 구간을 어떻게 분석할지 스스로 결정하며, 구글에 따르면 이를 통해 토큰을 최대 88% 절감하고 비용을 66% 줄이며 정확도까지 향상됩니다.
  • 이 기능은 현재 Gemini API를 통해 긴 영상에서 특정 장면을 찾는 데 사용할 수 있으며, 제미나이 앱과 유튜브 통합은 추후 계획되어 있습니다.

구글이 여러 제미나이 모델에 에이전트 기반 영상 분석을 추가하고 있습니다. 고정된 프레임 속도로 영상을 프레임 단위로 스캔하는 대신, 모델이 스스로 관련 구간을 탐색해 구글이 밝힌 바에 따르면 토큰 사용량과 비용을 크게 줄일 수 있습니다.

최신 모델인 제미나이 3.7 Flash, 3.6 Flash, 3.5 Flash-Lite는 1초 미만의 순간, 즉 초당 1프레임 방식에서는 놓칠 수 있는 상태 변화나 컷 전환까지 포착할 수 있습니다. 구글은 이를 통해 자동 영상 편집이 훨씬 정밀해진다고 말합니다. 이 시스템은 수백만 개의 토큰을 소모하지 않고도 수 시간 분량의 영상에서 개별 장면을 추적할 수 있습니다. 의심스러운 시간 구간을 더 높은 프레임 속도로 재샘플링하여 이상 징후를 감지하고, 시간 경과에 따른 반복 동작과 개별 객체를 정확하게 카운트합니다.

지금까지 제미나이는 정적 처리 방식에 의존했으며, 기본적으로 초당 1프레임(API로 조정 가능)으로 영상을 샘플링했습니다. 2025년 네이티브 영상 분석이 출시된 이래, 제미나이는 오디오 트랙을 텍스트로 변환하고 프레임을 초 단위로 분석해 왔습니다.

구글에 따르면 에이전트 기반 방식은 모델의 추론을 네이티브 영상 도구와 직접 연결합니다. 모델은 어떤 구간을, 어떤 속도로, 어떤 모달리티(프레임, 오디오, 트랜스크립트)로 볼지 스스로 결정합니다. 주어진 작업에 실제로 필요한 순간과 신호만 가져옵니다.

제미나이는 이제 영상에서 중요한 부분만 로드합니다. 내부 도구를 사용해 관련 있는 부분만 영상 파일에서 가져오는 것입니다. 개발자가 이런 선택적 접근 방식을 이전에 수동으로 구현할 수도 있었지만, 이제는 모델이 자체적으로 처리합니다.

이 접근법은 구글이 1월에 제미나이 3 Flash에 출시한 '에이전틱 비전(agentic vision)'에 기반합니다. 해당 기능은 모델이 파이썬 코드를 작성·실행하여 이미지를 확대, 자르기, 주석 처리하고, 응답 전에 생각-행동-관찰(think-act-observe) 루프로 각 결과를 확인하도록 했습니다. 출시 당시 모든 경우에 자동으로 작동하지는 않았지만 기반은 이미 마련되어 있었습니다. 구글이 12월 제미나이 3 Flash를 발표할 때 영상에 대한 시각적·공간적 추론을 향후 기능으로 언급한 바 있습니다.

효율성 향상은 긴 영상에서 가장 두드러지며, 10분짜리 튜토리얼부터 90분 강의, 수 시간 녹화본까지 해당됩니다. 정적 처리에서는 개발자가 높은 토큰 비용과 중요한 디테일을 버리는 방식 중 하나를 선택해야 했습니다. LongVideoBench 등 구글 자체 벤치마크에서 에이전트 기반 분석을 적용한 제미나이 3.7 Flash가 전반적인 품질에서 최고 점수를 기록했으며, 정확도와 비용 효율성의 최적 조합을 제공했습니다.

API를 통한 추가 요금 없음 이 기능은 구글 AI 스튜디오와 제미나이 엔터프라이즈 에이전트 플랫폼의 Gemini API를 통해 영상 업로드 및 유튜브 영상에 대해 사용할 수 있습니다. 개발자는 API 설정에서 처리 모드를 'agentic'으로 지정하기만 하면 되며, 표준 제미나이 API 토큰 요율이 적용되고 추가 요금은 없습니다. 자세한 내용은 개발자 가이드에서 확인할 수 있습니다.

구글은 이러한 개선 사항을 자체 제품에도 적용할 계획입니다. 이 기능은 곧 Flash 및 Flash Lite 기기의 모든 제미나이 앱 사용자에게 rollout될 예정입니다. 향후 몇 달간 에이전트 기반 영상 분석은 재생 페이지의 'Ask YouTube' 기능에도 적용되어, 영상에 실제로 보이는 내용과 더 밀접하게 연결된 답변을 제공하게 됩니다.

원문 보기
원문 보기 (영어)
Google Gemini's new agent-based video analysis cuts token usage by up to 88 percent Jonathan Kemper View the LinkedIn Profile of Jonathan Kemper Sep 2, 2026 Google Key Points Google is equipping its Gemini Flash models with agent-based video analysis that dynamically searches through video footage instead of scanning it frame by frame in a fixed pattern. The models autonomously decide which sections of a video to analyze and how, which, according to Google, saves up to 88 percent in tokens, cuts costs by 66 percent, and improves accuracy. The feature is now available through the Gemini API for finding specific scenes in long videos, with integration into the Gemini app and YouTube planned for a later date. Ask about this article… Search Google is adding agent-based video analysis to several Gemini models. Instead of scanning a video frame by frame at a fixed rate, the model hunts for relevant sections on its own, which Google says cuts token usage and costs by a wide margin. The latest models, Gemini 3.7 Flash , 3.6 Flash, and 3.5 Flash-Lite , can pick up moments shorter than one second, including state changes or cuts that would slip through at one frame per second. Google says this makes automated video editing far more precise. The system can also track down individual scenes in hours of footage without burning through millions of tokens. It spots anomalies by resampling suspicious time windows at a higher frame rate and accurately counts repeated movements and individual objects over time. Ad Until now, Gemini relied on static processing, sampling video at a fixed frame rate of one frame per second by default and adjustable through the API. Since native video analysis launched in 2025 , Gemini has been transcribing the audio track and analyzing frames on a per-second basis. Ad Google says the agent-based variant ties the model's reasoning directly to native video tools. The model decides on its own which sections to look at, at what speed, and through which modality, whether that means frames, audio, or transcript. It only pulls the moments and signals it actually needs for a given task. Gemini now loads only the parts of the video that matter Gemini now uses an internal tool to grab just the relevant portion of the video file. Developers could build this kind of selective approach manually before, but the model now handles it on its own. Ad The approach builds on "agentic vision," which Google shipped for Gemini 3 Flash in January . That feature let the model write and run Python code to zoom, crop, and annotate images, checking each result in a think-act-observe loop before responding. It didn't work automatically in every case at launch, but the groundwork was already there. When Google announced Gemini 3 Flash back in December , the company flagged visual and spatial reasoning for video as a coming capability. The efficiency gains show up most with long videos, anywhere from 10-minute tutorials to 90-minute lectures and multi-hour recordings. With static processing, developers had to pick between high token costs and methods that throw away important details. On Google's own benchmarks, including LongVideoBench, Gemini 3.7 Flash with agent-based analysis scores the highest overall quality and delivers the best mix of accuracy and cost efficiency. Ad No extra charge through the API The feature is live for video uploads and YouTube videos through the Gemini API in Google AI Studio and on the Gemini Enterprise Agent Platform . Developers set the processing mode to "agentic" in the API config and pay standard Gemini API token rates with no added fee. More details are in the Developer Guide . Ad Google plans to bring these improvements to its own products too. The feature should roll out soon to all Gemini app users on Flash and Flash Lite devices. Over the coming months, agent-based video analysis will also power the "Ask YouTube" feature on the playback page, giving answers more closely tied to what's actually visible in the video. AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now Source: Google