메뉴
BL
MarkTechPost • 10일 전

구글, 프로덕션급 음성 에이전트용 '제미나이 3.8 라이브' 공개

IMP
8/10
핵심 요약

구글이 실시간 대화 모델 최신버전인 Gemini 3.8 Live와 Gemini 3.8 Live Extended Thinking을 출시했습니다. 두 모델은 대화가 이어지는 동안 백그라운드에서 도구·API 호출을 실행하고, 실시간 영상 입력을 처리하며, 대화 중 97개 언어 간 전환이 가능합니다. Extended Thinking은 Artificial Analysis 음성-음성 품질 지수에서 82.6점으로 1위를 차지했으며, 분당 0.005달러의 오디오 입력 요금으로 Gemini API와 Google AI Studio에서 즉시 사용할 수 있습니다.

번역된 본문

구글은 지금까지 공개된 실시간 대화 모델 중 가장 발전된 형태인 Gemini 3.8 Live와 Gemini 3.8 Live Extended Thinking을 출시했다고 발표했습니다. 이 모델들은 대화가 계속 흘러가는 동안에도 백그라운드에서 도구 실행 및 API 호출을 수행하고, 실시간 시각 입력을 처리하며, 대화 도중에 97개 언어 사이를 자유롭게 전환할 수 있습니다. Extended Thinking은 Artificial Analysis의 음성-음성(Speech to Speech) 품질 지수에서 82.6점으로 1위를 기록했고, Big Bench Audio에서는 97.7%의 점수를 달성했습니다. 두 모델은 오늘부터 Gemini API와 Google AI Studio에서 이용할 수 있으며, 오디오 입력 요금은 분당 0.005달러입니다. 또한 생성된 모든 오디오에는 구글 딥마인드의 SynthID 워터마크가 삽입됩니다.

원문 보기
원문 보기 (영어)
Google has released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, its most advanced live dialogue models to date. The models execute tools and API calls in the background while the conversation keeps flowing, process live visual inputs, and switch between 97 languages mid conversation. Extended Thinking ranks #1 on Artificial Analysis' Speech to Speech Quality Index with 82.6 and scores 97.7% on Big Bench Audio. Both are available today in the Gemini API and Google AI Studio at $0.005/min for audio input, with all generated audio carrying Google DeepMind's SynthID watermark. The post Google Releases Gemini 3.8 Live and 3.8 Live Extended Thinking for Production Grade Voice Agents appeared first on MarkTechPost.