메뉴
HN
Hacker News • 10일 전

구글, 젬미니 3.8 라이브 및 확장 사고 모델 공개

IMP
8/10
핵심 요약

구글이 음성 대화 특화 AI 모델인 Gemini 3.8 Live와 Gemini 3.8 Live Extended Thinking을 공개했습니다. 두 모델은 실시간 시각 맥락 이해, 병렬 추론, 대화 중단 없는 백그라운드 작업 처리를 지원하며, Artificial Analysis 음성-음성 품질 지수에서 1위(82.6점)를 차지했습니다. Gemini API, 구글 워크스페이스, 젬미니 앱에서 즉시 사용할 수 있습니다.

번역된 본문

Gemini 3.8 Live 및 3.8 Live Extended Thinking 소개 (2026년 9월 15일)

Gemini 3.8 Live와 Gemini 3.8 Live Extended Thinking은 구글의 가장 진보된 실시간 대화 모델입니다. 지능과 병렬 추론의 대대적인 업그레이드를 통해 음성으로 복잡한 작업을 수행하고 협업할 때 더욱 직관적인 경험을 제공합니다.

— Tom Ouyang(수석 엔지니어), Malini Jaganathan(기술 스태프), Gemini 오디오 팀 대표

구글은 Gemini 3.8 Live와 Gemini 3.8 Live Extended Thinking을 출시하여 음성 상호작용을 더욱 자연스럽고 유연하며 지능적으로 만들었습니다. 이 모델들은 복잡한 추론, 실시간 시각 맥락 처리, 대화를 방해하지 않는 백그라운드 작업 실행을 지원합니다. 오늘부터 Gemini API, Google Workspace, Gemini 앱에서 이 기능들을 사용할 수 있습니다.

주요 특징은 다음과 같습니다:

  • Gemini 3.8 Live: 대규모 서비스와 비용 효율성을 위해 설계되었으며, 대화형 지능과 유연한 대화, 시각적 근거(visual grounding)를 결합한 모델입니다.
  • Gemini 3.8 Live Extended Thinking: 높은 복잡도의 작업을 위해 설계되었으며, 향상된 지능과 다단계 추론 능력을 갖추고 있습니다.

개발자와 기업을 위해 이 모델들은 프로덕션 수준의 안정적인 음성 에이전트를 구축할 수 있는 기반 요소를 제공합니다. 또한 Gemini 앱, Google Workspace, 검색에서 젬미니와의 대화를 더 유연하고 협업적으로 만들어, 음성만으로 복잡한 작업을 처리할 수 있도록 돕습니다.

더 유연하고 지능적인 대화 경험

Gemini 3.8 Live Extended Thinking은 기업급 작업 완수 능력과 지능을 제공하며, Artificial Analysis의 음성-음성(Speech to Speech) 품질 지수에서 전체 1위(82.6점)를 기록했습니다. 또한 에이전트형 작업 완수 부문에서도 τ-Voice 벤치마크 68.6%, Sierra의 τ-Voice-banking 벤치마크 35.1%로 선두를 유지하고 있습니다. 추론 능력 면에서도 Big Bench Audio에서 97.7%의 높은 점수를 기록했으며, 다른 프론티어 모델들과 비교해 매우 경쟁력 있는 가격을 유지합니다.

Gemini 3.8 Live는 사용자 선호도에서 높은 평가를 받아 Speech Agent Arena에서 2위를 차지했습니다. 이러한 성능 외에도 높은 비용 효율성을 유지하여, 개발자와 기업에 대규모 서비스에 적합한 유능하고 효율적인 모델을 제공합니다.

ServiceNow의 EVA-Bench(음성 에이전트 평가 벤치마크)에서는 이 모델들이 복잡한 워크플로에서 성능과 비용의 균형을 성공적으로 이루며 파레토 프론티어를 확장했습니다.

원문 보기
원문 보기 (영어)
Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking Sep 15, 2026 | x.com Facebook LinkedIn Mail Copy link Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking are our most advanced live dialogue models yet. Major upgrades in intelligence and parallel reasoning make them more intuitive to collaborate with and use to execute complex tasks using your voice. Tom Ouyang Principal Engineer Malini Jaganathan Member of Technical Staff, on behalf of the Gemini Audio Team Share x.com Facebook LinkedIn Mail Copy link . Inlining them here makes them available in the DOM for the page. --> Your browser does not support the audio element. Listen to article [[duration]] minutes This content is generated by Google AI. Generative AI is experimental Voice Speed Voice Speed 0.75X 1X 1.5X 2X Read AI-generated summary We are launching Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking to make voice interactions more natural, fluid, and intelligent. These models handle complex reasoning, real-time visual context, and background task execution without interrupting your conversation. You can start using these features today through the Gemini API, Google Workspace, and the Gemini app. Summaries were generated by Google AI. Generative AI is experimental. Check out "Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking" for smarter voice AI. Gemini 3.8 Live offers fast, fluid conversations with real-time visual and language support. Use 3.8 Live Extended Thinking to handle complex tasks while keeping the conversation flowing. These models work in the background to manage tools while you keep chatting. You can try these new features in Google Workspace, Search, and the Gemini app. Summaries were generated by Google AI. Generative AI is experimental. Google just launched two new AI models that make talking to your devices feel way more natural. They can handle interruptions, switch between languages, and even explain their thought process while they work. Whether you're solving a complex problem or just chatting, the AI now feels like it's actually listening and thinking along with you. It’s a big step toward making AI feel like a real conversation partner. Summaries were generated by Google AI. Generative AI is experimental. Explore other styles: General summary Bullet points Basic explainer Today, we’re introducing two new models that bring advancements in near real-time reasoning to more effectively enable voice agents and make conversing with AI feel more intuitive and intelligent. Gemini 3.8 Live : Built for scale and cost efficiency, combining conversational intelligence with fluid dialogue and visual grounding. Gemini 3.8 Live Extended Thinking : Built for high-complexity tasks, with increased intelligence and multi-step reasoning. For developers and enterprises, these models deliver the building blocks for reliable, production-ready voice agents. They also make speaking with Gemini across the Gemini app, Google Workspace, and Search more fluid and collaborative — helping you tackle complex tasks using just your voice. Experience more fluid, intelligent conversations Gemini 3.8 Live Extended Thinking provides enterprise-grade task completion and intelligence, capturing the #1 overall spot on Artificial Analysis' Speech to Speech Quality Index (82.6), and leads in agentic task completion with 68.6% on τ -Voice and 35.1% on Sierra’s τ -Voice-banking benchmark. It also provides strong reasoning capabilities, scoring 97.7% on Big Bench Audio, while maintaining a highly competitive price point compared to other frontier models. Gemini 3.8 Live has shown a high preference among users, securing a second place in the Speech Agent Arena . In addition to this performance, it remains highly cost-effective — providing developers and enterprises with a capable and efficient model built for scale. On ServiceNow’s EVA-Bench , a benchmark for evaluating voice agents, our models push the Pareto Frontier for complex workflows by successfully balancing accuracy with conversational quality. Note: This was run on the Live API on Gemini Enterprise Agent Platform. Gemini 3.8 Live processes visual inputs in near real-time, enriching conversations with context for more helpful responses. It automatically detects and transitions between 97 supported languages mid-conversation. It executes tools and API calls in the background while continuing the conversation, so the model can acknowledge requests and keep chatting while tasks finish in the background. Gemini 3.8 Live guides employee onboarding in real time, using visual context to answer live questions. Watch Gemini 3.8 Live play chess in near real-time using visual context, reasoning, and natural conversational flow. For tasks that require deeper reasoning, 3.8 Live Extended Thinking reasons and speaks simultaneously. It delivers increased intelligence for complex workflows while maintaining an uninterrupted conversational flow — using early verbal cues like “Let me check that…” to acknowledge prompts naturally, and live progress narration to walk users through multi-step background tasks as they progress. Watch Gemini 3.8 Live Extended Thinking transform raw sketches and near real-time voice feedback into functional React components. See Gemini 3.8 Live Extended Thinking coordinate multi-step bookings and asynchronous function calls — all without interrupting natural live conversation. Watch Gemini 3.8 Live build complete business plans and custom marketing toolkits on the fly through natural speech. Across Google Workspace and Search, our Live models deliver more intuitive, collaborative experiences — especially when tackling your most complex tasks. Try Gemini 3.8 Live Extended Thinking in Google Workspace with Docs Live, Gmail Live, and Keep Live. Get step-by-step, real-time troubleshooting help powered by Gemini 3.8 Live — right inside Search Live. Empowering the developer and enterprise voice ecosystem By using the Gemini Live API , developer platforms such as Agora , Fishjam , LangChain , LiveKit , Pipecat , Vercel , and Vision Agents enable developers to build and deploy high-performance voice-driven interfaces with ease. These platforms manage complex real-time media streaming infrastructure behind the scenes, allowing developers to focus entirely on crafting the user experience. We’re also partnering with companies like Salesforce, Genspark, and Lumeris who are excited about 3.8 Live and 3.8 Live Extended Thinking, highlighting its impressive latency, fluidity, and tool-calling capabilities. Ensure transparency with SynthID watermarking All audio generated by our AI products is watermarked with SynthID . This imperceptible watermark is woven directly into the audio output, ensuring AI-generated content remains detectable to help prevent misinformation. For details on our approach to safety and responsibility, review the model card . Start using our latest Gemini Audio models: 3.8 Live is rolling out starting today: For developers : In the Gemini API and Google AI Studio For enterprises : In private preview in Gemini Enterprise and coming soon to Gemini Enterprise for Customer Experience For everyone : In Search Live 3.8 Live Extended Thinking is rolling out starting today: For developers : In the Gemini API and Google AI Studio For enterprises : In private preview in Gemini Enterprise and coming soon to Gemini Enterprise for Customer Experience and Google Workspace business customers For everyone : In Gemini Live and for Google AI Pro and Ultra subscribers in Workspace in Docs, and all Google AI subscribers in Gmail and Keep Get the latest news from Google in your inbox Sign up for our newsletters with product updates, event information, special offers, and more. Done. Just one step more. Check your inbox to confirm your subscription. You can also subscribe with a different email address . Your information will be used in accordance with Google's privacy policy. You may opt out at any time. Posted in: