메뉴
HN
Hacker News • 2일 전

구글, 제미나이 3.8 TTS 음성 생성 모델 공개

IMP
7/10
핵심 요약

구글이 감정, 억양, 속도 등을 자연어 프롬프트로 세밀하게 조절할 수 있는 텍스트 음성 변환(TTS) 모델 'Gemini 3.8 Flash TTS'와 'Gemini 3.8 Flash-Lite TTS'를 공개했습니다. 100개 이상 언어로 맞춤 음성을 처음부터 만들거나 30초 샘플로 기존 음성을 재현할 수 있어 오디오북, 팟캐스트, 게임, 실시간 음성 에이전트 등에 활용 가능합니다.

번역된 본문

Gemini 3.8 텍스트 음성 변환 모델이 인사드립니다 (2026년 9월 23일)

Gemini 3.8 Flash TTS와 Gemini 3.8 Flash-Lite TTS는 지금까지 개발된 가장 표현력이 풍부한 오디오 생성 모델입니다. Google AI Studio, Gemini API, Gemini Enterprise, Gemini Notebook, Google Vids에서 맞춤 캐릭터 음성을 생성하고 장면 대화를 연출할 수 있습니다.

— Leland Rechis (그룹 프로덕트 매니저), Alan Cowen (리서치 사이언스 디렉터), Gemini 오디오 팀

오늘 우리는 Gemini 제품군에 두 개의 새로운 텍스트 음성 변환(TTS) 모델을 선보입니다. 이는 음성 생성을 정적인 프리셋에서 동적인 창작 스튜디오로 탈바꿈시킵니다. 이 모델들은 크리에이터, 개발자, 기업이 더 풍부하고 표현력 있는 오디오 경험을 만들 수 있게 하며, Gemini Notebook과 Google Vids 같은 제품의 사용자 경험도 개선합니다.

Gemini 3.8 Flash TTS: 깊이 있는 창작 지시와 캐릭터 디자인을 위해 설계되었습니다. 자연어 프롬프트로 완전히 새로운 음성을 처음부터 만들어 게임, 몰입형 오디오북, 팟캐스트, 인터랙티브 미디어 전반에서 캐릭터에 생명을 불어넣을 수 있습니다. 연기 지시, 속도, 방언 변화, 백채널링(듣는 사람의 반응 소리)에 대한 세밀한 제어와 함께 한 줄씩 연출할 수 있습니다.

Gemini 3.8 Flash-Lite TTS: 대규모 처리량과 비용 효율성을 위해 설계되었습니다. 대량 더빙, 오디오 콘텐츠 제작, 표현적인 음성 에이전트에 최적화되어 있으며 어조, 속도, 표현의 뉘앙스를 정밀하게 제어할 수 있습니다.

이 모델들은 빠르게 성장하는 Gemini 오디오 제품군인 3.5 Live Translate, 3.5 Transcribe, 3.8 Live, 3.8 Live Extended Thinking에 이어 추가된 것입니다.

나만의 음성 만들기와 맞춤 설정 30개의 기존 음성에서 무한한 음성 라이브러리로 확장하세요. 완전히 새로운 캐릭터 음성이든 일관된 브랜드 앰버서더든, 3.8 Flash TTS 모델이 완전한 보컬 스튜디오를 제공합니다.

  • 생성형 음성 디자인: Gemini 3.8 Flash TTS로 100개 이상의 언어와 방언에 걸쳐 역할, 억양, 음성 특성을 자연어 프롬프트로 맞춤 설정하여 새로운 음성을 처음부터 만들 수 있습니다. 불을 뿜는 드라마틱한 용에게 생명을 불어넣거나 독특한 지역 억양의 매력적인 내레이터를 만들 수도 있습니다. 멜버른 출신의 에너지 넘치는 DJ 음성 생성 예시, 초고음의 단조로운 로봇 음성 생성 예시, 일본식 용 캐릭터 구현 예시도 확인할 수 있습니다.
  • 방대한 음성 라이브러리: 멕시코 스페인어, 퀘벡 프랑스어, 스코트 영어 같은 지역 변형을 포함해 폭넓은 언어를 지원하는 2,000개 이상의 바로 사용 가능한 음성에 접근할 수 있습니다.
  • 음성 복제: 본인 또는 사용 권한이 있는 음성의 30초 오디오 샘플만으로 일관된 음성 프로필을 재현할 수 있으며, 내장된 동의 검증 기능이 함께 제공됩니다.

또한 생성된 오디오의 안전을 지키기 위한 워터마킹 같은 안전 도구가 내장되어 있어, 고품질 오디오북, 팟캐스트, 대규모 실시간 음성 에이전트 등에 책임감 있게 활용할 수 있습니다.

원문 보기
원문 보기 (영어)
Gemini 3.8 text-to-speech says hello Sep 23, 2026 | x.com Facebook LinkedIn Mail Copy link Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS are our most expressive audio generation models yet. Generate custom character voices and direct scene dialogue across Google AI Studio, Gemini API, Gemini Enterprise, Gemini Notebook, and Google Vids. Leland Rechis Group Product Manager Alan Cowen Director, Research Science, on Behalf of the Gemini Audio Team Share x.com Facebook LinkedIn Mail Copy link Read AI-generated summary Check out "Gemini 3.8 text-to-speech says hello" to see how our new models work. Create custom voices from scratch or replicate existing ones with simple natural language prompts. Direct your audio line-by-line to control pacing, emotion, and even realistic conversational sounds. Use these models for high-quality audiobooks, podcasts, or real-time voice agents at scale. We’ve included built-in safety tools like watermarking to keep your generated audio secure. Summaries were generated by Google AI. Generative AI is experimental. Google just launched new AI tools that let you create and customize realistic voices from scratch. You can direct these voices to sound exactly how you want, from their accent to their emotional tone. It’s perfect for making audiobooks, games, or podcasts that sound like real people talking. Plus, they added safety features to make sure these voices are used responsibly. Summaries were generated by Google AI. Generative AI is experimental. Explore other styles: Bullet points Basic explainer Today, we’re introducing two new text-to-speech models to the Gemini family, transforming voice generation from static presets into a dynamic creative studio. These models enable creators, developers, and enterprises to create richer, more expressive audio experiences, while enabling improved user experiences in products like Gemini Notebook and Google Vids . Gemini 3.8 Flash TTS: Built for deep creative direction and character design. Create entirely new voices from scratch using natural language prompts to bring characters to life across gaming, immersive audiobooks, podcasts, and interactive media. Direct every performance line by line with granular control over acting cues, pacing, dialect shifts, and backchanneling. Gemini 3.8 Flash-Lite TTS: Built for high-volume, cost-efficient scale. Optimized for high-volume dubbing, audio content creation, and expressive voice agents with fine-grained control over tone, pacing, and expressive nuance. These models complement our fast-growing Gemini Audio family, following 3.5 Live Translate , 3.5 Transcribe , 3.8 Live, and 3.8 Live Extended Thinking . Create and customize your own voices Scale up from 30 original voices to an infinite library. Whether you need an entirely original character voice or a consistent brand ambassador, our 3.8 Flash TTS model powers a full vocal studio. This enables you to create and use expressive, natural-sounding voices for every moment, while empowering developers and enterprises to easily build custom audio experiences. Generative voice design: With Gemini 3.8 Flash TTS, create bespoke voices from scratch by customizing role, accent and voice characteristics across more than 100 languages and dialects using natural language prompting — whether you're bringing a dramatic, fire-breathing dragon to life or crafting a charismatic narrator with a distinct regional cadence. Hear how Gemini 3.8 Flash TTS generates a high-energy DJ voice from Melbourne. Hear how Gemini 3.8 Flash TTS generates a super-tinny, monotone robot voice. Hear how Gemini 3.8 Flash TTS brings a Japanese dragon to life. Expansive voice library: Access 2,000+ production-ready voices with broad language coverage — including regional varieties like Mexican Spanish, Quebec French, and Scots English. Voice replication: Recreate consistent vocal profiles from just a 30-second audio sample of your voice or a voice you have the rights to use, backed by built-in consent verification, SynthID watermarking, and C2PA credentials to protect both developers and their vocal talent. Save and scale: Save and manage the custom voices you designed to ensure consistent performance and minimal drift across ongoing projects. Voice remixing: Coming soon, pick a voice from our voice library and fine-tune timbre, pitch, pace, and accent. Use prompts to dial in characteristics (e.g. “add subtle Southern US accent” or “soften the delivery”). Direct the performance, line by line Once you've selected your voices, both TTS models give you precise control over how each line is delivered. Direct performance line by line: Write your own stage directions or let Gemini steer delivery with natural script cues — from a calm customer service agent to a whispered suspense scene. Hear how Gemini 3.8 Flash TTS enables natural, highly expressive conversations for interactive voice agents. Watch and hear how Gemini 3.8 Flash TTS uses granular script control to build a deeply engaging, immersive audio experience. Long-form generation: Maintain high voice quality, natural pacing, and character timbre across hours of continuous audio with minimal speaker drift — ideal for podcasts and audiobooks. Native two-speaker scene staging: Direct multi-turn conversations seamlessly from a single script —whether for a podcast or dramatic storytelling—while keeping both voices distinctly separated with natural conversational turn-taking. Scripted vocal bursts & backchanneling: Add realistic conversational texture using non verbal cues (like <laughs>, <sigh>, <gasp> and active-listening interjections (like |mhm| or|yeah|) for precise comedic timing and reaction beats. See how Gemini 3.8 Flash TTS turns natural language prompts into bespoke vocal personas from scratch. Watch how Gemini 3.8 Flash TTS enables creators to design custom scenes to bring animated dialogue to life. See how Gemini 3.8 Flash TTS turns scripts into fully performed dialogue scenes, letting creators direct vocal delivery, and natural turn-taking. Get expressive high-quality speech generation built for global scale Gemini 3.8 Flash TTS delivers leading voice customization capabilities, securing the #1 overall spot on Hume AI’s Voice Design Benchmark (71.4) and also leading in accent modeling (60.8). Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS enable truly expressive performances without sacrificing reliability, also securing the #1 and #2 spots respectively on Hume AI’s Overall Quality Index. The model shows major improvements on a wide range of use cases such as long-form content and dual-speaker screenplay control compared to Gemini 3.1 Flash TTS. In blind human preference evaluations on Voice Arena , Gemini 3.8 Flash and Flash-Lite TTS secure top positions amongst competitors in key global languages, including Japanese, Brazilian Portuguese, Vietnamese, Modern Standard Arabic (MSA), Mexican Spanish and Hindi. With support for over 100 languages, these models empower creators, developers, and enterprises to build high-quality, multilingual voice experiences worldwide. Build with trust, consent, and transparency We built our voice creation and replication capabilities with strict safeguards to help protect voice talent, respect identity, and ensure content transparency. For voice replication our system leverages consent verification: users must provide a verbal consent recording from the voice owner that matches the reference speaker before a voice can be created. More broadly, every audio clip generated by our Gemini Audio models is watermarked with SynthID . This imperceptible watermark is woven directly into the audio output, ensuring AI-generated speech remains detectable to help prevent misinformation. For more details on our approach to safety and responsibility, review the model card . Try our new Google AI Studio audio playground Starting today, developers can experience these new speech generation capabilities in Google AI Studio . Built like a voice desig