메뉴
BL
The Decoder • 2일 전

구글, 텍스트 설명만으로 AI 음성을 만드는 Flash TTS 모델 공개

IMP
8/10
핵심 요약

구글이 100개 이상 언어를 지원하는 Gemini 3.8 Flash TTS와 Flash-Lite TTS 두 음성 생성 모델을 공개했습니다. Flash TTS는 텍스트 설명만으로 새로운 음성을 처음부터 설계할 수 있고, 30초 음성 샘플로 음성 복제도 가능합니다. 두 모델 모두 대사별 연출 지시, 두 음성 대화 모드, 웃음·한숨 같은 비언어적 표현을 지원하며 Gemini API와 Google AI Studio를 통해 제공됩니다.

번역된 본문

구글이 텍스트 설명만으로 AI 음성을 처음부터 설계할 수 있는 새로운 Flash TTS 모델 공개 마티아스 바스티안 | 2026년 9월 23일

핵심 요약

  • 구글이 100개 이상 언어를 지원하는 두 개의 새로운 텍스트-투-스피치(TTS) 모델, Gemini 3.8 Flash TTS와 Flash-Lite TTS를 공개했다.
  • Flash TTS는 텍스트 설명만으로 새로운 음성을 만들 수 있으며, 30초짜리 오디오 샘플로 음성 프로필을 생성하는 음성 복제 기능도 제공한다.
  • Flash TTS는 팟캐스트, 오디오북, 게임 캐릭터 등 창작 용도를 겨냥하고, Flash-Lite TTS는 더빙, 오디오 콘텐츠, 음성 에이전트를 위한 대규모 저비용 음성 생성용으로 설계되었다.
  • 두 모델 모두 대사별 연출 지시(stage direction), 두 음성 대화, 웃음이나 한숨 같은 비언어적 소리를 지원한다.
  • Gemini API와 Google AI Studio를 통해 순차 공개되며, 이후 Gemini Enterprise API 접근도 이어질 예정이다.

구글은 음성 생성을 위한 두 개의 새로운 모델인 Gemini 3.8 Flash TTS와 Flash-Lite TTS를 선보였다. Flash TTS는 텍스트 설명만으로 새로운 음성을 만들 수 있으며, 두 모델 모두 100개 이상 언어를 지원하고 사용자가 대화 대사 각 줄에 연출 지시를 추가할 수 있다.

구글에 따르면 Gemini 3.8 Flash TTS는 게임 캐릭터, 오디오북, 팟캐스트 같은 창작 프로젝트용으로 설계됐고, Gemini 3.8 Flash-Lite TTS는 더빙, 오디오 콘텐츠, 음성 에이전트를 위한 대규모 저비용 음성 생성에 초점을 맞추고 있다.

Flash TTS로 텍스트 설명만으로 음성 생성 Gemini 3.8 Flash TTS를 사용하면 사용자가 음성을 처음부터 설계할 수 있다. 구글에 따르면 텍스트 프롬프트 하나로 다양한 언어와 방언에 걸쳐 음성의 역할, 억양, 발성 특징을 정의할 수 있다. 처음부터 시작하고 싶지 않은 사용자를 위해 구글은 멕시코식 스페인어, 퀘벡 프랑스어, 스코틀랜드 영어 같은 지역 변형을 포함해 2,000개 이상의 프리셋 음성 라이브러리도 제공한다.

음성 복제 기능은 30초짜리 오디오 샘플로 음성 프로필을 만들 수 있다. 이 기능을 사용하려면 복제되는 목소리의 당사자가 음성 동의 발언을 녹음해야 하며, 해당 녹음의 목소리가 샘플과 일치해야 한다. 구글에 따르면 Gemini 오디오 모델이 생성하는 모든 클립에는 AI 생성 음성을 감지하는 데 도움이 되는 청각적으로 들리지 않는 SynthID 워터마크가 포함된다.

구글은 라이브러리 음성의 음색, 높낮이, 속도, 억양을 조정할 수 있는 '보이스 리믹싱(Voice Remixing)' 기능도 발표했지만, 아직 사용할 수는 없다.

스크립트 연출 지시로 대화와 전달 방식 제어 두 모델 모두 사용자가 각 대사에 대한 지시를 작성하거나 모델이 스스로 스크립트 단서를 해석하게 할 수 있다. 구글은 두 모델이 최소한의 '스피커 드리프트(화자 변화)'로 수 시간 분량의 오디오를 생성할 수 있다고 말했다. 즉, 시간이 지나도 음성이 거의 변하지 않는다는 뜻이다.

회사에 따르면 두 음성 모드는 하나의 스크립트에서 대화를 생성하면서도 각 음성을 구별되게 유지한다. 사용자는 웃음, 한숨, "음흠" 같은 소리도 스크립트에 넣어 반응과 멈춤을 원하는 정확한 위치에 배치할 수 있다.

기자가 직접 진행한 두 번의 테스트에서는 프리셋 음성에 짙은 독일 억양으로 영어를 말하는 짜증 난 베를린 시민을 흉내 내달라는 스타일 프롬프트를 사용했다. 스타일 제어 기능은 설득력 있는 억양과 억양새를 만들어냈지만, 두 테스트 모두 일부 구간에서 배경에 고음의 잡음이 섞였다. 그중 하나에서는 클립 끝에서 목소리가 바뀌기도 했다.

스타일 프롬프트 예시: 외국어로 영어를 말하는 베를린 출신 독일인 남성. 짙고 확실한 독일 억양으로 말하며, 명백히 원어민이 아니다. 영어 단어를 독일식으로 발음하고, 독일어의 리듬과 억양새를 영어 문장에 적용한다. "Th"는 "z"나 "d"로 발음하고("ze", "sink", "dat"), "w"는 "v"로("vat", "vell"), 단어 끝 자음은 더 강하게 발음한다("bad"를 "bat"처럼). "r"은 목구멍에서 나는 걸걸한 소리이며 영어의 "r"이 결코 아니다. 모음은 평평하고 짧아서 미국식이나 영국식 영어의 부드러움이 없다. 목소리는 비음이 섞이고 약간 힘이 들어간 중음역대로, 허스키한 질감이 있다.

원문 보기
원문 보기 (영어)
Google's new Flash TTS models let you design AI voices from scratch using text descriptions Matthias Bastian View the LinkedIn Profile of Matthias Bastian Sep 23, 2026 Nano Banana Pro prompted by THE DECODER Key Points Google has released two new text-to-speech models, Gemini 3.8 Flash TTS and Flash-Lite TTS, which support more than 100 languages. Flash TTS can create new voices from text descriptions, and a voice cloning feature builds voice profiles from 30-second audio samples. Flash TTS is aimed at creative uses such as podcasts, audiobooks, and game characters, while Flash-Lite TTS is designed for low-cost speech generation at scale for dubbing, audio content, and voice agents. Both models support stage directions for each line, two-voice dialogue, and nonverbal sounds like laughter and sighs. They're rolling out through the Gemini API and Google AI Studio, and Gemini Enterprise API access will follow. Ask about this article… Search Google is introducing Gemini 3.8 Flash TTS and Flash-Lite TTS, two new models for speech generation. Flash TTS can create new voices from text descriptions, and both models support more than 100 languages and let users add stage directions to individual lines of dialogue. Gemini 3.8 Flash TTS is designed for creative projects such as game characters, audiobooks, and podcasts, while Gemini 3.8 Flash-Lite TTS focuses on low-cost speech generation at scale for dubbing, audio content, and voice agents, according to Google. Flash TTS lets users create voices from text descriptions With Gemini 3.8 Flash TTS, users can design voices from scratch. According to Google, a text prompt can define a voice's role, accent, and vocal traits across a wide range of languages and dialects. For users who don't want to start from zero, Google offers a library of more than 2,000 preset voices, including regional variants such as Mexican Spanish, Quebec French, and Scottish English. Ad A voice cloning feature can build a voice profile from a 30-second audio sample. To use it, the person whose voice is being cloned has to record a spoken statement of consent, and the voice in that recording must match the sample. Every clip the Gemini audio models generate carries an inaudible SynthID watermark to help detect AI-generated speech, according to Google . Ad Google has also announced "Voice Remixing," a feature that will let users adjust the timbre, pitch, tempo, and accent of library voices, but it isn't available yet. Script directions give users control over dialogue and delivery Both models let users write directions for each line or have the model interpret script cues on its own. Google says the models can generate hours of audio with minimal "speaker drift," meaning the voice barely changes over time. Ad A two-voice mode generates dialogue from a single script while keeping the voices distinct, according to the company. Users can also script laughter, sighs, and sounds like "mhm" to place reactions and pauses exactly where they want them. In two of my own tests, I used a preset voice with a style prompt asking it to imitate an annoyed Berliner speaking English with a thick German accent. The style controls produced a convincing accent and intonation, but both tests had a high-pitched whine in the background at some points. In one of them, the voice also changed at the end of the clip. Ad Ad Style prompt: A native German man from Berlin speaking English as a foreign language, with a thick, unmistakable German accent. He is clearly not a native English speaker: he pronounces English words the German way, applying German rhythm and intonation to English sentences. "Th" becomes "z" or "d" ("ze," "sink," "dat"), "w" becomes "v" ("vat," "vell"), and final consonants become harder ("goot," "bat" for "bad"). The "r" is guttural and throaty, never the English "r." Vowels are flat and short, lacking the softness found in American or British English. The voice is nasal and slightly strained, in the mid-range, with a raspy quiver and a Kermit-like wobble, but grittier and less puppet-like. It sounds like a tired Berliner at 2 a.m. Delivery: quick, clipped, choppy. He hesitates briefly when searching for an English word and sometimes inserts the German word flatly without translating it. Occasional exasperated upward pitch breaks. Dry, sardonic, unimpressed, and slightly annoyed that he has to explain anything at all—and even more annoyed that he has to do it in English. Google starts rolling out both models, with enterprise API access to follow Google is rolling out both models through the Gemini API and Google AI Studio . Flash TTS is also available in Gemini Notebook , while Flash-Lite TTS is available in Google Vids . Google says access through the Gemini Enterprise API will follow soon for both models. Developer platforms including Agora , LiveKit , Pipecat , and Vercel already support integration through the Gemini API. Google hasn't listed regional endpoints for the new models yet, though earlier TTS models offered EU data processing . According to Google, data from the free tier is used to improve its products, while data from the paid tier isn't. Paid pricing is listed in US dollars per million tokens, with text tokens billed for input and audio tokens for output. Billing Flash TTS (through the end of 2026) Flash TTS (starting in 2027) Flash-Lite TTS (through the end of 2026) Flash-Lite TTS (starting in 2027) Text Input $0.50 $1.00 $0.50 $1.00 Audio Output $9.00 $18.00 $6.00 $12.00 Google says one second of generated audio equals 25 audio tokens, which puts an hour at 90,000 tokens. That means an hour of audio output costs $0.81 with Flash TTS and $0.54 with Flash-Lite TTS through the end of 2026. On January 1, 2027, those costs rise to $1.62 and $1.08, with text input billed separately. AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now Source: Google
관련 소식