메뉴
BL
Google AI Blog • 10일 전

모든 언어를 위한 AI: 300개 언어, 70억 인구를 잇다

IMP
8/10
핵심 요약

구글은 AI 번역 기술을 300개 이상 언어(전 세계 인구의 86%)로 확장하며, 단순 텍스트 번역을 넘어 감정·어조·속어, 코드 스위칭까지 이해하는 원어음 오디오 인텔리전스로 발전시키고 있습니다. Gemini 3.5 Live Translate와 Transcribe, 1,000개 언어 지원 이니셔티브를 통해 데이터가 부족한 언어에도 교차 언어 전이 학습 기법으로 접근성을 높이는 것이 핵심입니다.

번역된 본문

모든 언어로, 모두를 위한 AI

2026년 9월 15일 | 제임스 매니카(James Manyika), Research, Labs, Technology & Society 부문 수석 부사장

AI 생성 요약: 구글은 AI를 활용해 수백 개 언어로 사람들이 소통할 수 있도록 돕고 있으며, 그중에는 그동안 기술에서 소외되었던 언어들도 포함됩니다. 단순히 텍스트를 번역하는 것을 넘어, 사람들이 실제 삶에서 사용하는 감정, 어조, 속어까지 AI가 이해하도록 훈련하고 있습니다. 또한 전 세계 지역 사회와 협력해 안정적인 인터넷이 없는 사람들까지도 도구를 사용할 수 있도록 하고 있습니다. 이를 통해 기술이 다양한 문화를 존중하고, 모든 사람이 자신의 방식대로 이해받을 수 있도록 돕습니다.

오늘날 우리의 기술과 제품은 전 세계 인구의 86%에 해당하는 70억 명 이상이 사용하는 300개 이상의 언어로 일상적인 소통을 지원합니다. 이정표를 달성한 것은 의미가 있지만, 동시에 우리 미션에 필수적인 작업이 아직 남아 있음을 보여줍니다. 수십 년 동안 기술은 소수의 지배적 언어에서만 잘 작동했고, 수천 개의 살아있는 언어와 방언은 디지털 세계에서 제대로 대표되지 못하거나 아예 존재하지 않았습니다.

2006년 구글 번역(Google Translate)을 출시했을 때 우리의 목표는 단순했습니다. 언어 간 장벽을 허무는 것이었습니다. AI의 발전 덕분에 이 비전을 더 많은 사람들에게 확장할 수 있었고, 번역은 소수 언어에서 오늘날 250개 이상의 언어로 늘어났습니다. 하지만 텍스트를 번역하는 것만으로는 부족합니다. 기술은 사람들이 실제 세계에서 어떻게 소통하는지 이해해야 합니다. 그래서 우리는 문화적 뉘앙스와 인간 언어의 풍부함을 존중하는 시스템을 구축하는 데 연구개발을 집중하여, 모든 사람이 자신의 방식대로 참여하고 이해받을 수 있도록 하고 있습니다. 실제로 이런 작업이 어떻게 이루어지는지 소개합니다.

텍스트에서 진정한 이해로 지난날의 음성 인식 시스템은 딱딱한 다단계 과정을 따랐습니다. 오디오를 텍스트로 전사하고, 그 텍스트를 처리한 다음, 다시 오디오로 합성하는 방식이었습니다. 작동은 했지만, 이 파이프라인은 인간 소통에서 가장 풍부한 요소, 즉 어조, 속도, 감정, 맥락을 제거해 버렸습니다. 사람들은 완벽하게 정돈된 문법적 문장으로 말하지 않습니다. 우리는 웃고, 말을 겹치고, 머뭇거리고, 문장 중간에 여러 언어를 섞어 씁니다. 스팽글리시(Spanglish)나 힝글리시(Hinglish)를 말할 때처럼요.

이를 반영하기 위해 우리는 텍스트 전사를 넘어 네이티브 오디오 인텔리전스로 나아갔습니다. Gemini 같은 모델이 오디오를 있는 그대로 직접 처리하면서 소리와 의도를 모두 파악하도록 훈련한 것입니다. 이러한 노력에는 다음이 포함됩니다:

유려한 실시간 대화 도구: 오늘날 Gemini 3.5 Live Translate는 70개 언어와 2,000개 이상의 언어 쌍에서 실시간 음성 번역을 지원하며, 그 과정에서 코드 스위칭(언어 혼용)과 감정 신호까지 자연스럽게 포착합니다. Gemini 3.5 Transcribe는 지금까지 개발된 가장 정확한 음성-텍스트 변환 모델로, 시끄러운 환경이나 복잡한 전문 용어가 포함된 경우에도 원시 오디오를 다듬어진 정형화된 텍스트로 변환합니다. 또한 안드로이드 Gboard의 'Rambler' 기능을 지원하는데, 이 기능은 군더더기 표현을 제거하고 문법과 문장 부호를 고쳐주며, 음성 명령으로 편집·재작성하고 언어 간 매끄럽게 전환할 수 있게 합니다.

1,000개 언어 이니셔티브: AI는 이전에는 상상할 수 없었던 규모로 언어 장벽을 허물고 있습니다. 하지만 더 많은 사람들이 선호하는 언어로 다가가려면, 오늘날 AI가 가장 잘 수행하는 언어를 넘어서야 합니다. 우리의 목표는 세계에서 가장 많이 사용되는 1,000개 언어를 지원하는 것입니다. 이를 실현하기 위해 1,200만 시간의 오디오로 훈련된 유니버설 음성 모델(Universal Speech Model)은 교차 언어 전이 학습(cross-lingual transfer learning) 기법을 활용했습니다. 이는 데이터가 풍부한 언어에서 학습한 내용을 훈련 데이터가 훨씬 부족한 언어의 음성 이해력 향상에 전이할 수 있게 하는 기술로, 모델이 데이터가 풍부한 언어에서 배운 패턴을 적용할 수 있게 합니다.

원문 보기
원문 보기 (영어)
AI for everyone in every language Sep 15, 2026 | x.com Facebook LinkedIn Mail Copy link We’re moving beyond traditional text translation to build models that understand the world’s rich, living languages exactly as they are expressed. James Manyika SVP, Research, Labs, Technology & Society Share x.com Facebook LinkedIn Mail Copy link Read AI-generated summary Google is using AI to help people communicate in hundreds of languages, including those that were previously left out of technology. Instead of just translating text, they’re training AI to understand the emotion, tone, and slang people actually use in real life. They’re also working with local communities around the world to make sure their tools work for everyone, even those without reliable internet. This helps make sure that technology respects different cultures and helps people be understood on their own terms. Summaries were generated by Google AI. Generative AI is experimental. Today, our technologies and products power everyday interactions in more than 300 languages, spoken by more than 7 billion people — representing 86% of the global population. Reaching this milestone is meaningful, but it also underscores work that is critical to our mission . For decades, technology has worked best for a handful of dominant languages, leaving thousands of living languages and dialects poorly represented or absent altogether from the digital world. When we launched Google Translate in 2006, our goal was simple: to break down the barriers between languages. Advances in AI have helped us bring that vision to more people, expanding Translate from a handful of languages to more than 250 today. But translating text isn’t enough. Technology needs to understand how people actually communicate in the real world. So we focus our research and development on building systems that honor cultural nuance and the richness of human language, enabling everyone to participate and be understood on their own terms. Here’s what that work looks like in practice. Going from text to true understanding Historically, speech recognition systems followed a rigid, multi-step process: transcribing audio into text, processing that text, and then synthesizing it back into audio. While functional, this pipeline strips away the richest parts of human communication: tone, pacing, emotion, and context. People don't speak in perfectly neat, grammatical sentences. We laugh, overlap, hesitate, and weave multiple languages together mid-sentence, like when we speak Spanglish or Hinglish. To capture this, we moved beyond text transcripts to native audio intelligence — training models like Gemini to process audio directly as is, while also grasping both sound and intent. These efforts include: Fluid real-time dialogue tools: Today, Gemini 3.5 Live Translate powers real-time spoken translation across 70 languages and 2,000+ language pairs, naturally capturing code-switching and emotional cues along the way. Gemini 3.5 Transcribe is our most precise speech-to-text model yet, turning raw audio into polished, formatted text, even in noisy environments or with complex jargon. It also powers features like Rambler on Android Gboard, which removes filler words, fixes grammar and punctuation, and lets you edit or rewrite with voice commands and switch seamlessly between languages. The 1,000 Languages Initiative: AI is helping us break down language barriers at a scale that was previously unimaginable. But reaching more people in their preferred language means going beyond the languages where AI performs best today: Our goal is to support the world’s 1,000 most-spoken languages. To help make that possible, our Universal Speech Model — trained on 12 million hours of audio — used cross-lingual transfer learning , techniques that enable models to transfer what they learn from data-rich languages, to improve speech understanding in languages with far less training data. This allows models to apply patterns learned from data-rich languages to under-resourced ones. Rigorous foundational research: This work builds on 25 years of open research and more than 400 peer-reviewed speech papers , which have helped push the frontier and advance speech models. Putting communities at the heart of language data Because the web disproportionately represents a few dominant languages, teaching AI to understand underrepresented languages required us to rethink how we gather data. The solution is local grassroots partnerships. This localized approach has driven three of our most ambitious open-data partnerships: WAXAL (Wolof for “speaking,” pronounced "Wah-hal"): Built with partners including Makerere University and Digital Umuganda, WAXAL is a large-scale, open speech dataset covering 27 Sub-Saharan African languages spoken by more than 100 million people across more than 26 countries, capturing tonal variation and conversational rhythms often missing from traditional datasets. Project Vaani : In partnership with the Indian Institute of Science (IISc) and Bhashini, Project Vaani is mapping India’s linguistic diversity through a region-anchored rather than language-anchored approach, enabling it to collect to date more than 30,000 hours of speech across 109 languages from more than 155,000 speakers. Amplify Initiative : We teamed up with more than 1,600 local experts and 20 universities across four continents, including Brazil’s UFMG, India’s IIT Kharagpur, and Uganda’s Makerere University, to contribute 15,000 multimodal data points capturing local nuance. We’re also building on our work prioritizing open-source language innovation through our new tool Language Explorer . It’s an interactive tool that visualizes LinguaMeta, the world’s largest open-source language data repository. Recognized by Fast Company for design innovation , it continuously maps more than 7,000 spoken, written, and signed languages. The impact of these innovations and partnerships is greatest when they reach the people who can turn new data and insights into meaningful change in their communities. Google.org-supported efforts, including the Centre for Digital Language Inclusion and AI Singapore’s Project Aquarium , are helping bring multilingual tools to farmers, healthcare workers, teachers, and other essential community members around the world. Overcoming real-world constraints For more than 3 billion people 1 , reliable internet access is still out of reach. Technology is only truly accessible if it works where people live, including areas with limited or intermittent connectivity. To help address this, we developed TranslateGemma , a family of lightweight open translation models built from Gemini and trained across 55 languages. Because TranslateGemma runs efficiently on-device, high quality translation no longer requires a connection to the cloud or the internet. Still, running powerful AI models requires capable hardware, which excludes the hundreds of millions of people still using feature phones in low-resource regions. To bridge this divide, we’re supporting organizations like Viamo to power “Ask Viamo Anything” (AVA), a voice AI assistant that brings the power of Gemini to standard feature phones. Viamo has successfully piloted AVA in Rwanda with its existing interactive voice response users, and the service has already used Gemini to answer more than 2 million questions. Designing for accessibility Language isn’t just about regional dialects or vocabulary. It’s also about the many other ways people communicate. Conventional speech tools frequently fail people with non-standard speech, making them adapt to the technology rather than the other way around. We’re working to change that by designing for accessibility from the ground up, for example, with Sign Language-to-Text (SL2T) . Trained across 50+ sign languages, SL2T powers sign-to-text dictation in Gboard and Live Transcribe on Pixel 11, starting with American Sign Language (ASL) to English. This is an important first step toward ma