메뉴
BL
TechCrunch AI 20일 전

오픈AI, 동시 대화 가능한 새 음성 모델 공개

IMP
9/10
핵심 요약

오픈AI가 사용자의 말을 끊고 자연스럽게 대화할 수 있는 풀 듀플렉스(Full-duplex) 기반의 새로운 음성 모델 'GPT-Live-1'과 'mini'를 발표했습니다. 이 모델들은 실시간 번역 기능을 지원하며 대화 중단점 처리가 개선되어, 장시간 프롬프트 없이 복잡한 작업을 수행하는 핵심 인터페이스로 자리 잡을 전망입니다. 챗GPT(ChatGPT)의 기본 음성 모드도 이 중 미니(Mini) 모델로 교체되어, 모든 사용자가 훨씬 자연스러운 AI 음성 대화를 경험할 수 있게 되었습니다.

번역된 본문

오픈AI는 오늘 더 자연스럽고 대화의 순서를 더 잘 처리할 수 있는 'GPT-Live-1'과 'GPT-Live-1 mini'라는 새로운 대화형 모델을 공개했습니다. 이들은 풀 듀플렉스(Full-duplex, 전이중 통신 방식) 모델입니다. 즉, 동시에 말하고 들을 수 있어 사용자가 자연스럽게 대화를 끊고 들어갈 수 있으며 실시간 번역 기능도 가능하게 합니다. 또한 회사는 기존 챗GPT(ChatGPT)의 고급 음성 모드(Advanced Voice Mode)를 기본적으로 GPT-Live-1 mini로 교체하고 있습니다. 유료 결제 사용자는 더 큰 규모의 GPT-Live-1 모델에 접근할 수 있습니다.

이전 모델은 음성을 텍스트로 변환하는 음성 인식 모델, 답변을 생성하는 대형 언어 모델, 최종 답변을 음성으로 전달하는 텍스트 음성 변환 모델을 결합한 형태였습니다. 회사는 언론 브리핑에서 새로운 모델이 사용자가 말하는 동안 끼어드는 문제와 질문에 답변할 만큼 충분한 지능이 부족했던 문제를 해결했다고 밝혔습니다. 오픈AI의 새 모델은 대화를 나누는 동시에 검색, 추론 또는 에이전트(Agentic) 기능을 위해 최신 텍스트 모델인 GPT-5.5 등에 질의를 보냅니다. 또한, 오픈AI는 이 모델이 부름을 받을 때까지 오랜 시간 침묵을 유지하며 대화의 맥락을 파악할 수 있음을 보여주었습니다. 게다가 새로운 음성 모드는 최신 GPT 모델에 액세스할 수 있어 일부 정보를 시각적 형태로 제공할 수도 있습니다. DST와 럭스 캐피탈(Lux Capital)로부터 4천만 달러의 시드 자금을 조달한 모노그램(Monogram)과 같은 다른 스타트업들 역시 어시스턴트를 더 상호작용적으로 만들기 위해 시각적 답변에 주력하고 있습니다.

회사는 챗GPT의 새로운 음성 모드가 더 긴 대화를 나누도록 설계되었다고 밝혔습니다. 브리핑 도중 챗GPT 보이스의 프로덕트 책임자인 아티 엘레티(Atty Eleti)는 산책 중에 이 음성 기능을 이용해 30~40분 동안 대화를 나눴다고 말했습니다. 오픈AI는 음성이 복잡한 작업을 위한 컴퓨팅의 주요 인터페이스가 될 수 있다고 생각합니다. 보도에 따르면 올해 AI 기능이 탑재된 이어폰을 출시할 수도 있지만, 이번 브리핑에서는 하드웨어 제품에 대한 어떠한 정보도 제공하지 않았습니다.

엘레티는 "시간이 지나면서 이것이 음성을 일종의 주요 컴퓨팅 인터페이스로 사용하고, 점점 더 복잡해지는 장기 에이전트 작업을 관리할 수 있는 능력을 잠금 해제할 것이라고 생각합니다. 우리가 사람들이 코덱스(Codex)와 챗GPT를 사용하여 달성하는 놀라운 사용 사례를 보듯, 음성은 모든 종류의 작업에 대한 미래의 인터페이스가 될 수 있다고 생각합니다."라고 말했습니다.

오픈AI는 지난 몇 년 동안 챗GPT의 음성 모드가 더 자연스럽게 들리도록 음성 기반 기능을 강화하는 작업을 진행해 왔습니다. 회사에 따르면 1억 5천만 명 이상이 음성(Voice) 및 받아쓰기(Dictation)와 같은 기능을 사용하여 챗GPT와 대화합니다.

경쟁사들 역시 어시스턴트를 더 표현력 있게 만들려고 시도하고 있습니다. 애플(Apple)과 아마존(Amazon)은 모두 맥락 처리를 개선하여 어시스턴트가 더 대화형으로 업데이트했습니다. 오큘러스(Oculus)의 공동 창립자인 브렌든 아이리브(Brendan Iribe)와 앙킷 쿠마르(Ankit Kumar)가 설립한 세same과 같은 스타트업은 백그라운드에서 작업을 완료하면서도 더 자연스러운 대화를 나눌 수 있는 AI 어시스턴트를 출시했습니다. 오픈AI 역시 같은 방향으로 나아가며 사용자가 오랫동안 핸즈프리로 어시스턴트와 대화할 수 있도록 하는 것을 목표로 하고 있습니다.

새로운 음성 모드가 더 자연스럽게 들린다고 주장함에도 불구하고, 회사는 이것을 AI 동반자(AI companion)로 만드는 것을 목표로 하지 않는다고 강조했습니다. 회사는 새로운 모델에 청소년에게 연령에 맞는 답변을 제공하고, 자해와 같은 주제로 대화가 흘러갈 경우 관련 리소스를 제공하기 위한 안전장치가 내장되어 있다고 언급했습니다.

물론 새로운 음성 모드는 여전히 개선의 여지가 남아있습니다. 데모 당시, 회사가 힌디어로 실시간 번역 기능을 선보였을 때 어시스턴트는 짙은 미국식 억양으로, 다소 부자연스럽고 책에서 읽는 듯한 톤의 힌디어를 구사했습니다. 회사는 새로운 모드가 '가장 많이 사용되는 언어'에 최적화되어 있다고 밝혔지만, 구체적으로 어떤 언어인지는 명시하지 않았습니다.

원문 보기
원문 보기 (영어)
OpenAI today released new conversational models, called GPT-Live-1 and GPT-Live-1 mini, claiming that they sound more natural and can handle turn-taking better. These are full-duplex models, meaning they can speak and listen at the same time, allowing users to interrupt naturally and enabling features like live translation. The company is also replacing its current Advanced Voice Mode in ChatGPT with GPT-Live-1 mini by default. Users of paid tiers will be able to access the larger GPT-Live-1 model. The previous model combined a speech-to-text model to transcribe speech, a large language model to generate responses, and a text-to-speech model to deliver the final answer. The company said in a press briefing that the new models solve issues like interrupting users while they're talking and not having enough intelligence to answer questions. OpenAI's new models will send the query to its latest text models like GPT-5.5 for search, reasoning, or agentic capabilities while continuing the conversation. OpenAI also showed that the model can stay silent for a long time and absorb the context of the conversation until it's called upon. Plus, as the new voice mode has access to newer GPT models, it can also present some information in a visual format. Other startups like Monogram, which raised $40 million in seed funding from DST and Lux Capital , are also leaning into visual responses to make assistants more interactive. The company said the new voice mode in ChatGPT is designed to have longer conversations. During the briefing, ChatGPT Voice's product lead, Atty Eleti, said he has had 30- to 40-minute-long conversations with the voice feature during walks. OpenAI thinks that voice could be the primary interface to computing for complex work. Reports have suggested that it could launch a pair of earbuds with AI capabilities this year . However, it didn't provide any information on hardware products. "Over time, we think this will also unlock the ability to use voice as a kind of primary interface to computing, and to manage increasingly complex long-running agentic work. The kind of amazing use cases that we see people using Codex and ChatGPT to accomplish, we think voice can be the future interface to all kinds of work," Eleti said. OpenAI has worked on bolstering voice-based features over the past few years to make ChatGPT's voice mode sound more natural. The company said that more than 150 million people talk to ChatGPT using features like Voice and Dictation. Rivals are also attempting to make assistants more expressive. Both Apple and Amazon have updated their assistants to be more conversational with better context handling. Startups like Sesame , founded by Oculus co-founder Brendan Iribe and Ankit Kumar, also launched AI assistants with more natural conversation while completing tasks in the background. OpenAI is moving in the same direction, aiming to let users talk to its assistant hands-free for a longer time. Despite its claim that the new voice mode sounds more natural, the company emphasized that it's not aiming to make this an AI companion. It noted that the new models have safeguards built in to give age-appropriate responses to teens and provide resources if the conversation turns to topics like self-harm. The new voice mode still needs work. During the demo, when the company showed its live translation feature in Hindi, the assistant had a heavy American accent and spoke in Hindi that was unnatural sounding and had slightly bookish tone. The company said the new mode is optimized for "most spoken languages" but didn't specify which ones. Topics AI , AI , Apps , ChatGPT , voice model When you purchase through links in our articles, we may earn a small commission . This doesn’t affect our editorial independence. Ivan Mehta Ivan covers global consumer tech developments at TechCrunch. He is based out of India and has previously worked at publications including Huffington Post and The Next Web. You can contact or verify outreach from Ivan by emailing im@ivanmehta.com or via encrypted message at ivan.42 on Signal. View Bio November 4 Boston Last chance to save up to $190 on TechCrunch Founder Summit. Join 1,000+ founders and VCs at all stages for real-world scaling insights and connections that move the needle. Savings end June 26, 11:59 p.m. PT . REGISTER NOW Most Popular If you use Google, you're training its AI. Here's how to opt out. Sarah Perez Reddit is using LLMs to solve a problem LLMs largely created Amanda Silberling Amazon will stop accepting new customers for Mechanical Turk Anthony Ha 5 desk gadgets that can make your workday better Aisha Malik Chevy built an all-American EV truck — why is nobody buying it? Tim De Chant Mark Zuckerberg tells staff that AI agents haven't progressed as quickly as he'd hoped Lucas Ropek Jersey Mike's IPO illustrates how bad the AI hype has become Julie Bort