메뉴
BL
The Decoder 20일 전

동시에 듣고 말하는 오픈AI 'GPT-Live' 등장

IMP
9/10
핵심 요약

OpenAI가 사람처럼 대화하며 듣고 말하는 동시 처리가 가능한 새로운 음성 모델 GPT-Live를 공개했습니다. 이 모델은 대화를 주도하는 역할만 하며, 복잡한 질문이나 검색이 필요한 작업은 백그라운드에서 GPT-5.5 모델에 위임하여 처리함으로써 기존 음성 AI의 지능적 한계를 극복했습니다. 실시간 자연스러운 대화와 최고 수준의 추론 능력을 결합함으로써 AI 음성 비서의 혁신을 이끌고 있다는 점에서 매우 중요합니다.

번역된 본문

ChatGPT가 이제 동시에 듣고 말할 수 있게 되어, AI 대화가 더욱 인간적으로 느껴지도록 만들었습니다. 마티아스 바스티안 (Matthias Bastian) - 2026년 7월 8일

핵심 요약: OpenAI는 채팅GPT(ChatGPT)에서 라이브 대화를 위한 새로운 언어 모델인 'GPT-Live'를 도입했습니다. 이 모델은 풀 듀플렉스(Full-duplex) 아키텍처 덕분에 듣고 말하는 것을 동시에 할 수 있습니다. 이 모델은 사용자의 말 끊김(인터럽트)에 반응하며, 대화가 자연스럽게 흘러가도록 '어어', '네' 같은 추임새를 사용합니다. 유료 사용자용과 무료 사용자용 두 가지 버전으로 제공됩니다. 웹 검색이나 논리적 추론이 필요한 복잡한 작업의 경우, GPT-Live는 대화를 유지하면서 백그라운드에서 GPT-5.5 모델로 요청을 전달하여 답변의 품질을 크게 향상시킵니다.

OpenAI는 풀 듀플렉스 아키텍처를 적용한 차세대 음성 모델, GPT-Live를 선보였습니다. 이 모델은 듣고 말하는 것을 동시에 수행할 수 있으며, 복잡한 작업은 백그라운드에서 GPT-5.5에 넘깁니다.

GPT-Live를 통해 OpenAI는 채팅GPT와의 대화가 실제 사람과 대화하는 것처럼 느껴지도록 설계된 새로운 클래스의 AI 음성 모델을 출시했습니다. 두 가지 버전이 전 세계적으로 즉시 배포되고 있습니다. GPT-Live-1은 Go, Plus, Pro 요금제를 사용하는 유료 사용자를 위한 것이며, GPT-Live-1 mini는 무료 계정에서 사용할 수 있습니다. 두 모델 모두 iOS, Android, ChatGPT.com에서 작동합니다. OpenAI는 조만간 API 액세스를 추가할 계획이며, 개발자는 양식을 통해 신청할 수 있습니다.

과거의 사용자가 말하고 AI가 대답하는 경직된 주고받기 방식과 달리, GPT-Live는 풀 듀플렉스 아키텍처를 사용합니다. 이 모델은 사람 간의 실제 대화처럼 듣는 동시에 말할 수 있습니다. 올해 초, 엔비디아(Nvidia)는 PersonaPlex라는 유사한 오픈소스 모델을 출시한 바 있습니다.

OpenAI에 따르면 GPT-Live는 말할지, 계속 들을지, 멈출지, 대화를 끊고 말할지, 도구를 호출할지 결정하기 위해 1초에 여러 번 의사 결정을 내립니다. 이 모델은 대화를 잘 따라가고 있다는 신호를 보내기 위해 '네' 또는 '알겠습니다'와 같은 추임새를 사용할 수 있습니다. 사용자는 말을 끊고 들어가거나, 생각할 시간을 잠깐 가지거나, 모델에게 천천히 말해달라고 요청할 수 있습니다.

OpenAI는 이전의 '고급 음성 모드(Advanced Voice Mode)'와 비교했을 때, 사람들은 75.7%의 경우에서 GPT-Live-1을 선호했고 69.2%의 경우에서 GPT-Live-1 mini를 선호했다고 밝혔습니다.

새로운 대화 기능 외에도 이번 업데이트는 날씨, 주가, 스포츠 점수 등과 같은 정보를 보여주기 위해 채팅GPT 음성이 대화 중에 표시할 수 있는 시각적 카드를 추가했습니다. 또한 OpenAI는 GPT-Live에 사용할 수 있는 9가지 음성을 전면 개선했습니다. 출시와 동시에 GPT-Live는 채팅GPT에서 비디오나 화면 공유 기능이 있는 음성을 지원하지 않지만, OpenAI는 해당 기능이 곧 추가될 예정이라고 밝혔습니다. 이러한 기능이 포함된 기존의 표준 및 고급 음성 모드는 당분간 유지됩니다.

GPT-5.5에 대한 백그라운드 위임, 지능 격차 해소 가장 큰 변화는 실시간 대화 관리와 실제 추론 기능의 분리입니다. 질문에 웹 검색, 추론 또는 에이전트와 같은 기능이 필요할 경우, GPT-Live는 이를 백그라운드 모델인 현재의 GPT-5.5에 넘깁니다. 백그라운드 모델이 작동하는 동안에도 GPT-Live는 대화를 계속 이어갑니다. OpenAI는 GPT-Live가 항상 최신 최첨단 모델에 연결된 상태를 유지하도록 아키텍처가 구축되었다고 밝혔습니다. 또한 사용자는 빠른 답변을 위한 '즉시(Instant)'부터 채팅GPT가 더 많은 시간을 들여 생각해야 하는 작업을 위한 '중간(Medium)' 및 '높음(High)'까지 추론 수준을 선택할 수 있습니다.

이러한 위임 방식은 이전 실시간 모델의 주요 약점을 해결합니다. 기존 모델은 자체 기능만을 바탕으로 질문에 대답했습니다. 이로 인해 최첨단 모델들보다 크게 뒤처졌고 심각한 작업에는 거의 사용할 수 없었습니다. GPT-Live를 통해 이러한 격차가 좁혀지고 있으며, 지식에 대한 벤치마크 점수가 비약적으로 향상되었습니다. 과학적 추론을 위한 GPQA 테스트에서 GPT-Live-1은 높은 추론 수준에서 84.2%의 정확도를 기록한 반면, 기존 고급 음성 모드는 45.3%에 그쳤습니다. 에이전트 기반 웹 검색 벤치마크인 BrowseComp에서 그 격차는 훨씬 더 큽니다. GPT-Live-1은 75.2%의 점수를 받은 반면, 기존 고급 음성 모드는 0.7%에 불과했습니다. GPT-Live는 또한 OpenAI의 내부 통신 테스트인 tau3 Voice Telecom 테스트에서도 선두를 달리고 있습니다.

원문 보기
원문 보기 (영어)
ChatGPT can now listen and talk at the same time, making AI conversations seem more human Matthias Bastian View the LinkedIn Profile of Matthias Bastian Jul 8, 2026 Key Points OpenAI has introduced GPT-Live, a new language model for live conversations in ChatGPT that can listen and speak simultaneously thanks to a full-duplex architecture. The model responds to interruptions, uses filler words like "mhmm" to keep the conversation flowing naturally, and is available in two versions: one for paying users and one for free users. For complex tasks that require web searches or logical reasoning, GPT-Live delegates the requests to the GPT-5.5 model in the background while maintaining the conversation, significantly improving answer quality. Ask about this article… Search OpenAI introduces GPT-Live, a new generation of voice models with full-duplex architecture. The model can listen and speak at the same time and hands off complex tasks to GPT-5.5 in the background. With GPT-Live, OpenAI is shipping a new class of AI voice models designed to make talking to ChatGPT feel more like a real conversation. Two versions are rolling out worldwide right away. GPT-Live-1 is for paying users on Go, Plus, and Pro plans, while GPT-Live-1 mini is available on free accounts. Both work on iOS, Android, and ChatGPT.com . OpenAI plans to add API access soon, and developers can sign up through a form . Instead of the old rigid back-and-forth where users speak and the AI responds, GPT-Live uses a full-duplex architecture. The model can listen and talk at the same time, much like a real conversation between people. Earlier this year, Nvidia released a similar open-source model called PersonaPlex . Ad OpenAI says GPT-Live makes decisions multiple times per second about whether to speak, keep listening, pause, interrupt, or call a tool. The model can use filler phrases like "mhmm" or "got it" to signal it's following along. Users can interrupt, take a moment to think, or ask the model to slow down. Ad DEC_D_Incontent-1 Compared to the previous "Advanced Voice Mode," people preferred GPT-Live-1 in 75.7 percent of cases and GPT-Live-1 mini in 69.2 percent, OpenAI says. Beyond the new conversational features, the update adds visual cards that ChatGPT Voice can show during a conversation for things like weather, stock prices, or sports scores. OpenAI also revamped the nine available voices for GPT-Live. At launch, GPT-Live doesn't support Voice with video or screen sharing in ChatGPT, though OpenAI says those features are coming soon. The older Standard and Advanced Voice Mode with these features will stick around for now. Ad Background delegation to GPT-5.5 closes the intelligence gap The biggest change is the split between live conversation management and actual reasoning. When a question needs a web search, reasoning, or agent-like capabilities, GPT-Live hands it off to a background model. Right now, that's GPT-5.5. While the background model works, GPT-Live keeps the conversation going. OpenAI says the architecture is built so GPT-Live stays connected to the latest frontier models at all times. Users can also pick a reasoning level, from "Instant" for quick answers to "Medium" and "High" for tasks where ChatGPT should spend more time thinking. Ad DEC_D_Incontent-2 This delegation fixes a major weakness of earlier live models, which answered questions based only on their own capabilities. That made them fall far behind frontier models and left them barely usable for serious work. Ad With GPT-Live, that gap appears to be closing, and benchmark scores for knowledge are drastically better. On the GPQA test for scientific reasoning, GPT-Live-1 hits 84.2 percent accuracy at the high reasoning level, compared to 45.3 percent in Advanced Voice Mode. The gap is even wider on BrowseComp, a benchmark for agent-based web search. GPT-Live-1 scores 75.2 percent. Advanced Voice Mode manages 0.7 percent. GPT-Live also leads on OpenAI's internal tau3 Voice Telecom test, which evaluates full-duplex voice agents on realistic telecom support tasks. Depending on reasoning level, GPT-Live-1 completes tasks at a much higher rate while also finishing faster. Advanced Voice Mode does worst, with the lowest success rate and the longest processing time. More human-sounding AI raises safety stakes People have a well-documented tendency to treat AI models like humans and let themselves be talked into all sorts of actions , sometimes with serious consequences . Research shows that heavy voice model users are even more prone to emotional dependency on AI . An AI system that sounds and responds even more like a person will likely amplify that risk. OpenAI says it has built safety measures that kick in even while the user is still speaking. The system can steer the model toward safer responses, show extra safety information, or end the conversation entirely in high-risk situations. Crisis hotlines appear when topics related to self-harm come up. For younger users, OpenAI trained the model to behave in an age-appropriate way. Parents can use parental controls to decide whether their child can access ChatGPT Voice and will get notified in high-risk situations. GPT-Live also isn't designed to mimic real voices and only uses predefined ones. Full details on the safety measures are available in the System Card . AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now Source: OpenAI
관련 소식