메뉴
HN
Hacker News 20일 전

GPT-Live 발표: 자연스러운 양방향 음성 대화의 새 시대

IMP
9/10
핵심 요약

AI가 사람처럼 자연스럽게 대화하며 동시에 듣고 말할 수 있는 새로운 양방향(Full-duplex) 음성 모델, GPT-Live가 공개되어 ChatGPT 음성 기능을 강화합니다. 이 모델은 복잡한 질문에는 백그라운드에서 최신 프론티어 모델(GPT-5.5)을 활용해 추론 및 작업을 수행하면서도 끊김 없는 대화를 유지합니다. 이는 기존의 STT-LLM-TTS를 순차적으로 거치던 방식의 한계를 넘어, 실제 사람과 대화하는 듯한 매끄러운 인간-AI 상호작용을 구현했다는 점에서 매우 중요합니다.

번역된 본문

2026년 7월 8일 제품 출시: GPT-Live 소개. 자연스러운 인간-AI 상호작용을 위한 차세대 음성 모델로, 현재 ChatGPT 음성(Voice)을 구동하고 있습니다.

오늘 우리는 AI와 대화하는 것이 실제 사람과 대화하는 것과 훨씬 더 비슷하게 느껴지게 만드는 차세대 음성 모델, GPT-Live를 출시합니다. GPT-Live는 전이중(Full-duplex) 아키텍처를 기반으로 구축되었으며, 이는 동시에 듣고 말할 수 있음을 의미합니다. 대화 중에 GPT-Live는 "음" 또는 "네"와 같은 말로 주의를 기울이고 있음을 보여주거나, 빠르고 자연스럽게 대화를 주고받거나, 사용자가 생각할 시간이 필요할 때는 조용히 기다릴 수도 있습니다. 그 결과, 매우 편안하게 대화할 수 있는 새로운 음성 경험을 제공합니다.

또한 GPT-Live는 우리가 개발한 가장 스마트한 음성 모델이기도 합니다. 웹 검색, 더 깊은 추론 또는 더 복잡한 작업이 필요한 질문의 경우, 무대 뒤에서 최신 프론티어 모델(Frontier model)에 작업을 위임하고 결과가 준비되면 다시 대화로 가져옵니다. 백그라운드에서 작업을 수행하는 동안에도 GPT-Live는 사용자와 계속 대화를 나누며 흐름을 유지할 수 있습니다. 출시와 함께 GPT-Live는 백그라운드에서 GPT-5.5를 사용하게 됩니다. 새로운 프론티어 모델이 출시됨에 따라 GPT-Live에 사용되는 모델도 지속적으로 업데이트될 예정입니다. 이러한 발전을 통해 더 지능적이고 자연스럽게 사용할 수 있는 새로운 ChatGPT 음성 경험을 제공합니다. 시간이 지남에 따라 이 연구가 점점 더 복잡하고 오래 실행되며 에이전트 중심적인 작업에 음성을 사용하는 능력도 함께 열어줄 것이라 믿습니다.

오늘부터 전 세계 ChatGPT 사용자에게 두 가지 버전의 GPT-Live, 즉 GPT-Live-1 및 GPT-Live-1 mini의 배포를 시작합니다. 또한 조만간 API를 통해서도 이 기능을 제공할 계획이며, 개발자 및 기업 고객은 제공된 양식을 작성하여 알림을 신청할 수 있습니다.

인간-AI 상호작용의 새로운 시대로의 진입 우리의 비전은 진정으로 자연스러운 인간-AI 상호작용을 가능하게 하는 것입니다. 즉, 추론 및 복잡한 작업 실행이 배경에서 원활하게 이루어지는 동안, AI와 협업하는 것이 다른 사람과 일하는 것만큼이나 유연하고 반응이 빠르게 느껴지는 세상을 만드는 것입니다.

이전 접근 방식 이전 세대의 음성 AI 시스템은 우리를 그 비전에 한 걸음 더 가까이 다가가게 했지만, 중요한 타협점들을 동반했습니다.

계단식 음성 시스템 (Cascaded voice systems) 계단식 음성 시스템은 각 대화 차례를 처리하기 위해 여러 모델이 연달아 작동하는 방식에 의존합니다. 초기의 ChatGPT 음성은 세 가지 모델을 함께 연결했습니다. 사용자의 음성을 텍스트로 변환하는 STT(Speech-to-Text) 모델, 응답을 생성하는 대형 언어 모델(LLM), 그리고 이를 다시 음성으로 변환하는 TTS(Text-to-Speech) 모델입니다. 이러한 접근 방식 덕분에 우리는 처음으로 최첨단 AI 모델과 대화할 수 있었습니다. 하지만 이러한 복잡성은 대가를 치르게 했습니다. 모델 간에 정보가 손실될 수 있었고, 응답은 느리고 부자연스러웠습니다.

원문 보기
원문 보기 (영어)
July 8, 2026 Product Release Introducing GPT‑Live A new generation of voice models for natural human-AI interaction, now powering ChatGPT Voice. Loading… Share We’re launching GPT‑Live, a new generation of voice models that make talking with AI feel much more like having a real conversation. GPT‑Live is built on a full-duplex architecture, meaning it can listen and speak at the same time. During conversations, GPT‑Live can show it’s paying attention with phrases like “mhmm” or “yeah”, engage in quick back-and-forth, or just stay quiet when you need a moment to think. The result is a voice experience that is refreshingly easy to talk to. GPT‑Live is also our smartest voice model yet. For questions that require web search, deeper reasoning, or more complex work, it delegates to our latest frontier model behind the scenes and brings the result back into the conversation when it’s ready. While it works, GPT‑Live can keep talking with you and maintain the flow of conversation. At launch, GPT‑Live will use GPT‑5.5 in the background. As we release new frontier models, we’ll continuously update the model used by GPT‑Live. These advances power a new ChatGPT Voice experience that is more intelligent and natural to use. Over time, we believe this research will also unlock the ability to use voice for increasingly complex, longer-running, and more agentic work. We’re beginning to roll out two versions of GPT‑Live – GPT‑Live‑1 and GPT‑Live‑1 mini – to ChatGPT users globally today. We also plan to bring them to the API soon, and developers and enterprises can sign up to be notified using this form ⁠ . Entering a new era of human-AI interaction Our vision is to enable truly natural human–AI interaction: a world where collaborating with AI feels as fluid and responsive as working with another person, while reasoning and complex task execution happen seamlessly in the background. Previous approaches Older generations of voice AI systems brought us closer to that vision, but with important tradeoffs. Cascaded voice systems Cascaded voice systems rely on a series of models acting one after another to process each turn. The original ChatGPT Voice chained three models together: a speech-to-text model to transcribe your speech, a large language model to produce a response, and a text-to-speech model to convert it back into speech. This approach enabled us to talk to frontier AI models for the first time, but the complexity came at a cost: information could be lost across models, and responses were slow and stilted. STT GPT-5.5 TTS STT GPT-5.5 TTS Transcript Example conversation with Standard Voice Mode, using GPT-5.5 Instant
관련 소식