메뉴
BL
The Decoder • 15일 전

OpenAI, 동시에 듣고 말하는 GPT-Live-1 API 공개

IMP
8/10
핵심 요약

OpenAI가 음성 모델 GPT-Live-1을 API로 개방했습니다. 이 모델은 '풀듀플렉스(full-duplex)' 방식으로 말을 하면서 동시에 들을 수 있어, 개발자가 자연스러운 실시간 음성 앱을 만들 수 있습니다. 분당 0.05달러로 저렴하지는 않지만, 벤치마크에서 이전 모델을 크게 앞서며 Yelp 등이 이미 전화 예약 서비스에 적용했습니다.

번역된 본문

OpenAI가 GPT-Live-1을 개발자용 API로 공개했습니다. 이 음성 모델은 '풀듀플렉스(full-duplex)'라 불리는 기능으로 말하고 동시에 들을 수 있으며, 이미 ChatGPT 내부에서 구동 중입니다. 개발자는 작업에 따라 다양한 백엔드 모델과 결합해 용도별로 추론 깊이, 속도, 비용을 조정할 수 있습니다. 분당 0.05달러로 저렴한 편은 아닙니다.

Yelp는 이 모델을 전화 기반 예약 서비스에 활용하고 있으며, CTO 알렉스 레비(Alex Levy)에 따르면 통화 처리 품질이 향상되었다고 합니다.

OpenAI의 벤치마크에서 GPT-Live-1은 이전 모델들을 크게 앞섭니다. 풀듀플렉스 상호작용 테스트에서 GPT-Realtime-2.1의 45.4%에 비해 80.1%를 기록했습니다. 순서 교체 지연시간(turn-taking latency)은 1.4초에서 0.8초로 줄었고, 도구 호출 정확도는 60%에서 87%로 향상됐습니다. 은행 음성 지원 벤치마크에서는 통과율이 이전 모델의 12.4%에서 32%로 상승했습니다.

GPT-Live-1은 다양한 억양, 방언, 언어를 아우르는 12가지 새 음성도 함께 제공하며, ASR 전사본과 응답 텍스트를 기본으로 제공합니다. 자세한 내용은 API 문서에서 확인할 수 있습니다.

출처: OpenAI

원문 보기
원문 보기 (영어)
OpenAI's GPT-Live-1 API lets developers build apps that talk and listen at the same time Matthias Bastian View the LinkedIn Profile of Matthias Bastian Sep 10, 2026 OpenAI is making GPT-Live-1 available to developers as an API. The speech model can listen and talk at the same time, a feature known as "full-duplex," and is already running inside ChatGPT. Developers can pair it with different backend models depending on the task, matching reasoning depth, speed, and cost to each use case. At $0.05 per minute, it's not cheap. Yelp is using the model for phone-based reservations and reports better call handling, according to CTO Alex Levy . On OpenAI's benchmarks, GPT-Live-1 pulls well ahead of its predecessors. In full-duplex interactivity tests, it scores 80.1 percent compared to 45.4 percent for GPT-Realtime-2.1 . Turn-taking latency drops to 0.8 seconds from 1.4 seconds. Tool-calling accuracy jumps to 87 percent from 60 percent. In a banking voice support benchmark, GPT-Live-1 hits a 32 percent pass rate, up from 12.4 percent for the previous model. GPT-Live-1 also ships with twelve new voices spanning different accents, dialects, and languages. It provides ASR transcripts and response text out of the box. Full details will be available in the API documentation . Ad Ad AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now Source: OpenAI Ask about this article… Search
관련 소식