메뉴
BL
The Decoder 8일 전

알리바바 '큐원 오디오 3.0', TTS 리더보드 1위 달성

IMP
7/10
핵심 요약

알리바바의 새로운 텍스트 음성 변환(TTS) 모델인 Qwen-Audio-3.0-TTS-Plus가 인공지능 벤치마크 플랫폼의 오디오 리더보드에서 경쟁 모델들을 제치고 1위를 차지했습니다. 이 모델은 16개 국어를 지원하며, 태그를 통한 감정 표현과 뛰어난 목소리 복제 능력을 자랑하지만 생성 속도 측면에서는 경쟁사에 다소 뒤처진다는 특징이 있습니다.

번역된 본문

알리바바의 새로운 텍스트 음성 변환(TTS) 모델, 'Qwen-Audio-3.0-TTS-Plus'가 Artificial Analysis의 Speech Arena 리더보드에서 서비스 제공자 보이스 부문 1위를 차지했습니다. 이 모델은 1,236점의 Elo 점수를 기록하며 근소한 차이로 Simba 3.2(1,234점)를 제치는 성과를 보였습니다. Gemini 3.1 Flash TTS(1,214점)와 Sonic 3.5(1,207점)가 그 뒤를 잇고 있습니다.

해당 모델은 두 가지 버전으로 제공됩니다. 'Flash' 버전은 약 300밀리초의 짧은 지연 시간(latency)으로 실시간 상호작용에 최적화되어 있으며, 'Plus' 버전은 최고 수준의 고품질 음성 출력을 목표로 합니다. 타갈로그어, 말레이어, 태국어, 베트남어 및 여러 중국어 방언을 포함해 총 16개 국어를 지원합니다. 또한 사용자는 자연어로 말투를 조절하거나, "[angry]" 또는 "[giggles]"와 같은 태그를 사용해 비언어적인 단서를 추가할 수 있습니다.

알리바바는 또한 이 모델이 이전 버전들에 비해 잡음이나 에코가 심한 참조 음원을 사용해 목소리를 복제(Clone)할 때 훨씬 뛰어난 처리 능력을 보여준다고 밝혔습니다. 반면 음성 생성 속도는 아쉬운 부분으로 꼽힙니다. 초당 16자를 처리하여 초당 120자인 Sonic 3.5나 30.2자인 Simba 3.2에 비해 큰 격차로 뒤처집니다. 가격은 알리바바 클라우드 모델 스튜디오를 통해 100만 자당 27.60달러로 책정되었습니다.

원문 보기
원문 보기 (영어)
Alibaba's Qwen Audio 3.0 TTS Plus tops the competition in the text-to-speech rankings Matthias Bastian View the LinkedIn Profile of Matthias Bastian Jul 21, 2026 Alibaba's new text-to-speech model, Qwen-Audio-3.0-TTS-Plus, leads Artificial Analysis' Speech Arena leaderboard for provider voices. With an Elo score of 1,236, it sits just ahead of Simba 3.2 (1,234). Gemini 3.1 Flash TTS (1,214) and Sonic 3.5 (1,207) follow behind. The model comes in two versions. Flash is built for real-time interaction with about 300 milliseconds of latency, while Plus targets high-quality speech output. It supports 16 languages, including less commonly covered ones like Tagalog, Malay, Thai, and Vietnamese, along with several Chinese dialects. Users can steer the speaking style with natural language or add nonverbal cues using tags like "[angry]" or "[giggles]." Alibaba also says the model handles noisy or echo-heavy reference recordings better than previous versions when cloning voices. Speed is a weak spot: At 16 characters per second, it trails Sonic 3.5 (120) and Simba 3.2 (30.2) by a wide margin. Pricing lands at $27.60 per million characters through Alibaba Cloud Model Studio . A collection of audio samples is available here . Ad DEC_D_Incontent-1 Ad AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now Source: Artificial Analysis Ask about this article… Search