메뉴
BL
The Decoder • 2일 전

알리바바, Qwen Audio 3.1 출시… AI 오디오 가격 최대 95% 인하

IMP
7/10
핵심 요약

알리바바의 AI팀 Qwen이 음성 인식(ASR), 음성 합성(TTS), 실시간 상호작용용 모델 5종으로 구성된 Qwen-Audio-3.1을 공개했습니다. 다국어·방언 인식 개선, 감정·환경음 감지, 텍스트 프롬프트만으로 감정과 스타일 제어, 동시 듣기·말하기 지원 등이 핵심 기능입니다. 아울러 TTS 약 70%, 실시간 약 85%, ASR 최대 95% 가격을 인하해 AI 오디오 시장에서 공격적인 경쟁에 나섰습니다.

번역된 본문

알리바바, Qwen Audio 3.1 출시… AI 오디오 가격 최대 95% 인하 Matthias Bastian | 2026년 9월 23일

알리바바의 AI팀 Qwen이 음성 인식(ASR), 텍스트 음성 변환(TTS), 실시간 상호작용을 위한 5개 모델 라인업인 Qwen-Audio-3.1을 공개했다.

ASR 모델은 다국어 및 방언 인식 성능이 개선되었으며, 군더더기 말과 반복을 자동으로 정리해준다. ASR-Next는 타임스탬프가 포함된 다중 화자 식별 기능을 추가했고, 감정, 주변 소음, 기계 소음도 감지할 수 있다.

TTS는 다국어 합성과 자연스러운 크로스 언어 음성 전이를 지원한다. 사용자는 "이걸 날카롭고 위엄 있는 톤으로, 존경을 요구하듯 읽어줘" 같은 간단한 텍스트 프롬프트로 감정, 속도, 스타일을 조절할 수 있다. TTS-Next는 언어 모델과 디퓨전 방식을 결합해 음성, 효과음, 배경 오디오를 한 번에 생성한다.

실시간 모델은 동시에 듣고 말하는 것과 즉시 끼어들기를 지원한다. Qwen에 따르면 상대방의 기분이 가라앉아 있다고 감지하면 더 천천히, 더 공감하는 방식으로 응답한다.

알리바바는 가격도 대폭 인하한다. TTS는 약 70%, 실시간은 약 85%, ASR은 최대 95%까지 내린다. 자세한 내용은 블로그와 Qwen Cloud에서 확인할 수 있다.

광고

과장 없는 AI 뉴스 – 사람이 직접 엄선합니다 THE DECODER 구독: 광고 없는 읽기, 주간 AI 뉴스레터, 연 6회 발간되는 독점 프론티어 리포트 "AI Radar", 전체 아카이브 접근, 댓글 기능 이용.

지금 구독하기

출처: X 경유

원문 보기
원문 보기 (영어)
Alibaba launches Qwen Audio 3.1 with new models and slashes AI audio prices by up to 95 percent Matthias Bastian View the LinkedIn Profile of Matthias Bastian Sep 23, 2026 Alibaba's AI team Qwen has released Qwen-Audio-3.1, a lineup of five models for speech recognition (ASR), text-to-speech (TTS), and real-time interaction. The ASR model improves multilingual and dialect recognition and automatically cleans up filler words and repetitions. ASR-Next adds multi-speaker identification with timestamps and detects emotions, ambient sounds, and machine noise. TTS handles multilingual synthesis with natural cross-language voice transfer. Users control emotion, speed, and style through simple text prompts like "Read this with a sharp, commanding tone, demanding respect." TTS-Next pairs a language model with a diffusion approach to generate voice, sound effects, and background audio in a single pass. The real-time model supports simultaneous speaking and listening with instant interruption. When it detects a low mood, it responds more slowly and with more empathy, according to Qwen. Alibaba is also slashing prices. TTS drops about 70 percent, Realtime roughly 85 percent, and ASR up to 95 percent. More details on the blog and on Qwen Cloud . Ad Ad AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now Source: via X Ask about this article… Search