메뉴
BL
MarkTechPost • 39일 전

카르테시아, Sonic-3.6 공개 — 스피치 아레나 양대 리더보드 1위

IMP
6/10
핵심 요약

Cartesia(카르테시아)가 트랜스포머가 아닌 상태 공간 모델(state space model) 기반의 스트리밍 TTS(텍스트 음성 변환) 모델 Sonic-3.6을 공개했습니다. 이 모델은 Artificial Analysis의 두 음성 리더보드에서 모두 1위를 차지했으며(Provider Voice 1,283 Elo, Controlled Voice 1,123 Elo), 첫 오디오 출력까지 90ms 미만의 지연 시간을 자랑합니다. 현재 Cartesia 자체 API에서 베타로 이용할 수 있습니다.

번역된 본문

Cartesia(카르테시아)가 Sonic-3.6을 출시했습니다. 이는 트랜스포머(transformer)가 아닌 상태 공간 모델(state space model)을 기반으로 구축된 스트리밍 텍스트 음성 변환(TTS) 모델입니다. 현재 Artificial Analysis의 음성 리더보드 양쪽에서 모두 1위를 기록하고 있습니다. Provider Voice에서는 1,283 Elo, Controlled Voice에서는 1,123 Elo를 달성했는데, 후자는 모든 모델을 동일한 8개의 기준 음성(reference voice)에 복제(cloning)하여 합성 엔진 자체의 성능만을 비교하는 리더보드입니다. Cartesia는 첫 오디오 출력까지의 시간(time-to-first-audio)이 90ms 미만이라고 밝혔습니다. 이 모델은 현재 Cartesia 자체 API에서 베타 버전으로 이용할 수 있습니다.

원문 보기
원문 보기 (영어)
Cartesia has released Sonic-3.6, a streaming text-to-speech model built on state space models rather than transformers. It now ranks #1 on both Artificial Analysis speech leaderboards — 1,283 Elo on Provider Voice and 1,123 on Controlled Voice, the board that clones every model onto the same eight reference voices to isolate the synthesis engine. Cartesia states sub-90ms time-to-first-audio. The model is available in beta on Cartesia's own API The post Cartesia Ships Sonic-3.6: A Streaming TTS Model That Now Leads Both Artificial Analysis Speech Arenas appeared first on MarkTechPost.