메뉴
BL
MarkTechPost • 24일 전

메타, 음성 인식·화자 구분 하나로 통합한 'Muse Voice Transcribe' 공개

IMP
6/10
핵심 요약

메타 슈퍼인텔리전스 연구소(Meta Superintelligence Labs)가 스트리밍 음성 인식(ASR), 화자 분리(Diarization), 발화 종료 감지(Endpointing)를 단일 자기회귀(autoregressive) 모델로 통합한 'Muse Voice Transcribe'를 공개했습니다. 기존 음성 시스템이 세 개의 모델을 연결해 지연과 장애 지점을 키웠다면, 이 모델은 하나로 통합해 실시간 처리에 유리합니다.

번역된 본문

대부분의 상용 음성 처리 스택은 세 개의 시스템을 이어 붙인 구조입니다. 하나의 모델이 텍스트를 전사하고, 두 번째 모델이 화자를 구분하며, 별도의 감지기가 사용자가 말을 멈췄는지 판단합니다. 각 단계가 연결될 때마다 지연 시간이 늘어나고 새로운 장애 지점이 생깁니다. 메타 슈퍼인텔리전스 연구소가 이번 주 발표한 'Muse Voice Transcribe'는 이 세 가지 작업을 단일 자기회귀(autoregressive) 모델 하나로 통합했습니다. 메타는 이를 ... [원문이 여기서 요약 형태로 종료됨]

이 글 '메타 슈퍼인텔리전스 연구소, Muse Voice Transcribe 공개: 스트리밍 ASR·화자 분리·발화 종료 감지를 위한 단일 실시간 모델'은 MarkTechPost에 처음 게재되었습니다.

원문 보기
원문 보기 (영어)
Most production voice stacks are three systems stitched together. One model transcribes, a second separates speakers, and a detector decides when the user stopped talking. Each hand-off adds latency and a new failure mode. Muse Voice Transcribe, announced by Meta Superintelligence Labs this week, collapses those three jobs into a single autoregressive model. Meta calls […] The post Meta Superintelligence Labs Releases Muse Voice Transcribe: One Real-Time Model for Streaming ASR, Diarization, and Endpointing appeared first on MarkTechPost.