BL
MarkTechPost • 2일 전
NVIDIA, 화자 8명 실시간 추적하는 오픈 웨이트 화자 구분 모델 공개
IMP 6/10
핵심 요약
NVIDIA가 1억 파라미터 규모의 오픈 웨이트 화자 구분(speaker diarization) 모델 'Nemotron 3 Diarization'을 허깅페이스에 공개했습니다. 이 모델은 대화에서 '누가 언제 발언했는지'를 추적하며, 최대 8명의 화자를 음성이 겹치는 상황에서도 실시간으로 구분할 수 있습니다. 하나의 체크포인트로 오프라인 녹음과 실시간 스트리밍 모두 처리 가능하며, 웨이트가 공개되어 실무 배포가 가능하다는 점이 중요합니다.
번역된 본문
NVIDIA가 오픈 웨이트 화자 구분(diarization) 모델인 Nemotron 3 Diarization을 허깅페이스(Hugging Face)에 공개했습니다. 이 모델은 모든 대화에 대해 단 하나의 질문에 답합니다. 즉, '누가 언제 발언했는가'입니다. 1억(100M) 파라미터 규모의 이 모델은 최대 8명의 화자를 추적할 수 있으며, 음성이 겹치는 구간에서도 가능합니다. 하나의 체크포인트로 오프라인 녹음과 실시간 스트리밍 모두를 처리합니다. 배포 가능할까요? 네. 웨이트는 다음 라이선스 하에 공개되었습니다[…]
원문 보기 (영어)
NVIDIA has released Nemotron 3 Diarization, an open-weight speaker diarization model on Hugging Face. It answers one question about any conversation: who spoke when. The 100M-parameter model tracks up to 8 speakers, including when voices overlap. One checkpoint handles both offline recordings and real-time streaming. Is it deployable? Yes. The weights are released under the […]
The post NVIDIA Releases Nemotron 3 Diarization: A 100M-Parameter Open-Weight Model That Tracks 8 Speakers in Real Time appeared first on MarkTechPost.