메뉴
HN
Hacker News • 8일 전

칸토(Canto): 실제 환경에 최적화된 음성 인식 모델

IMP
6/10
핵심 요약

Wispr AI Lab이 실시간 받아쓰기용 음성 인식 모델 '칸토(Canto)'를 공개했습니다. 잡음, 저볼륨, 짧은 발화 등 실제 사용 환경의 받아쓰기 평가에서 Google, OpenAI, AssemblyAI, Deepgram 등 경쟁 모델 대비 최저 단어 오류율(WER)을 기록했습니다. 공개 벤치마크에서도 LibriSpeech 최저 오류율에并列했으며, 실시간 저지연 응용에 적합한 모델이라는 점이 의미 있습니다.

번역된 본문

2026년 9월 17일 • 연구 • 10분

칸토(Canto): 실제 환경에 맞춰 만들어진 음성 모델

Wispr AI Lab

음성 인식 모델은 통제된 환경에서 녹음된 깨끗한 오디오를 전사하는 데는 놀라울 정도로 뛰어난 성능을 보입니다. 하지만 실제 받아쓰기는 그런 조건에서 이루어지는 경우가 거의 없습니다. 수백만 명의 사용자가 Wispr Flow를 이용해 책상에서, 회의 사이에, 통근 중이나 붐비는 사무실에서 친구에게 메시지를 보내고, 이메일을 쓰고, 코드를 작성하고, 아이디어를 정리합니다. 이들은 노트북 마이크, 이어버드, 헤드셋으로 말하며, 배경에는 다른 사람의 목소리, 음악, 교통 소음이 섞여 있는 경우가 많습니다.

오늘 Wispr Advanced Interfaces Lab은 실시간 받아쓰기를 위한 최신 음성 모델인 칸토(Canto)를 소개합니다. 실제 받아쓰기 데이터에 대한 평가에서 칸토는 우리가 테스트한 모든 모델 중 가장 낮은 단어 오류율을 기록했습니다. 우리는 칸토를 Google, OpenAI, AssemblyAI, Deepgram의 모델과 비교했습니다. 칸토는 Wispr Advanced Interfaces Lab의 더 폭넓은 연구개발 프로그램의 첫 번째 모델입니다. 이 글에서는 칸토의 성능, 도전적인 실제 환경을 처리하도록 학습시킨 방법, 그리고 다음 단계를 이미 만들어가고 있는 연구를 공유합니다.

"오늘 Wispr Advanced Interfaces Lab은 실시간 받아쓰기를 위한 최신 음성 모델인 칸토(Canto)를 소개합니다."

아리야 라스트로우(Ariya Rastrow), Wispr Flow CSO

실제 환경에서 칸토 평가하기

실제 환경에서 칸토를 테스트하기 위해, 우리는 2,300명 이상의 고유 화자로부터 수집된 10시간 분량의 영어 Wispr Flow 받아쓰기 데이터로 평가 세트를 구성했으며, 다양한 애플리케이션과 사용 사례에서 무작위로 샘플링했습니다. 학습 세트와 테스트 세트에 포함된 화자 간에 엄격한 분리를 적용하여 화자 특성에 대한 과적합을 방지했습니다. 모든 샘플은 Wispr의 데이터 공유 설정에 동의한 사용자에게서 나온 것으로, 해당 데이터는 익명으로 모델 평가 및 개선에 사용됩니다. 이 설정에 대한 자세한 내용은 데이터 관리 문서에서 확인할 수 있습니다.

칸토는 비교 대상 모델들 중 가장 낮은 단어 오류율(WER)을 달성했습니다. WER은 사람이 전사한 정답 대비 단어 치환, 생략, 삽입을 측정하는 지표로, 낮을수록 좋습니다.

가장 도전적인 조건에서의 성능

무작위 샘플링 데이터에서 좋은 성능을 보였지만, 우리는 가장 어려운 상황에서의 성능도 연구하고자 했습니다. 받아쓰기가 실패하기 가장 쉬운 조건들을 담은 별도의 3시간 분량 챌린지 평가 세트를 구축했습니다. 이 세트에는 근처의 말소리, 음악, 교통 소음, 바람, 낮은 녹음 볼륨의 영향을 받은 오디오와 속삭이거나 먼 거리에서의 발화가 포함되어 있습니다. 또한 모델이 모호성을 해결할 수 있도록 해주는 주변 맥락과 언어 정보가 거의 없는 짧은 받아쓰기도 포함됩니다.

전체 챌린지 세트에서 칸토는 Gemini 3.1 Pro에 이어 2위를 기록했습니다. Gemini 3.1 Pro는 실시간 저지연 애플리케이션에는 적합하지 않은 훨씬 큰 프론티어급 멀티모달 모델입니다. 우리가 평가한 실시간 전사 모델들 중에서는 칸토가 가장 낮은 WER을 달성했습니다.

우리는 이 데이터셋을 더 깊이 분석해 모델별 차이를 살펴보았습니다. Gemini 3.1 Pro는 잡음이 있는 오디오에서 가장 낮은 WER을 기록했습니다. 칸토는 저볼륨 발화와 짧은 받아쓰기에서 최저 WER을 공동으로 기록했습니다. 짧은 받아쓰기는 비교 전체에서 가장 높은 오류율을 보였습니다. 발화에 단어가 몇 개 없을 때는 실수 하나가 WER에 미치는 영향이 크고, 모델이 모호성을 해결할 수 있는 언어적 맥락도 적기 때문입니다.

공개 벤치마크에서의 성능 비교

우리는 칸토를 세 가지 공개 영어 데이터셋(LibriSpeech, FLEURS, Common Voice)에서도 평가했습니다. 칸토는 LibriSpeech에서 최저 WER을 공동 기록했습니다. FLEURS와 Common Voice에서도 경쟁력 있는 성능을 유지했지만 두 평가에서 1위를 하지는 못했습니다. 이러한 공개 데이터셋은 유용하고 재현 가능한 비교를 제공하지만, 대부분 낭독 음성을 담고 있습니다. LibriSpeech는 오디오북에서 추출되었고, FLEURS와 Common Voice는 주로 사람들이 준비된 문장을 읽는 것으로 구성되어 있습니다. 이는 일상적인 '실제 환경(in-the-wild)'의 받아쓰기와는 다른 분포를 반영합니다.

원문 보기
원문 보기 (영어)
17.09.2026 • Research • 10 min Canto: a speech model built for the real world Wispr AI Lab Speech recognition models have become remarkably good at transcribing clean audio recorded under controlled conditions. But real dictation rarely happens under those conditions. Millions of people use Wispr Flow to message friends, write emails, code, and work through ideas at their desks, between meetings, during commutes, and in busy offices. They speak through laptop microphones, earbuds, and headsets, often with other voices, music, or traffic in the background. Today, the Wispr Advanced Interfaces Lab is introducing Canto, our latest speech model for real-time dictation. On an evaluation of real-world dictations, Canto achieved the lowest word error rate among all the models we tested. We compared Canto with models from Google, OpenAI, AssemblyAI, and Deepgram. Canto is the first model in a broader research and development program at Wispr Advanced Interfaces Lab. In this post, we share how it performs, how we trained it to handle challenging real-world conditions, and the research already shaping what comes next. “ Today, the Wispr Advanced Interfaces Lab is introducing Canto, our latest speech model for real-time dictation. ” Ariya Rastrow CSO, Wispr Flow Evaluating Canto in real-world conditions To test Canto in real-world conditions, we created an evaluation set composed of 10 hours of English-language Wispr Flow dictations from more than 2,300 unique speakers, randomly sampled across applications and use cases. We were careful to enforce a strict separation between speakers represented in the train and test sets to avoid overfitting on speaker characteristics. Every sample came from a user who opted in to Wispr’s data-sharing setting, which allows their data to be used anonymously to evaluate and improve our models. More information about this setting is available in our data-controls documentation . Canto achieved the lowest Word Error Rate (WER) of the models in our comparison. WER measures word substitutions, omissions, and insertions relative to a human-transcribed reference (lower is better). Performance under the most challenging conditions The model performed well on randomly sampled data, but we were also interested in studying its performance in the most challenging situations. We built a separate, 3-hour challenge evaluation set that featured those conditions most likely to cause dictation to fail. The set includes audio affected by nearby speech, music, traffic, wind, low recording volume, and whispered or far-field speech. It also includes short dictations that give a model very little surrounding context and language to help resolve ambiguity. On the full challenge set, Canto ranked second behind Gemini 3.1 Pro, a much larger frontier-size multi-modal model that is not suitable for real-time low-latency applications. Among the real-time transcription models we evaluated, Canto achieved the lowest WER. We further examined this dataset to understand how the models behave differently. Gemini 3.1 Pro achieved the lowest WER on noisy audio. Canto tied for the lowest WER on low-volume speech and short dictations. Short dictations produced the highest error rates across the comparison. A single mistake has a larger effect on WER when an utterance contains only a few words, and the models also have less linguistic context available to resolve ambiguity. Comparing performance on public benchmarks We also evaluated Canto on three public English datasets: LibriSpeech, FLEURS, and Common Voice. Canto tied for the lowest WER on LibriSpeech. It remained competitive on FLEURS and Common Voice, although it did not lead either evaluation. These public datasets provide useful, reproducible comparisons, but they largely contain read speech. LibriSpeech is drawn from audiobooks, while FLEURS and Common Voice consist primarily of people reading prepared sentences. They capture a different distribution from everyday “in-the-wild” dictation, where people speak spontaneously, pause, revise their thoughts, and often provide very little linguistic context. The contrast helps explain why we evaluate Canto on both public datasets and real Wispr usage. Post-training Canto with Supervised Fine-Tuning and Reinforcement Learning Canto starts from a model pretrained on millions of hours of speech and text. We then train Canto in two stages. First, we show it audio paired with reference transcripts. The model learns to predict the words in each transcript, one step at a time. This stage, called supervised fine-tuning, teaches it how to perform the transcription task. Next, we train it to compare the quality of complete transcripts. For the same audio, the model generates several possible transcriptions. We score each one against a reference, then use those scores to make better transcriptions more likely in future training. This is known as reinforcement learning, or RL. The distinction is in how the model receives feedback. Supervised fine-tuning provides the expected words at each step. Reinforcement learning evaluates the completed transcription (generated by the model). That lets us train around the outcomes we care about, such as reducing recognition errors under difficult recording conditions. Our approach uses Group Relative Policy Optimization, or GRPO. The central idea is simple: compare the candidate transcripts within each group and learn from their relative scores. For each audio example, we sample several candidate transcripts from the model being trained. We call these candidates rollouts. Each receives a reward: a numerical score measuring how well it satisfies the training objective. GRPO compares each candidate’s reward with the rewards of the other candidates for the same audio. The resulting relative score, called an advantage, indicates whether that candidate performed better or worse than its group. These advantages guide the training update. Sequence-level training has a long history in speech recognition. Researchers have previously optimized ASR systems using minimum word error rate training and policy-gradient methods . GRPO was introduced more recently in DeepSeekMath , and direct applications to autoregressive speech recognition have only begun to appear in recent work on GRPO for ASR and speech-domain adaptation . “We built infrastructure that can generate and score speech rollouts at scale, allowing us to construct training environments around specific behaviors and failure modes.” For Wispr, the value of RL is not just a lower aggregate error rate. We built infrastructure that can generate and score speech rollouts at scale, allowing us to construct training environments around specific behaviors and failure modes. That system now supports experiments in contextual speech recognition, personalization, diarization, and difficult audio conditions. The research described below extends beyond the training used for the released Canto model and is already informing the next generation of our models. Learning from corrections The usage of names and vocabulary changes faster than speech models can be retrained. A word such as “Claude” might once have been an uncommon alternative to “cloud.” As its usage changes, a practical dictation model needs a way to learn the distinction. For users who have opted into data sharing, corrections can provide a valuable signal for improving recognition. But not every edit points to a transcription error. People also change formatting, rewrite sentences, or simply change their minds. A hypothetical example below illustrates this problem. The model transcribed “Claude” as “cloud,” which the person corrected. They also changed their mind and rewrote the end of the sentence. Only the first change represents a recognition error. We use signals from the audio and the edit itself to identify corrections that are most likely to represent recognition errors. These include forced-alignment confidence, which measures how wel