Speech recognition models have become remarkably good at transcribing clean audio recorded under controlled conditions. But real dictation rarely happens under those conditions. Millions of people use Wispr Flow to message friends, write emails, code, and work through ideas at their desks, between meetings, during commutes, and in busy offices. They speak through laptop microphones, earbuds, and headsets, often with other voices, music, or traffic in the background.

Today, the Wispr Advanced Interfaces Lab is introducing Canto, our latest speech model for real-time dictation. On an evaluation of real-world dictations, Canto achieved the lowest word error rate among all the models we tested. We compared Canto with models from Google, OpenAI, AssemblyAI, and Deepgram.

Canto is the first model in a broader research and development program at Wispr Advanced Interfaces Lab. In this post, we share how it performs, how we trained it to handle challenging real-world conditions, and the research already shaping what comes next.

Evaluating Canto in real-world conditions required creating an evaluation set composed of 10 hours of English-language Wispr Flow dictations from more than 2,300 unique speakers, randomly sampled across applications and use cases. We were careful to enforce a strict separation between speakers represented in the train and test sets to avoid overfitting on speaker characteristics. Every sample came from a user who opted in to Wispr’s data-sharing setting, which allows their data to be used anonymously to evaluate and improve our models.

Canto achieved the lowest Word Error Rate (WER) of the models in our comparison. WER measures word substitutions, omissions, and insertions relative to a human-transcribed reference (lower is better). The model performed well on randomly sampled data, but we were also interested in studying its performance in the most challenging situations. We built a separate, 3-hour challenge evaluation set that featured those conditions most likely to cause dictation to fail.

The set includes audio affected by nearby speech, music, traffic, wind, low recording volume, and whispered or far-field speech. It also includes short dictations that give a model very little surrounding context and language to help resolve ambiguity. On the full challenge set, Canto ranked second behind Gemini 3.1 Pro, a much larger frontier-size multi-modal model that is not suitable for real-time low-latency applications. Among the real-time transcription models we evaluated, Canto achieved the lowest WER.

We further examined this dataset to understand how the models behave differently. Gemini 3.1 Pro achieved the lowest WER on noisy audio. Canto tied for the lowest WER on low-volume speech and short dictations. Short dictations produced the highest error rates across the comparison. A single mistake has a larger effect on WER when an utterance contains only a few words, and the models also have less linguistic context available to resolve ambiguity.

Canto starts from a model pretrained on millions of hours of speech and text. We then train Canto in two stages. First, we show it audio paired with reference transcripts. The model learns to predict the words in each transcript, one step at a time. This stage, called supervised fine-tuning, teaches it how to perform the transcription task. Next, we train it to compare the quality of complete transcripts. For the same audio, the model generates several possible transcriptions. We score each one against a reference, then use those scores to make better transcriptions more likely in future training. This is known as reinforcement learning, or RL.

The distinction is in how the model receives feedback. Supervised fine-tuning provides the expected words at each step. Reinforcement learning evaluates the completed transcription (generated by the model). That lets us train around the outcomes we care about, such as reducing recognition errors under difficult recording conditions. Our approach uses Group Relative Policy Optimization, or GRPO.

Canto is the first model in a larger family. We are already training its successor at more than ten times Canto’s scale, with the goal of improving recognition under difficult audio conditions, improving multi-speaker recognition, making better use of contextual vocabulary, and reaching the same standard across more languages. The post-training approach described here will also let us introduce more specialized rewards and harder training environments as the models grow.

Questo articolo è stato scritto con l'assistenza dell'IA.
News Factory APP - notizie agentiche per potenziare il tuo SEO e AEO.