Skip to content
Tech News
← Back to articles

Canto: A speech model built for the real world

read original get Jabra Evolve2 65 Headset → more articles
Why This Matters

Wispr Flow has unveiled Canto, a speech recognition model optimized for messy, real-world dictation conditions like background noise and varied microphones, rather than the clean studio audio most models are trained on. This matters because voice dictation is increasingly used for everyday tasks like messaging, coding, and note-taking, and accuracy in noisy real-life settings directly affects usability and adoption of voice-first interfaces.

Key Takeaways
Worth a Look

Jabra Evolve2 65 Headset — If you're dictating on the go through busy offices, commutes, or coffee shops like the scenarios Canto is built for, a headset with a noise-isolating boom mic can dramatically improve transcription accuracy. The Jabra Evolve2 65 is designed for exactly this kind of real-world use, with background noise suppression and a clear mic pickup that pairs well with speech-to-text tools like Wispr Flow. It's a practical companion for anyone relying on voice dictation throughout their workday.

See Jabra Evolve2 65 Headset on Amazon → Affiliate link — we may earn a commission on purchases, at no extra cost to you. Product picked by AI based on this article; it is not a tested recommendation.

Speech recognition models have become remarkably good at transcribing clean audio recorded under controlled conditions. But real dictation rarely happens under those conditions. Millions of people use Wispr Flow to message friends, write emails, code, and work through ideas at their desks, between meetings, during commutes, and in busy offices. They speak through laptop microphones, earbuds, and headsets, often with other voices, music, or traffic in the background.

Today, the Wispr Advanced Interfaces Lab is introducing Canto, our latest speech model for real-time dictation. On an evaluation of real-world dictations, Canto achieved the lowest word error rate among all the models we tested. We compared Canto with models from Google, OpenAI, AssemblyAI, and Deepgram.

Canto is the first model in a broader research and development program at Wispr Advanced Interfaces Lab. In this post, we share how it performs, how we trained it to handle challenging real-world conditions, and the research already shaping what comes next.

“Today, the Wispr Advanced Interfaces Lab is introducing Canto, our latest speech model for real-time dictation.” Ariya Rastrow CSO, Wispr Flow

Evaluating Canto in real-world conditions

To test Canto in real-world conditions, we created an evaluation set composed of 10 hours of English-language Wispr Flow dictations from more than 2,300 unique speakers, randomly sampled across applications and use cases. We were careful to enforce a strict separation between speakers represented in the train and test sets to avoid overfitting on speaker characteristics. Every sample came from a user who opted in to Wispr’s data-sharing setting, which allows their data to be used anonymously to evaluate and improve our models. More information about this setting is available in our data-controls documentation.

Canto achieved the lowest Word Error Rate (WER) of the models in our comparison. WER measures word substitutions, omissions, and insertions relative to a human-transcribed reference (lower is better).

Word Error Rates on 10 hours of Wispr Flow dictations, randomly sampled across applications and use cases.

Performance under the most challenging conditions

The model performed well on randomly sampled data, but we were also interested in studying its performance in the most challenging situations. We built a separate, 3-hour challenge evaluation set that featured those conditions most likely to cause dictation to fail.

... continue reading