Tech News
← Home  ·  All topics

Speech Recognition

4 GoKawiil briefs on this topic

Wispr Flow launches Canto speech model for real-world dictation

Wispr Advanced Interfaces Lab has released Canto, a new speech recognition model built specifically for real-time dictation in noisy, everyday conditions rather than clean studio audio. In tests using 10 hours of real Wispr Flow dictations from over 2,300 speakers, Canto posted the lowest word error rate compared with models from Google, OpenAI, AssemblyAI, and Deepgram.

Microsoft cuts MAI-Transcribe-2 pricing to 10 cents per hour, undercutting rivals

Microsoft AI launched MAI-Transcribe-2, a speech-to-text model it says beats OpenAI, Google and ElevenLabs on speed, accuracy and cost, pricing it at 10 cents per audio hour—a 72% cut from the $0.36 rate it charged for its first model just five months ago. The new version supports 60 languages, up from 43, and adds built-in speaker diarization and word-level timestamps, targeting noisy real-world business audio rather than clean recordings.

Google launches Gemini 3.5 Transcribe with cleanup and speaker-detection features

Google unveiled Gemini 3.5 Transcribe, a new speech-to-text AI model that detects over 85 languages, strips filler words, and reformats casual speech into clean, structured text. It can also learn custom vocabulary, capture alphanumeric strings like order numbers, and assign dialogue to up to three speakers with timestamps from recorded audio.

Google launches Gemini Audio models to power live dialogue and transcription

Google has released Gemini Audio, a set of three new speech-focused AI models: Gemini 3.5 Live, 3.5 Live Experimental, and 3.5 Transcribe. These models are meant to handle real-time conversation and speech-to-text tasks and are rolling out across Search Live, Gemini Live, Docs, Keep, Gmail, the Gemini app, and Gboard. Notably, Gemini 3.5 Transcribe replaces Google's older Chirp 3 model and adds features like filler-word removal, auto-formatting, and voice editing.