Joe Maring / Android Authority
TL;DR Google has introduced Gemini Audio, a trio of new models built to improve dialogue and speech recognition.
The models include Gemini 3.5 Live, 3.5 Live Experimental, and 3.5 Transcribe.
The new family of AI models is rolling out for everyone in Search Live, Gemini Live, Docs, Keep, Gmail, the Gemini app, and Gboard.
It’s been a few months since Google debuted the first of its Gemini 3.5 models. With the rapid pace of AI development, the company has since released Gemini 3.6 and Gemini 3.7. However, the tech giant hasn’t completely moved on from the 3.5 family just yet, as it is introducing new members to the group today.
Google has announced a trio of new speech-related models: Gemini 3.5 Transcribe, Gemini 3.5 Live, and Gemini 3.5 Live Experimental. Together, these models make up what the company calls Gemini Audio. According to Google, these models were built to power real-time dialogue and speech recognition for more natural and responsive conversations.
Gemini 3.5 Transcribe: Advanced transcription Gemini 3.5 Transcribe is described as the Mountain View-based firm’s most precise speech-to-text model yet. It will replace the previous transcription model, Chirp 3. Google claims that Gemini 3.5 Transcribe offers better precision, deep context awareness, and automatic language detection across more than 85 languages.
Transcribe delivers a handful of key benefits, such as: Smart transcription: Removes filler words (like “ums” and “ahs”), auto-formats your text, and edits naturally with just your voice.
Removes filler words (like “ums” and “ahs”), auto-formats your text, and edits naturally with just your voice. Function calling: The model can delegate complex tasks (such as image generation) to other Gemini models via function calls. While function calls are already available in the Gemini macOS app for developers, support is expected to come to the API soon.
The model can delegate complex tasks (such as image generation) to other Gemini models via function calls. While function calls are already available in the Gemini macOS app for developers, support is expected to come to the API soon. More precise transcriptions: Google claims its model delivers a 4% WER (word-error-rate) on streaming audio and 2.6% WER on pre-recorded files across diverse real-world conditions, including background noise and conversational AI interactions.
... continue reading