Google has launched beta 'Live' voice features across Gmail, Docs and Keep, letting subscribers talk to these apps instead of typing. Users can ask Gmail to find or summarize emails, have Keep organize notes from spoken input, and tell Docs to draft or pull in content from Gmail, Chat and Drive. The features, first shown at Google's I/O conference in May, are currently limited to AI Plus, Pro and Ultra subscribers, with Docs support restricted to Pro and Ultra tiers.
cnet.com
· 2026-09-04
Meta unveiled Muse Voice Transcribe, its first real-time audio perception model, capable of transcribing over 20 speakers simultaneously while switching between more than 20 validated languages, including mid-sentence code-switching. CEO Mark Zuckerberg demoed the tool on X, noting it uses adaptive delay to balance speed and accuracy. The model is now available through Meta's Mac AI app, Muse Code, and its Model API at $3 per 1,000 audio minutes.
engadget.com
· 2026-09-01
Google has launched Gemini 3.5 Transcribe, a new voice-to-text AI model that automatically strips out filler words and stumbles to produce cleaner transcriptions. Already running Gboard's Rambler feature on the Pixel 11, the model is set to expand across Google's products. Google says it is roughly 70 percent faster than its predecessor, Chirp 3, and slightly more accurate, with a live-speech error rate of 5.5 percent versus 7.32 percent.
arstechnica.com
· 2026-08-26
Google unveiled Gemini 3.5 Transcribe, a new speech-to-text AI model that detects over 85 languages, strips filler words, and reformats casual speech into clean, structured text. It can also learn custom vocabulary, capture alphanumeric strings like order numbers, and assign dialogue to up to three speakers with timestamps from recorded audio.
engadget.com
· 2026-08-26
Google has released three new Gemini Audio models: 3.5 Transcribe, 3.5 Live, and 3.5 Live Experimental. Transcribe automatically strips out filler words like 'um' and 'uh', formats text, adapts to custom vocabulary and jargon, and can identify up to three speakers with word-level timestamps. The Live models improve handling of interruptions, multilingual recognition, and real-time reasoning narration for voice chat.
theverge.com
· 2026-08-26