Google launches Gemini 3.8 Flash and Flash-Lite TTS voice-cloning models
Google has released Gemini 3.8 Flash TTS and Flash-Lite TTS, new text-to-speech models that can generate custom voices from text descriptions, replicate a voice from a 30-second sample with permission, and follow directions for accent, pacing, emotion and multi-speaker dialogue. The models support over 100 languages and more than 2,000 ready-made voices, and are rolling out in Gemini Notebook and Google Vids as well as to developers via Google AI Studio and the Gemini API.
GoKawiil's interpretation of the reporting above, not reported fact.
By emphasizing performance-level control—pacing, whispers, sighs, two-speaker dialogue—Google appears to be positioning these models for uses like audiobooks, dubbing, podcasts and voice agents that previously required more manual production. Strong scores on Hume AI's benchmarks suggest Google is trying to establish a competitive edge in synthetic voice quality, though real-world performance across languages and use cases remains to be tested by developers.
- Gemini 3.8 Flash TTS and Flash-Lite TTS are now rolling out to Gemini Notebook, Google Vids, AI Studio, and the Gemini API.
- The models can clone a voice from a 30-second sample and support over 100 languages with 2,000+ preset voices.
- Google reports top rankings for the models on Hume AI's Voice Design and quality benchmarks.
Source: androidauthority.com, 2026-09-24
Published there as: “Gemini can now clone your voice and perform scripts like an actor”
Read the original report → The summary and analysis above are GoKawiil's own, written from reporting by the source above. Facts and quotes belong to the original publisher.