Google launches Gemini 3.8 flash and flash-lite text-to-speech models for enhanced audio creation
Google has introduced two new text-to-speech models, Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, designed to transform voice generation from static presets into dynamic creative tools.
These models enable creators, developers, and enterprises to produce more expressive audio experiences, enhancing products like Gemini Notebook and Google Vids.
The Flash TTS model is tailored for deep creative and character design, allowing users to generate custom voices with natural language prompts, while Flash-Lite TTS is optimized for large-scale, cost-effective use cases. Both models offer fine-grained control over tone, rhythm, and performance nuances.
The Flash TTS model achieved top scores in Hume AI’s benchmarks, outperforming competitors in accent modeling and overall quality. Voice replication features include SynthID watermarks and C2PA credentials to ensure transparency and protect voice talent.
Developers can now access these tools via Google AI Studio, with enterprise access coming soon through Gemini Enterprise APIs.