Google has launched Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, two new text-to-speech models designed to expand voice generation capabilities beyond static presets.

According to Google, Gemini 3.8 Flash TTS is built for deep creative direction and character design, allowing creators to generate entirely new voices from scratch using natural language prompts. The model supports voice creation across gaming, audiobooks, podcasts, and interactive media, with granular control over acting cues, pacing, dialect shifts, and conversational elements like backchanneling. Gemini 3.8 Flash-Lite TTS is optimized for high-volume, cost-efficient applications such as dubbing and voice agents, with fine-grained control over tone and pacing.
Both models offer expansive voice libraries and customization options. Users can access over 2,000 production-ready voices across more than 100 languages and dialects, including regional varieties like Mexican Spanish, Quebec French, and Scots English. For original content, the Gemini 3.8 Flash TTS model enables generative voice design to create bespoke voices by customizing role, accent, and voice characteristics through natural language prompting.
A voice replication feature allows users to recreate consistent vocal profiles from a 30-second audio sample, supported by consent verification, SynthID watermarking, and C2PA credentials. Google says a voice remixing feature is coming soon, enabling fine-tuning of timbre, pitch, pace, and accent using prompts.
Both models include performance direction capabilities, allowing line-by-line control over delivery. Long-form generation maintains voice quality and natural pacing across hours of continuous audio with minimal speaker drift. Native two-speaker scene staging enables multi-turn conversations from a single script, while scripted vocal bursts and backchanneling add realistic conversational texture using non-verbal cues.
According to Google, Gemini 3.8 Flash TTS secured the #1 overall spot on Hume AI’s Voice Design Benchmark and leading positions in accent modeling. Both models ranked #1 and #2 respectively on Hume AI’s Overall Quality Index. In blind human preference evaluations on Voice Arena, the models secured top positions in global languages including Japanese, Brazilian Portuguese, Vietnamese, Modern Standard Arabic, Mexican Spanish, and Hindi.
Google has integrated consent verification and transparency measures into the voice creation system. Users must provide verbal consent recordings from voice owners before creating replicated voices. All audio generated by Gemini Audio models receives SynthID watermarking, an imperceptible watermark designed to keep AI-generated speech detectable.
Developers can access the models through Google AI Studio and the Gemini API starting today. Gemini 3.8 Flash TTS is rolling out to Gemini Notebook for all users, with enterprise API access coming soon. Gemini 3.8 Flash-Lite TTS is available in the Gemini API and Google AI Studio, with enterprise access forthcoming. Integration partners include Agora, LiveKit, Pipecat, Vercel, Figma, HeyGen, Linguana, Wondercraft, 99.co, and Ollang.
Key facts
- Google introduced Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS for expressive voice generation
- Models support generative voice design across over 100 languages and dialects with access to 2,000+ production-ready voices
- Voice replication requires consent verification from the voice owner via a 30-second audio sample
- Gemini 3.8 Flash TTS ranked #1 on Hume AI’s Voice Design Benchmark and #1 on Overall Quality Index
- All generated audio is watermarked with SynthID to remain detectable as AI-generated content
- Models enable line-by-line performance direction, long-form generation, and dual-speaker screenplay control
