
Google introduces Gemini 3.8 text-to-speech models with granular audio controls
Google has expanded its audio generation capabilities with two new text-to-speech models designed for expressive voice design and high-volume deployment. The release includes advanced scripting tools, multi-speaker staging, and built-in safety safeguards like SynthID watermarking.
Published by Jin · 2 min read · 24 SEPT 2026
Google has introduced two new text-to-speech models to the Gemini family, expanding voice generation beyond static presets into dynamic creative tools. The newly released Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS models offer enhanced expressiveness, granular performance direction, and broad language support across more than 100 languages and dialects.
Custom voice creation and design
Gemini 3.8 Flash TTS functions as a vocal studio, enabling creators to build bespoke voices from scratch using natural language prompts. Users can define specific roles, accents, and emotional tones, transforming text prompts into distinct vocal personas.
In addition to generating entirely new voices, the system supports voice replication using a 30-second audio sample. To maintain security, replication requires consent verification, matching a verbal recording from the voice owner before generation begins. SynthID watermarking and C2PA credentials are also integrated to ensure content transparency and prevent unauthorized use.
Granular performance and multi-speaker staging
Both models give developers and creators precise control over line-by-line delivery. Users can write stage directions or rely on natural script cues to adjust pacing, emotion, and conversational nuances.
Key capabilities include:
- Long-form generation for audiobooks and podcasts with minimal speaker drift.
- Native two-speaker scene staging for seamless multi-turn conversations from a single script.
- Scripted vocal bursts and non-verbal cues, such as pauses and active-listening interjections.
Availability and benchmarks
In benchmark evaluations, Gemini 3.8 Flash TTS secured the top position on Hume AI’s Voice Design Benchmark with a score of 71.4, while both models ranked at the top of the Overall Quality Index for expressiveness and reliability.
The new models are rolling out across multiple platforms. Developers can access both variants immediately through Google AI Studio and the Gemini API. Enterprise integration is forthcoming via Gemini Enterprise, while consumer features are deploying to Gemini Notebook and Google Vids.
Source — Original announcement ↗
Worth a read?
Comments · 0