Gemini 3.8 Flash TTS
Shape each line's delivery with Gemini 3.8 Flash TTS — build custom voices, stage two-speaker scenes and export audio in 130 languages.
Support
Pro AI Tools
Explore elite tools including the Suno AI Music Generator.
MiniMax H3
MiniMax H3 AI Video Generator
Seedance 2.5
The Future of AI Video Is Here.

Seedance 2.0
The Future of AI Video Is Here.

Veo3.1
Create Stunning Videos with Veo3.1
FLUX 3 Video Generator

Kling 3.0
Next-Gen AI Video Generator
Grok Video Generator
Create Videos from Text or Images with AI
MiniMax H3 video generator
MiniMax H3 AI Video Generator

Gemini 3.8 Flash TTS: A New Standard for Directed Speech
Google's 23 September 2026 launch splits its speech line into a creative, direction-friendly model and a lean engine built for bulk audio.
- Two Models, Two Very Different JobsThe flagship tier rewards expressive, character-heavy reads, while Flash-Lite TTS keeps large batches of plain narration inexpensive.
- You Direct the Delivery Instead of Choosing a PresetPer-turn style notes, structured speech metadata and inline vocal cues shape tone, tempo, mood and accent.
- Design Voices From Text, or Clone With ConsentDescribe the voice you want in plain words, or mirror a real person using a reference clip paired with a consent recording.
How to Get Clean Results From Gemini 3.8 Flash TTS
Four habits that keep the spoken text clean and let the performance metadata do the acting.
What Gemini 3.8 Flash TTS Can Actually Do
From line-by-line direction to 130-language output, these are the capabilities that matter in daily production.
Line-by-Line Performance Control
Style per turn, plus in-script laughs, sighs, coughs, breaths and pauses — it feels like coaching a performer rather than picking a preset.
Voices Described in Plain Language
Prompt for age, personality, accent, texture or role, with more than 2,000 ready-made voices available through the Voices endpoint.
Replication That Requires Consent
You need a clean reference clip and a matching consent clip from the same adult, and every output carries SynthID watermarks plus C2PA credentials.
Built-In Two-Speaker Dialogue
Write the exchange once for podcasts, lessons, demos or game scenes — no manual line stitching afterwards.
Steady Voice Across Long Recordings
Google reports consistent identity, timbre, loudness and room tone through multi-minute narration and extended back-and-forth.
130 Languages With Regional Accents
The flagship reaches 130 languages versus 101 for Flash-Lite, including regional accents, minority dialects and IPA pronunciation overrides.
Gemini 3.8 Flash TTS: Questions Creators Ask
Quick answers about pricing, picking between the two tiers, benchmark scores and the safety rules behind the model.
What does Gemini 3.8 Flash TTS charge per minute of audio?
Roughly 1.35 cents for each audio minute, based on launch pricing of $0.50 per million input tokens and $9 per million output tokens.
When should I choose Flash TTS over Flash-Lite TTS?
Reach for the flagship when acting nuance and long recordings matter; Flash-Lite TTS wins on bulk jobs and low latency.
How does it score against rival voice models?
Google cites 71.4 on Hume's Voice Design Benchmark, while Voice Arena ranks it second with 1,260 Elo.
What changed from the Gemini 3.1 Flash TTS Preview?
Flash-Lite TTS takes the 3.1 preview's place and drops audio output pricing from $20 to $6 per million tokens.
Can I clone a voice, and what safeguards apply?
Cloning requires both a reference clip and a consent recording made by the same adult speaker.
Why does the model read my stage directions aloud?
Because the input is treated as a verbatim transcript — move lasting directions into the speech metadata instead.
Run Gemini 3.8 Flash TTS on Your Next Script
Try both tiers in the Gemini API or Google AI Studio; moving between Gemini 3.8 Flash TTS and Flash-Lite TTS is just a model ID swap. Weigh batch against priority inference before you lock in a production budget.
