Gemini 3.8 Flash TTS

Shape each line's delivery with Gemini 3.8 Flash TTS — build custom voices, stage two-speaker scenes and export audio in 130 languages.

Gemini 3.8 Flash TTS
Gemini 3.8 Flash TTS handles expressive, character-driven reads; its Flash-Lite sibling keeps bulk narration affordable.
AI Video Prompt Generator

Support

Pro AI Tools

Explore elite tools including the Suno AI Music Generator.

placeholder hero

Gemini 3.8 Flash TTS: A New Standard for Directed Speech

Google's 23 September 2026 launch splits its speech line into a creative, direction-friendly model and a lean engine built for bulk audio.

  • Two Models, Two Very Different Jobs
    The flagship tier rewards expressive, character-heavy reads, while Flash-Lite TTS keeps large batches of plain narration inexpensive.
  • You Direct the Delivery Instead of Choosing a Preset
    Per-turn style notes, structured speech metadata and inline vocal cues shape tone, tempo, mood and accent.
  • Design Voices From Text, or Clone With Consent
    Describe the voice you want in plain words, or mirror a real person using a reference clip paired with a consent recording.

How to Get Clean Results From Gemini 3.8 Flash TTS

Four habits that keep the spoken text clean and let the performance metadata do the acting.

What Gemini 3.8 Flash TTS Can Actually Do

From line-by-line direction to 130-language output, these are the capabilities that matter in daily production.

Line-by-Line Performance Control

Style per turn, plus in-script laughs, sighs, coughs, breaths and pauses — it feels like coaching a performer rather than picking a preset.

Voices Described in Plain Language

Prompt for age, personality, accent, texture or role, with more than 2,000 ready-made voices available through the Voices endpoint.

Replication That Requires Consent

You need a clean reference clip and a matching consent clip from the same adult, and every output carries SynthID watermarks plus C2PA credentials.

Built-In Two-Speaker Dialogue

Write the exchange once for podcasts, lessons, demos or game scenes — no manual line stitching afterwards.

Steady Voice Across Long Recordings

Google reports consistent identity, timbre, loudness and room tone through multi-minute narration and extended back-and-forth.

130 Languages With Regional Accents

The flagship reaches 130 languages versus 101 for Flash-Lite, including regional accents, minority dialects and IPA pronunciation overrides.

FAQ

Gemini 3.8 Flash TTS: Questions Creators Ask

Quick answers about pricing, picking between the two tiers, benchmark scores and the safety rules behind the model.

1

What does Gemini 3.8 Flash TTS charge per minute of audio?

Roughly 1.35 cents for each audio minute, based on launch pricing of $0.50 per million input tokens and $9 per million output tokens.

2

When should I choose Flash TTS over Flash-Lite TTS?

Reach for the flagship when acting nuance and long recordings matter; Flash-Lite TTS wins on bulk jobs and low latency.

3

How does it score against rival voice models?

Google cites 71.4 on Hume's Voice Design Benchmark, while Voice Arena ranks it second with 1,260 Elo.

4

What changed from the Gemini 3.1 Flash TTS Preview?

Flash-Lite TTS takes the 3.1 preview's place and drops audio output pricing from $20 to $6 per million tokens.

5

Can I clone a voice, and what safeguards apply?

Cloning requires both a reference clip and a consent recording made by the same adult speaker.

6

Why does the model read my stage directions aloud?

Because the input is treated as a verbatim transcript — move lasting directions into the speech metadata instead.

Run Gemini 3.8 Flash TTS on Your Next Script

Try both tiers in the Gemini API or Google AI Studio; moving between Gemini 3.8 Flash TTS and Flash-Lite TTS is just a model ID swap. Weigh batch against priority inference before you lock in a production budget.