Model
Gemini 3.8 Flash TTS
The expressive tier of Google's September 23, 2026 TTS release. Model code gemini-3.8-flash-tts, built for audiobooks, dialogue and dialect work across 130 languages.
What it's for
Gemini 3.8 Flash TTS is the expressive model in the 3.8 lineup. Where Flash-Lite optimizes for latency and cost, Flash optimizes for performance quality - the version you want for long-form narration, character dialogue and accented or dialect-heavy reads.
Key facts
- Model code:
gemini-3.8-flash-tts - Languages: 130
- Google list price: $0.50 per 1M text input tokens, $9.00 per 1M audio output tokens (launch rate through Dec 31, 2026; doubles Jan 1, 2027). Audio bills at 25 tokens/second.
- Output: WAV by default
- On this site: 2.12 credits per 1,000 characters, billed per 100 characters (a 200-character line costs 1 credit); MP3 and WAV downloads
Controls on this site
The playground exposes the expressive surface: 30 prebuilt voices, per-speaker accent (from Neutral to British RP to Australian), style (Vocal Smile, Newscaster, Whisper, Empathetic, Promo/Hype, Deadpan), pace (Natural, Rapid Fire, The Drift, Staccato), an audio_profile free-text field, and a Director panel for scene, sample_context and temperature (0-2).
Inline tags
Gemini 3.8 understands inline performance tags inside the text: [laughs], [sighs], [whispers], [excited], [short pause]. Insert them with the chips under each text field.
Flash or Flash-Lite?
Use Flash when quality matters more than speed - narration, dialogue, creative content. Use Flash-Lite for high-volume or latency-sensitive jobs.