Developers

Gemini 3.8 TTS API Guide

What changed in the Gemini 3.8 text-to-speech release, the model codes to call, and Google's list pricing. This page describes Google's API - not this site's playground request format.

Model codes

  • gemini-3.8-flash-tts - expressive tier, 130 languages
  • gemini-3.8-flash-lite-tts - low-latency tier, 101 languages, replaces gemini-3.1-flash-tts-preview

What changed in 3.8

Style lives in speech_metadata

In the 3.8 API, style direction (delivery style, accent guidance) belongs in the speech_metadata field of the request rather than being embedded in the text prompt. Keep style instructions out of the spoken text.

WAV is returned by default

3.8 returns WAV audio by default - no explicit output format field is needed for studio-grade files.

Inline tone tags

Performance cues can be placed inline in the text: [laughs], [sighs], [whispers], [excited], [short pause]. The model performs them in context.

Google list pricing

  • Text input: $0.50 per 1M tokens
  • Audio output, Flash: $9.00 per 1M tokens
  • Audio output, Flash-Lite: $6.00 per 1M tokens
  • Audio is billed at 25 tokens per second of generated speech
  • These are launch rates through December 31, 2026; they double on January 1, 2027

Voices and speakers

The API supports the 30 prebuilt voices (see the full list) and multi-speaker dialogue. Google's API also offers voice design, voice cloning and a 2,000+ voice library; this site intentionally exposes only the 30 prebuilt voices, 1-2 speakers, and inline tone tags.

For request schemas and full field documentation, refer to Google's official Gemini API docs. This guide intentionally covers only what Google has documented publicly - check the docs before writing integration code.

Try before you integrate

The playground here runs the same models - a quick way to audition voices, accents and tags before you build against the API.