Gemini 3.8 TTS lands on the API: design any voice from a text prompt

Two models, GA on the Gemini API and AI Studio: steer delivery line by line, stage two speakers, clone from 30 seconds, all cheap enough to embed.

Nowline SEP 25 3:00 AM banner

Top AI stories from the last hour

Top AI stories from the last hour

Copy markdown

  • Voice design from plain English

    Describe the voice you want — accent, age, mood — and Gemini 3.8 Flash TTS builds it, no picking from a fixed menu. It also takes line-by-line delivery direction, two-speaker scene staging, and non-verbal cues, across 30 prebuilt studio voices (plus any you design) and 130 languages, generally available now on the Gemini API and AI Studio.

  • The price makes speech embeddable

    Flash TTS runs $0.50 in / $9.00 audio out per million tokens on a promo through Dec 31 (then $1/$18) — roughly 54 cents per hour of generated audio. Flash-Lite TTS drops that to $0.50/$6.00 (then $1/$12) with 101 languages, for high-volume or latency-sensitive work.

  • Clone a voice from 30 seconds

    Both models can clone a voice from a 30-second sample, gated behind consent verification, with a SynthID watermark baked into every clip. Enough to spin up a consistent narrator or character voice without booking a recording studio.

  • What to build this weekend

    Steerable, cheap TTS unlocks projects that were fiddly before: audiobook narration with directed pacing, multilingual dubbing of your own videos, or an in-app agent voice that shifts tone per line. No waitlist — it's live today.