Gemini 3.8 TTS lands on the API: design any voice from a text prompt
Two models, GA on the Gemini API and AI Studio: steer delivery line by line, stage two speakers, clone from 30 seconds, all cheap enough to embed.

Copy markdown
Voice design from plain English
Describe the voice you want — accent, age, mood — and Gemini 3.8 Flash TTS builds it, no picking from a fixed menu. It also takes line-by-line delivery direction, two-speaker scene staging, and non-verbal cues, across 30 prebuilt studio voices (plus any you design) and 130 languages, generally available now on the Gemini API and AI Studio.
The price makes speech embeddable
Flash TTS runs $0.50 in / $9.00 audio out per million tokens on a promo through Dec 31 (then $1/$18) — roughly 54 cents per hour of generated audio. Flash-Lite TTS drops that to $0.50/$6.00 (then $1/$12) with 101 languages, for high-volume or latency-sensitive work.
Clone a voice from 30 seconds
Both models can clone a voice from a 30-second sample, gated behind consent verification, with a SynthID watermark baked into every clip. Enough to spin up a consistent narrator or character voice without booking a recording studio.
What to build this weekend
Steerable, cheap TTS unlocks projects that were fiddly before: audiobook narration with directed pacing, multilingual dubbing of your own videos, or an in-app agent voice that shifts tone per line. No waitlist — it's live today.