Gemini 3.5 Transcribe: 2.6% WER speech-to-text now in the API
Google's new speech model runs from AI Studio at ~$0.005/min, adds speaker IDs and word-level timestamps, and streams live captions at sub-second latency.

Copy markdown
2.6% WER, and it cleans up your speech
Gemini 3.5 Transcribe posts a 2.6% word error rate non-streaming (4.0% streaming) and cuts time-to-final by roughly 70% versus Chirp 3. It drops filler words and honors self-corrections, so “send it to, uh, no—forward it to Sam” comes out clean.
Two endpoints, ~$0.005 a minute
Call gemini-3-5-transcribe for pre-recorded audio—speaker attribution up to three people, word-level timestamps—or gemini-3-5-transcribe-live over the Live API for sub-second streaming captions. Reported blended cost lands near $0.005/min pre-recorded and $0.009/min live, with a free tier, undercutting most hosted speech-to-text.
85+ languages, custom vocab, function calls
Automatic language detection spans 85+ languages with regional-accent handling, plus a custom-vocabulary slot for domain jargon and function calling to hand work to other Gemini models mid-transcript. It's in public preview now via the Gemini API in Google AI Studio.
Weekend build: a meeting-to-JSON agent
With speaker IDs and word timestamps on tap, a recorded standup can become structured, speaker-labeled action items for well under a dollar an hour—no separate diarization service and no per-seat SaaS in the loop.