Gemini 3.5 Transcribe: 2.6% WER speech-to-text, now in the API
Google's new speech model runs live or batch, cleans filler and formats text, diarizes up to 8 speakers across 85+ languages, and is free in preview.

Copy markdown
2.6% WER, and it cleans up the audio
Non-streaming transcription hits a 2.6% word error rate (4.0% streaming), with roughly 70% lower latency than Chirp 3. Smart mode strips ums, stutters, and false starts and auto-formats the text, so you ship clean output instead of a raw dump.
Two model IDs, batch or real-time
Call gemini-3.5-transcribe through the now-GA Interactions API for recorded files, or gemini-3.5-transcribe-live over the Live API for streaming audio. Both can trigger function calling to hand off tasks to other Gemini models.
8-speaker diarization, timestamps, 85+ languages
Flip on diarization_mode to label up to 8 speakers (3+ is experimental), request word-level timestamps, and auto-detect across 85+ locales, plus a custom-vocabulary hook for jargon and unusual spellings.
Free in preview — go build the notetaker
It's live in public preview via Google AI Studio and Antigravity at no cost right now, and already powers Gboard Rambler, the Gemini macOS app, and soon Chrome voice typing. Enough to wire up a meeting notetaker or a real-time voice agent this weekend.