Gemini 3.5 Transcribe: 2.6% WER speech-to-text, now in the API

Google's new speech model runs live or batch, cleans filler and formats text, diarizes up to 8 speakers across 85+ languages, and is free in preview.

Nowline AUG 27 5:00 AM banner

Top AI stories from the last hour

Top AI stories from the last hour

Copy markdown

  • 2.6% WER, and it cleans up the audio

    Non-streaming transcription hits a 2.6% word error rate (4.0% streaming), with roughly 70% lower latency than Chirp 3. Smart mode strips ums, stutters, and false starts and auto-formats the text, so you ship clean output instead of a raw dump.

  • Two model IDs, batch or real-time

    Call gemini-3.5-transcribe through the now-GA Interactions API for recorded files, or gemini-3.5-transcribe-live over the Live API for streaming audio. Both can trigger function calling to hand off tasks to other Gemini models.

  • 8-speaker diarization, timestamps, 85+ languages

    Flip on diarization_mode to label up to 8 speakers (3+ is experimental), request word-level timestamps, and auto-detect across 85+ locales, plus a custom-vocabulary hook for jargon and unusual spellings.

  • Free in preview — go build the notetaker

    It's live in public preview via Google AI Studio and Antigravity at no cost right now, and already powers Gboard Rambler, the Gemini macOS app, and soon Chrome voice typing. Enough to wire up a meeting notetaker or a real-time voice agent this weekend.