Grok Voice Think Fast 2.0: 0.70s speech-to-speech, default swaps Aug 5
xAI's voice model tops GPT-Realtime and Gemini on quality at $0.08/min — pin 1.0 first. Plus Google's Lyria 3.5 lands across Flow, Gemini and the API.

Copy markdown
0.70s to first audio — and it tops the field
Think Fast 2.0 replies in 0.70 seconds (down from 1.25s in 1.0) and scores 82.9% on Artificial Analysis's Speech-to-Speech Quality Index, ahead of GPT-Realtime-2.1 (79.1%) and Gemini 3.1 Flash (69.5%). Real-time voice agents that don't feel laggy are finally practical.
Your default swaps Aug 5 — pin 1.0 to opt out
The grok-voice-latest alias auto-upgrades to 2.0 on August 5. If your prompts or budgets are tuned to the old model, pin grok-voice-think-fast-1.0 now, or your production voice stack changes under you mid-week.
$0.08/min, with ~60% less reasoning overhead
Audio runs $0.08 per minute, and xAI says it cut reasoning-token use ~60% versus 1.0, so the smarter model isn't proportionally costlier per turn. Some trackers frame the $0.08 as a step up from an earlier $0.05 tier — watch your bill after the cutover.
OpenAI-shaped Realtime API, MCP tools, 24 languages
It ships a WebSocket Realtime API that mirrors OpenAI's event shape, plus MCP tool-calling, session resumption, and 24+ languages with keyterm biasing. You can port an existing Realtime voice agent with minimal rewiring.
Elsewhere: Google's Lyria 3.5 music model
Google shipped Lyria 3.5 (Jul 29) across Flow Music, Gemini, AI Studio, Vertex AI and the Gemini API — sharper vocals and lyrics, 48kHz stereo, and direct tempo/duration control past the old 3-minute cap. A weekend soundtrack generator just got usable.