Grok Voice Think Fast 2.0: 0.70s speech-to-speech, default swaps Aug 5

xAI's voice model tops GPT-Realtime and Gemini on quality at $0.08/min — pin 1.0 first. Plus Google's Lyria 3.5 lands across Flow, Gemini and the API.

Nowline AUG 1 2:00 PM banner

Top AI stories from the last hour

Top AI stories from the last hour

Copy markdown

  • 0.70s to first audio — and it tops the field

    Think Fast 2.0 replies in 0.70 seconds (down from 1.25s in 1.0) and scores 82.9% on Artificial Analysis's Speech-to-Speech Quality Index, ahead of GPT-Realtime-2.1 (79.1%) and Gemini 3.1 Flash (69.5%). Real-time voice agents that don't feel laggy are finally practical.

  • Your default swaps Aug 5 — pin 1.0 to opt out

    The grok-voice-latest alias auto-upgrades to 2.0 on August 5. If your prompts or budgets are tuned to the old model, pin grok-voice-think-fast-1.0 now, or your production voice stack changes under you mid-week.

  • $0.08/min, with ~60% less reasoning overhead

    Audio runs $0.08 per minute, and xAI says it cut reasoning-token use ~60% versus 1.0, so the smarter model isn't proportionally costlier per turn. Some trackers frame the $0.08 as a step up from an earlier $0.05 tier — watch your bill after the cutover.

  • OpenAI-shaped Realtime API, MCP tools, 24 languages

    It ships a WebSocket Realtime API that mirrors OpenAI's event shape, plus MCP tool-calling, session resumption, and 24+ languages with keyterm biasing. You can port an existing Realtime voice agent with minimal rewiring.

  • Elsewhere: Google's Lyria 3.5 music model

    Google shipped Lyria 3.5 (Jul 29) across Flow Music, Gemini, AI Studio, Vertex AI and the Gemini API — sharper vocals and lyrics, 48kHz stereo, and direct tempo/duration control past the old 3-minute cap. A weekend soundtrack generator just got usable.