MAI-Transcribe-2-Streaming: the new #1 real-time STT at 2.5% WER
Microsoft's streaming model hits 0.13s latency across 60 languages and ships beside a voice-TTS pair—enough to build a sub-second voice agent this weekend.

Copy markdown
#1 streaming STT, 2.5% WER, ~0.13s
MAI-Transcribe-2-Streaming tops all 38 models on Artificial Analysis's real-time WER index at 2.5% error, emitting first partial text in ~0.12s and committed final text ~0.13s after you stop speaking. Microsoft says words land roughly 2x faster than the next model—for a voice agent, that's the gap between feeling live and feeling laggy.
$0.54/hr, 60 languages, OpenAI-Realtime-compatible
Streaming runs $0.54 per audio-hour (intro rate through Dec 31), covers 60 languages with continuous auto language-detection, and exposes an OpenAI Realtime-compatible WebSocket plus the Azure Speech SDK and Vercel, with LiveKit 'coming soon.' If you already wired an OpenAI realtime endpoint, swapping this in is close to drop-in.
Pair it with MAI-Voice-2.1 for the full loop
The companion MAI-Voice-2.1 TTS shipped alongside it: the Flash tier renders 45s of audio at ~150ms latency for $15 per 1M characters across 23 languages ($22/1M for standard). Transcribe-in, synthesize-out gives you both halves of a real-time voice agent from a single vendor.
The catch: preview, no SLA, thin metadata
It's public preview with no SLA and 'not recommended for production,' and it omits word-level timestamps, confidence scores, and per-language accuracy breakdowns. If you don't need the top score, Grok Voice Transcribe 2.0 is cheaper at $0.20/hr for 2.73% WER—worth a bake-off before you commit.