Microsoft ships its first real-time streaming transcription model
MAI-Transcribe-2-Streaming returns first words in ~100ms across 60 languages, with paired low-latency voice and a reportedly cheaper local-AI box.

Copy markdown
Live transcription, first words in ~100ms
MAI-Transcribe-2-Streaming is Microsoft's first real-time speech-to-text model: partial results in just over 100ms, 60 languages with auto-detection, and words delivered 2x faster than its closest rival. At $0.54 per audio hour through year-end, live captions and voice agents are cheap to ship today.
A full voice loop from one vendor
The same drop adds MAI-Voice-2.1 (TTS in 23 languages, 26 locales) and MAI-Voice-2.1-Flash, which hits 150ms end-to-end latency at $15 per million characters ($22 for standard). Pair streaming STT with Flash TTS for sub-second turn-taking. Callable now via Microsoft Foundry, MAI Playground, OpenRouter and Vercel, with LiveKit coming.
A cheaper DGX Spark for local models
NVIDIA is reportedly adding a 64GB DGX Spark tier near $4,999, shipping late October, aimed at running ~100B-parameter models on your desk instead of the cloud. Treat the price as unconfirmed, but a lower entry point for local inference matters for anyone keeping data off hosted APIs.