Microsoft ships its first real-time streaming transcription model

MAI-Transcribe-2-Streaming returns first words in ~100ms across 60 languages, with paired low-latency voice and a reportedly cheaper local-AI box.

Nowline OCT 3 7:00 AM banner

Top AI stories from the last hour

Top AI stories from the last hour

Copy markdown

  • Live transcription, first words in ~100ms

    MAI-Transcribe-2-Streaming is Microsoft's first real-time speech-to-text model: partial results in just over 100ms, 60 languages with auto-detection, and words delivered 2x faster than its closest rival. At $0.54 per audio hour through year-end, live captions and voice agents are cheap to ship today.

  • A full voice loop from one vendor

    The same drop adds MAI-Voice-2.1 (TTS in 23 languages, 26 locales) and MAI-Voice-2.1-Flash, which hits 150ms end-to-end latency at $15 per million characters ($22 for standard). Pair streaming STT with Flash TTS for sub-second turn-taking. Callable now via Microsoft Foundry, MAI Playground, OpenRouter and Vercel, with LiveKit coming.

  • A cheaper DGX Spark for local models

    NVIDIA is reportedly adding a 64GB DGX Spark tier near $4,999, shipping late October, aimed at running ~100B-parameter models on your desk instead of the cloud. Treat the price as unconfirmed, but a lower entry point for local inference matters for anyone keeping data off hosted APIs.