Report: Microsoft tests MAI-Realtime, its own full-duplex voice model

An early-access scoop: bidirectional voice in 17 languages, built to slot in where Azure and Copilot voice now lean on OpenAI's GPT-Realtime.

Nowline AUG 4 4:00 AM banner

Top AI stories from the last hour

Top AI stories from the last hour

Copy markdown

  • Full-duplex, 17 languages, two voices

    MAI-Realtime reportedly listens and speaks at the same time — no turn-taking — across 17 languages, with two voices, Victoria and Grant, that early testers call noticeably more natural than Copilot's current voice mode.

  • The real story: cutting OpenAI out of the voice stack

    Azure Speech's Voice Live API leans on OpenAI's GPT-Realtime for its speech-to-speech layer today; MAI-Realtime is built to fill that slot in-house. If it ships, the model under Copilot voice — and the Azure endpoint you may already call — changes vendors without you touching a line of code.

  • Turn-taking you can actually tune

    Two endpointing setups reportedly ship: a 'Switchboard' mode driven by an MAI-Ears endpointer with inline control tokens, and a deterministic mode pairing silence-based endpointing with a Whisper semantic endpointer — real knobs for trading latency against accuracy in a live agent.

  • The catch: a preview sighting, not a launch

    This surfaced inside Microsoft's MAI Playground, not a release — no public API, no timeline, and it won't sing or make non-speech sounds. It's a single-source scoop, so treat the specs as provisional until Microsoft confirms.

  • Elsewhere: the realtime-voice race adds a third full-duplex player

    It lands a day after OpenAI's GPT Live went full-duplex. The pattern worth watching: Microsoft building its own realtime voice rather than reselling OpenAI's, the same decoupling it began with MAI-1 and MAI-Voice.