Report: Microsoft tests MAI-Realtime, its own full-duplex voice model
An early-access scoop: bidirectional voice in 17 languages, built to slot in where Azure and Copilot voice now lean on OpenAI's GPT-Realtime.

Copy markdown
Full-duplex, 17 languages, two voices
MAI-Realtime reportedly listens and speaks at the same time — no turn-taking — across 17 languages, with two voices, Victoria and Grant, that early testers call noticeably more natural than Copilot's current voice mode.
The real story: cutting OpenAI out of the voice stack
Azure Speech's Voice Live API leans on OpenAI's GPT-Realtime for its speech-to-speech layer today; MAI-Realtime is built to fill that slot in-house. If it ships, the model under Copilot voice — and the Azure endpoint you may already call — changes vendors without you touching a line of code.
Turn-taking you can actually tune
Two endpointing setups reportedly ship: a 'Switchboard' mode driven by an MAI-Ears endpointer with inline control tokens, and a deterministic mode pairing silence-based endpointing with a Whisper semantic endpointer — real knobs for trading latency against accuracy in a live agent.
The catch: a preview sighting, not a launch
This surfaced inside Microsoft's MAI Playground, not a release — no public API, no timeline, and it won't sing or make non-speech sounds. It's a single-source scoop, so treat the specs as provisional until Microsoft confirms.
Elsewhere: the realtime-voice race adds a third full-duplex player
It lands a day after OpenAI's GPT Live went full-duplex. The pattern worth watching: Microsoft building its own realtime voice rather than reselling OpenAI's, the same decoupling it began with MAI-1 and MAI-Voice.