Microsoft appears to be preparing its first native real-time voice model, referred to as MAI Realtime, which has surfaced as a hidden early-access entry in the company’s MAI Playground. The listing suggests a small group of partners already has hands-on access, and what is visible points to a bidirectional, full-duplex system, one that listens and speaks at the same time rather than trading turns, placing it in direct comparison with OpenAI’s GPT Live 1 or Sesame.
Two voices are present so far, Victoria and Grant, both noticeably more natural than what Copilot’s voice mode currently delivers. Language can be pinned explicitly or left on automatic detection, and the model switches languages mid-conversation without losing its footing. Turn-taking is configurable through two listener options: a Switchboard mode built around an MAI-Ears endpointer driven by inline control tokens, and a deterministic setup that pairs silence-based endpointing with a Whisper semantic endpointer.
The practical difference between them is subtle in use, though interruptions are handled cleanly and response latency is low. The model does not sing or produce non-speech sounds, which keeps it squarely a conversational system rather than a general audio generator. A debug panel exposes live latency figures, model thoughts and processing steps, and sample sharing looks set to arrive for playground users once access widens.
That would fill a conspicuous gap. Every MAI speech model shipped so far runs in one direction: MAI-Voice-2 and its Flash variant for synthesis, and MAI-Transcribe-1.5 for recognition, while the speech-to-speech layer in Azure Speech’s Voice Live API still relies on the GPT-Realtime model. A first-party full-duplex model would close that dependency for Mustafa Suleyman’s superintelligence team, which shipped seven in-house models at Build 2026 and has been steadily swapping OpenAI components out of Copilot, Teams and Bing. Microsoft Foundry is the likely developer destination, with Copilot voice the obvious consumer surface, though no timeline has been attached to either.