cd /news/artificial-intelligence/exclusive-microsoft-tests-new-mai-re… · home topics artificial-intelligence article
[ARTICLE · art-83762] src=testingcatalog.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Microsoft tests new MAI Realtime voice model

Microsoft is testing a new native real-time voice model, MAI Realtime, which has appeared as a hidden early-access entry in its MAI Playground. The bidirectional, full-duplex system, which listens and speaks simultaneously, is available to a small group of partners and features two natural-sounding voices, Victoria and Grant. This first-party model would fill a gap in Microsoft's speech offerings, potentially replacing the GPT-Realtime model in Azure Speech's Voice Live API and reducing reliance on OpenAI components.

read2 min views1 publishedAug 2, 2026
Microsoft tests new MAI Realtime voice model
Image: Testingcatalog (auto-discovered)

Microsoft appears to be preparing its first native real-time voice model, referred to as MAI Realtime, which has surfaced as a hidden early-access entry in the company’s MAI Playground. The listing suggests a small group of partners already has hands-on access, and what is visible points to a bidirectional, full-duplex system, one that listens and speaks at the same time rather than trading turns, placing it in direct comparison with OpenAI’s GPT Live 1 or Sesame.

Two voices are present so far, Victoria and Grant, both noticeably more natural than what Copilot’s voice mode currently delivers. Language can be pinned explicitly or left on automatic detection, and the model switches languages mid-conversation without losing its footing. Turn-taking is configurable through two listener options: a Switchboard mode built around an MAI-Ears endpointer driven by inline control tokens, and a deterministic setup that pairs silence-based endpointing with a Whisper semantic endpointer.

The practical difference between them is subtle in use, though interruptions are handled cleanly and response latency is low. The model does not sing or produce non-speech sounds, which keeps it squarely a conversational system rather than a general audio generator. A debug panel exposes live latency figures, model thoughts and processing steps, and sample sharing looks set to arrive for playground users once access widens.

That would fill a conspicuous gap. Every MAI speech model shipped so far runs in one direction: MAI-Voice-2 and its Flash variant for synthesis, and MAI-Transcribe-1.5 for recognition, while the speech-to-speech layer in Azure Speech’s Voice Live API still relies on the GPT-Realtime model. A first-party full-duplex model would close that dependency for Mustafa Suleyman’s superintelligence team, which shipped seven in-house models at Build 2026 and has been steadily swapping OpenAI components out of Copilot, Teams and Bing. Microsoft Foundry is the likely developer destination, with Copilot voice the obvious consumer surface, though no timeline has been attached to either.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @microsoft 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/exclusive-microsoft-…] indexed:0 read:2min 2026-08-02 ·