Microsoft Quietly Tests MAI-Realtime, a Voice AI That Talks While It Listens Microsoft is quietly testing MAI-Realtime, a full-duplex voice AI that can talk and listen simultaneously, in limited early access on its MAI Playground, according to an exclusive report from TestingCatalog published August 2, 2026. The model, developed under Mustafa Suleyman's Microsoft AI division, supports 17 languages and aims to reduce Copilot's reliance on OpenAI's GPT-Realtime, though Microsoft has not publicly announced it or confirmed pricing or launch details. Microsoft is quietly testing MAI-Realtime, a voice AI that can talk and listen at the same time instead of waiting its turn, built to cut Copilot's reliance on OpenAI. You've probably talked to a voice assistant that pauses awkwardly, waits for you to stop, then answers a beat too late. MAI-Realtime is Microsoft's attempt to kill that pause entirely. According to an exclusive report from TestingCatalog published August 2, 2026, the model is running in limited early access on Microsoft's MAI Playground, and it's the company's first native full-duplex voice system, one that listens while it's speaking, the way two people actually talk to each other. That sounds like a small engineering detail. It isn't. Every mainstream voice assistant today, including Microsoft's own MAI Voice 2, still runs on turn-taking: you speak, it processes, then it replies. Full-duplex removes the handoff. You can interrupt mid-sentence, change the subject, or talk over the model, and TestingCatalog reports the interruption handling is clean with low response latency. Two voices are available so far, Victoria and Grant, and TestingCatalog describes both as noticeably more natural than what Copilot's voice mode delivers today. That's just the surface. The model reportedly runs 17 languages, including English, German, Spanish, French, Italian, Portuguese, Japanese, Korean, Chinese, Dutch, Hindi, Indonesian, Arabic, Russian, Turkish, Vietnamese and Thai, and it can detect and switch between them mid-conversation without losing context. It can also call tools while you talk, including live web search, so a conversation can pull in a fact or a headline without breaking the exchange. What it won't do, per TestingCatalog, is sing or produce non-speech sounds. It's built for conversation, not for party tricks. MAI-Realtime doesn't exist in isolation. It's the newest piece of an in-house speech lineup that already includes MAI-Voice-1, MAI-Voice-2 and MAI-Transcribe-1, developed under Mustafa Suleyman's Microsoft AI division. MAI-Voice-2 launched earlier this year priced at $22 per million characters on Azure AI Foundry, positioned between OpenAI's own text-to-speech pricing at $15 per million characters and ElevenLabs' subscription tiers. Frankly, the pricing detail matters less than the intent behind it. Microsoft wants to own the model, the safety policy and the bill for voice, not lease all three from OpenAI. That intent has been building for a while. The Register reported in August 2025 that Microsoft was rolling out home-made models even as it kept negotiating its OpenAI relationship, and MAI-Realtime looks like the voice half of that plan reaching its most direct form yet. TestingCatalog's reporting frames the model as filling a specific gap: reducing Copilot's dependence on OpenAI's GPT-Realtime model, the system that currently powers full-duplex voice inside ChatGPT's Advanced Voice Mode. Google isn't standing still either. Gemini Live has offered its own full-duplex-style conversation mode for over a year. The real fight here isn't Microsoft against OpenAI in the abstract. It's three companies racing to own the moment when talking to an AI stops feeling like using a walkie-talkie. MAI-Realtime hasn't been announced publicly. Microsoft hasn't published a model card, confirmed pricing, named launch regions, or added it to the public MAI Playground catalog. TestingCatalog's account rests on access granted to a small group of partners testing a hidden listing, not on any Microsoft statement, so treat the feature list as a preview rather than a locked spec until Microsoft says otherwise. Still, the direction is clear enough. Chat got commoditized. Image generation got commoditized. Voice, the interface layer that decides whether an AI assistant feels like a tool or a conversation, is the next thing every major lab wants to own outright. Microsoft just showed its hand. Also read: OpenAI Fires Back at Apple Trade Secret Lawsuit With Its Own Email Evidence https://startupfortune.com/openai-fires-back-at-apple-trade-secret-lawsuit-with-its-own-email-evidence/ • China's Flood of Cheap AI Models Is Pricing Smaller US Labs Out https://startupfortune.com/chinas-flood-of-cheap-ai-models-is-pricing-smaller-us-labs-out/ • Alibaba-Backed 3D AI Startup Vast Is Weighing a Hong Kong Listing https://startupfortune.com/alibaba-backed-3d-ai-startup-vast-is-weighing-a-hong-kong-listing/