MAI-Transcribe-2-Streaming: Real-Time Voice Agents, Completed Microsoft launched MAI-Transcribe-2-Streaming on October 1, a WebSocket-based streaming speech-to-text model that ranked first for both Final Transcript and First Partial Transcript accuracy across 38 models on Artificial Analysis's AA-WER Streaming benchmark with a 2.5% word error rate. The model ships alongside the new MAI-Voice-2.1-Flash text-to-speech model and the generally available Azure Voice Live orchestration layer, letting Azure developers run the full real-time voice agent loop on first-party infrastructure instead of routing audio through Deepgram or AssemblyAI. MAI-Transcribe-2-Streaming delivers final transcripts 0.13 seconds after end of speech and begins partial hypotheses around 100 ms after audio starts, versus the 5-7% production word error rate posted by Deepgram Nova-3. Microsoft launched three voice models on October 1 that do something its Azure speech stack could not do before: run the entire real-time voice agent loop — hear, transcribe, respond, speak — entirely on first-party infrastructure. The new piece is MAI-Transcribe-2-Streaming , a WebSocket-based streaming speech-to-text model that ranked first for accuracy across 38 models in independent testing. Paired with the newly released MAI-Voice-2.1-Flash TTS and the generally available Azure Voice Live orchestration layer, Azure developers now have a complete voice agent pipeline without routing audio through Deepgram, AssemblyAI, or any other third-party vendor. The Gap That Existed Until Yesterday Building a real-time voice agent has always required assembling three components: STT, LLM, and TTS. The LLM part has been well-served for a while. The TTS part got serviceable. The STT piece was the persistent gap for Azure-native developers: OpenAI Whisper does not stream it is a batch endpoint with a 25 MB file cap , and Azure’s own previous neural speech models did not offer true streaming transcription with partial hypotheses. The standard workaround was to add Deepgram or AssemblyAI to the stack — which works fine but introduces a third-party dependency, a separate pricing agreement, and an additional latency hop. MAI-Transcribe-2-Streaming is the answer to that problem. It streams. It is first-party. And its accuracy numbers are not typical launch-day marketing figures. The Accuracy Numbers Are Independently Verified Artificial Analysis ran the model against 38 competitors on its AA-WER Streaming benchmark and put MAI-Transcribe-2-Streaming at number one for both Final Transcript accuracy and First Partial Transcript accuracy , with a 2.5% word error rate. The final transcript is delivered 0.13 seconds after end of speech. Partial hypotheses start arriving around 100 ms after audio begins — meaning an LLM sitting downstream can start reasoning before the speaker has even finished their sentence. For context: Deepgram Nova-3 https://deepgram.com/learn/voice-agent-architecture-stt-llm-tts-pipeline-design , currently the most widely deployed streaming STT in voice agents, posts 5–7% WER in production. That is a meaningful gap. Mishearing “cancel” as “confirm,” or “staging” as “staying,” at 6% WER happens regularly enough to matter in production workflows. The 2.5% figure is not a rounding improvement — it is a substantively different accuracy tier. How to Integrate There are two integration paths. The first is a direct WebSocket connection using Microsoft’s Realtime API: wss://{resource}.services.ai.azure.com/mai/v1/realtime?intent=transcription Header: api-key: {your-key} Push 100–200 ms audio packets into the socket, and the model returns two event types: partial updated hypothesis as speech continues and final committed segment after a pause . The second path is the Azure Speech SDK https://learn.microsoft.com/en-us/azure/ai-services/speech-service/mai-transcribe-2-streaming , which wraps the WebSocket with retry logic, connection management, and audio buffering — the better starting point for most teams. If you are on the Vercel AI SDK, support just landed via PR 21910 https://github.com/vercel/ai/pull/21910 : experimental streamTranscribe { model: 'microsoft/mai-transcribe-2-streaming' } The model is also accessible through OpenRouter https://openrouter.ai/microsoft/mai-transcribe-2 and the Vercel AI Gateway, which lowers the barrier for JS/TS teams not already on Azure. Where It Fits in the Latency Budget Voice agents live and die by round-trip latency. Twilio’s engineering team recommends budgeting roughly 350 ms for STT, 375 ms for LLM first token, and 100 ms for TTS first byte — a total target around 1.1 seconds mouth-to-ear. Beyond 800 ms, users start noticing the gap. At 130 ms to final transcript, MAI-Transcribe-2-Streaming comfortably fits the STT budget. What streaming STT gives you beyond accuracy is the ability to overlap stages: the LLM can begin inference while the speaker is still talking. That time saving is real and measurable — typically 200–400 ms off total round-trip time. The Honest Comparison | Model | WER | Streaming | Latency final | Price | SLA | |---|---|---|---|---|---| | MAI-Transcribe-2-Streaming | 2.5% | Yes | 0.13s | $9/1K min | No preview | | Deepgram Nova-3 | 5–7% | Yes | ~0.3–0.45s | ~$3–4/1K min | Yes | | Whisper Large v3 | ~12% multilingual | No | Batch only | $0.36/1K min | Yes | MAI-Transcribe-2-Streaming wins on accuracy. Deepgram wins on price, latency slightly , and production readiness. If you are launching a voice product in the next six months and need an SLA, Deepgram is still the pragmatic choice. If you are building on Azure and willing to stay in preview, MAI-Transcribe-2-Streaming gives you the best transcript quality available right now. The Caveats You Need to Know Microsoft lists all three new models as public preview — no SLA, not recommended for production. The $0.54/hour introductory price $9 per 1,000 minutes is 5.4× more expensive than the non-streaming batch version and roughly double what Deepgram charges. LiveKit integration — critical for the open-source voice agent ecosystem — is listed as “coming soon.” And while the STT model handles 60 languages, the TTS counterpart MAI-Voice-2.1 covers 23 — a mismatch that complicates fully multilingual deployments. None of these are deal-breakers for experimentation. They are all deal-breakers for production commitment right now. Bottom Line Microsoft needed a streaming STT model to complete its voice agent stack, and it shipped one with genuinely best-in-class accuracy. The documentation is live https://learn.microsoft.com/en-us/azure/ai-services/speech-service/mai-transcribe-2-streaming , the Vercel AI SDK integration is merged, and the MAI Playground lets you hear the full pipeline working in under five minutes. Build a proof of concept now. Do not pull Deepgram out of a production deployment until Microsoft ships a GA release with an SLA attached. For what it is — the most accurate streaming STT on the market, now fully integrated into the Azure Foundry voice pipeline — MAI-Transcribe-2-Streaming is worth the preview friction if you are building voice agents and accuracy is your primary concern.