Microsoft launched three voice models on October 1 that do something its Azure speech stack could not do before: run the entire real-time voice agent loop — hear, transcribe, respond, speak — entirely on first-party infrastructure. The new piece is MAI-Transcribe-2-Streaming, a WebSocket-based streaming speech-to-text model that ranked first for accuracy across 38 models in independent testing. Paired with the newly released MAI-Voice-2.1-Flash (TTS) and the generally available Azure Voice Live orchestration layer, Azure developers now have a complete voice agent pipeline without routing audio through Deepgram, AssemblyAI, or any other third-party vendor.
The Gap That Existed Until Yesterday #
Building a real-time voice agent has always required assembling three components: STT, LLM, and TTS. The LLM part has been well-served for a while. The TTS part got serviceable. The STT piece was the persistent gap for Azure-native developers: OpenAI Whisper does not stream (it is a batch endpoint with a 25 MB file cap), and Azure’s own previous neural speech models did not offer true streaming transcription with partial hypotheses. The standard workaround was to add Deepgram or AssemblyAI to the stack — which works fine but introduces a third-party dependency, a separate pricing agreement, and an additional latency hop.
MAI-Transcribe-2-Streaming is the answer to that problem. It streams. It is first-party. And its accuracy numbers are not typical launch-day marketing figures.
The Accuracy Numbers Are Independently Verified #
Artificial Analysis ran the model against 38 competitors on its AA-WER Streaming benchmark and put MAI-Transcribe-2-Streaming at number one for both Final Transcript accuracy and First Partial Transcript accuracy, with a 2.5% word error rate. The final transcript is delivered 0.13 seconds after end of speech. Partial hypotheses start arriving around 100 ms after audio begins — meaning an LLM sitting downstream can start reasoning before the speaker has even finished their sentence.
For context: Deepgram Nova-3, currently the most widely deployed streaming STT in voice agents, posts 5–7% WER in production. That is a meaningful gap. Mishearing “cancel” as “confirm,” or “staging” as “staying,” at 6% WER happens regularly enough to matter in production workflows. The 2.5% figure is not a rounding improvement — it is a substantively different accuracy tier.
How to Integrate #
There are two integration paths. The first is a direct WebSocket connection using Microsoft’s Realtime API:
wss://{resource}.services.ai.azure.com/mai/v1/realtime?intent=transcription
Header: api-key: {your-key}
Push 100–200 ms audio packets into the socket, and the model returns two event types: partial (updated hypothesis as speech continues) and final (committed segment after a ). The second path is the Azure Speech SDK, which wraps the WebSocket with retry logic, connection management, and audio buffering — the better starting point for most teams.
If you are on the Vercel AI SDK, support just landed via PR #21910:
experimental_streamTranscribe({
model: 'microsoft/mai-transcribe-2-streaming'
})
The model is also accessible through OpenRouter and the Vercel AI Gateway, which lowers the barrier for JS/TS teams not already on Azure.
Where It Fits in the Latency Budget #
Voice agents live and die by round-trip latency. Twilio’s engineering team recommends budgeting roughly 350 ms for STT, 375 ms for LLM first token, and 100 ms for TTS first byte — a total target around 1.1 seconds mouth-to-ear. Beyond 800 ms, users start noticing the gap.
At 130 ms to final transcript, MAI-Transcribe-2-Streaming comfortably fits the STT budget. What streaming STT gives you beyond accuracy is the ability to overlap stages: the LLM can begin inference while the speaker is still talking. That time saving is real and measurable — typically 200–400 ms off total round-trip time.
The Honest Comparison #
| Model | WER | Streaming | Latency (final) | Price | SLA |
|---|---|---|---|---|---|
| MAI-Transcribe-2-Streaming | 2.5% | Yes | 0.13s | $9/1K min | No (preview) |
| Deepgram Nova-3 | 5–7% | Yes | ~0.3–0.45s | ~$3–4/1K min | Yes |
| Whisper Large v3 | ~12% multilingual | No | Batch only | $0.36/1K min | Yes |
MAI-Transcribe-2-Streaming wins on accuracy. Deepgram wins on price, latency (slightly), and production readiness. If you are launching a voice product in the next six months and need an SLA, Deepgram is still the pragmatic choice. If you are building on Azure and willing to stay in preview, MAI-Transcribe-2-Streaming gives you the best transcript quality available right now.
The Caveats You Need to Know #
Microsoft lists all three new models as public preview — no SLA, not recommended for production. The $0.54/hour introductory price ($9 per 1,000 minutes) is 5.4× more expensive than the non-streaming batch version and roughly double what Deepgram charges. LiveKit integration — critical for the open-source voice agent ecosystem — is listed as “coming soon.” And while the STT model handles 60 languages, the TTS counterpart (MAI-Voice-2.1) covers 23 — a mismatch that complicates fully multilingual deployments.
None of these are deal-breakers for experimentation. They are all deal-breakers for production commitment right now.
Bottom Line #
Microsoft needed a streaming STT model to complete its voice agent stack, and it shipped one with genuinely best-in-class accuracy. The documentation is live, the Vercel AI SDK integration is merged, and the MAI Playground lets you hear the full pipeline working in under five minutes. Build a proof of concept now. Do not pull Deepgram out of a production deployment until Microsoft ships a GA release with an SLA attached.
For what it is — the most accurate streaming STT on the market, now fully integrated into the Azure Foundry voice pipeline — MAI-Transcribe-2-Streaming is worth the preview friction if you are building voice agents and accuracy is your primary concern.