cd /news/artificial-intelligence/mai-transcribe-2-streaming-real-time… · home › topics › artificial-intelligence › article
[ARTICLE · art-144230] src=byteiota.com ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

MAI-Transcribe-2-Streaming: Real-Time Voice Agents, Completed

Microsoft launched MAI-Transcribe-2-Streaming on October 1, a WebSocket-based streaming speech-to-text model that ranked first for both Final Transcript and First Partial Transcript accuracy across 38 models on Artificial Analysis's AA-WER Streaming benchmark with a 2.5% word error rate. The model ships alongside the new MAI-Voice-2.1-Flash text-to-speech model and the generally available Azure Voice Live orchestration layer, letting Azure developers run the full real-time voice agent loop on first-party infrastructure instead of routing audio through Deepgram or AssemblyAI. MAI-Transcribe-2-Streaming delivers final transcripts 0.13 seconds after end of speech and begins partial hypotheses around 100 ms after audio starts, versus the 5-7% production word error rate posted by Deepgram Nova-3.

read5 min views3 publishedOct 3, 2026
MAI-Transcribe-2-Streaming: Real-Time Voice Agents, Completed
Image: Byteiota (auto-discovered)

Microsoft launched three voice models on October 1 that do something its Azure speech stack could not do before: run the entire real-time voice agent loop — hear, transcribe, respond, speak — entirely on first-party infrastructure. The new piece is MAI-Transcribe-2-Streaming, a WebSocket-based streaming speech-to-text model that ranked first for accuracy across 38 models in independent testing. Paired with the newly released MAI-Voice-2.1-Flash (TTS) and the generally available Azure Voice Live orchestration layer, Azure developers now have a complete voice agent pipeline without routing audio through Deepgram, AssemblyAI, or any other third-party vendor.

The Gap That Existed Until Yesterday #

Building a real-time voice agent has always required assembling three components: STT, LLM, and TTS. The LLM part has been well-served for a while. The TTS part got serviceable. The STT piece was the persistent gap for Azure-native developers: OpenAI Whisper does not stream (it is a batch endpoint with a 25 MB file cap), and Azure’s own previous neural speech models did not offer true streaming transcription with partial hypotheses. The standard workaround was to add Deepgram or AssemblyAI to the stack — which works fine but introduces a third-party dependency, a separate pricing agreement, and an additional latency hop.

MAI-Transcribe-2-Streaming is the answer to that problem. It streams. It is first-party. And its accuracy numbers are not typical launch-day marketing figures.

The Accuracy Numbers Are Independently Verified #

Artificial Analysis ran the model against 38 competitors on its AA-WER Streaming benchmark and put MAI-Transcribe-2-Streaming at number one for both Final Transcript accuracy and First Partial Transcript accuracy, with a 2.5% word error rate. The final transcript is delivered 0.13 seconds after end of speech. Partial hypotheses start arriving around 100 ms after audio begins — meaning an LLM sitting downstream can start reasoning before the speaker has even finished their sentence.

For context: Deepgram Nova-3, currently the most widely deployed streaming STT in voice agents, posts 5–7% WER in production. That is a meaningful gap. Mishearing “cancel” as “confirm,” or “staging” as “staying,” at 6% WER happens regularly enough to matter in production workflows. The 2.5% figure is not a rounding improvement — it is a substantively different accuracy tier.

How to Integrate #

There are two integration paths. The first is a direct WebSocket connection using Microsoft’s Realtime API:

wss://{resource}.services.ai.azure.com/mai/v1/realtime?intent=transcription
Header: api-key: {your-key}

Push 100–200 ms audio packets into the socket, and the model returns two event types: partial (updated hypothesis as speech continues) and final (committed segment after a ). The second path is the Azure Speech SDK, which wraps the WebSocket with retry logic, connection management, and audio buffering — the better starting point for most teams.

If you are on the Vercel AI SDK, support just landed via PR #21910:

experimental_streamTranscribe({
  model: 'microsoft/mai-transcribe-2-streaming'
})

The model is also accessible through OpenRouter and the Vercel AI Gateway, which lowers the barrier for JS/TS teams not already on Azure.

Where It Fits in the Latency Budget #

Voice agents live and die by round-trip latency. Twilio’s engineering team recommends budgeting roughly 350 ms for STT, 375 ms for LLM first token, and 100 ms for TTS first byte — a total target around 1.1 seconds mouth-to-ear. Beyond 800 ms, users start noticing the gap.

At 130 ms to final transcript, MAI-Transcribe-2-Streaming comfortably fits the STT budget. What streaming STT gives you beyond accuracy is the ability to overlap stages: the LLM can begin inference while the speaker is still talking. That time saving is real and measurable — typically 200–400 ms off total round-trip time.

The Honest Comparison #

Model WER Streaming Latency (final) Price SLA
MAI-Transcribe-2-Streaming 2.5% Yes 0.13s $9/1K min No (preview)
Deepgram Nova-3 5–7% Yes ~0.3–0.45s ~$3–4/1K min Yes
Whisper Large v3 ~12% multilingual No Batch only $0.36/1K min Yes

MAI-Transcribe-2-Streaming wins on accuracy. Deepgram wins on price, latency (slightly), and production readiness. If you are launching a voice product in the next six months and need an SLA, Deepgram is still the pragmatic choice. If you are building on Azure and willing to stay in preview, MAI-Transcribe-2-Streaming gives you the best transcript quality available right now.

The Caveats You Need to Know #

Microsoft lists all three new models as public preview — no SLA, not recommended for production. The $0.54/hour introductory price ($9 per 1,000 minutes) is 5.4× more expensive than the non-streaming batch version and roughly double what Deepgram charges. LiveKit integration — critical for the open-source voice agent ecosystem — is listed as “coming soon.” And while the STT model handles 60 languages, the TTS counterpart (MAI-Voice-2.1) covers 23 — a mismatch that complicates fully multilingual deployments.

None of these are deal-breakers for experimentation. They are all deal-breakers for production commitment right now.

Bottom Line #

Microsoft needed a streaming STT model to complete its voice agent stack, and it shipped one with genuinely best-in-class accuracy. The documentation is live, the Vercel AI SDK integration is merged, and the MAI Playground lets you hear the full pipeline working in under five minutes. Build a proof of concept now. Do not pull Deepgram out of a production deployment until Microsoft ships a GA release with an SLA attached.

For what it is — the most accurate streaming STT on the market, now fully integrated into the Azure Foundry voice pipeline — MAI-Transcribe-2-Streaming is worth the preview friction if you are building voice agents and accuracy is your primary concern.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @microsoft 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/mai-transcribe-2-str…] indexed:0 read:5min 2026-10-03 · —