# MAI-Transcribe-2-Streaming: Real-Time Voice Agents, Completed

> Source: <https://byteiota.com/mai-transcribe-2-streaming-real-time-voice-agents-completed/>
> Published: 2026-10-03 01:09:00+00:00

Microsoft launched three voice models on October 1 that do something its Azure speech stack could not do before: run the entire real-time voice agent loop — hear, transcribe, respond, speak — entirely on first-party infrastructure. The new piece is **MAI-Transcribe-2-Streaming**, a WebSocket-based streaming speech-to-text model that ranked first for accuracy across 38 models in independent testing. Paired with the newly released **MAI-Voice-2.1-Flash** (TTS) and the generally available **Azure Voice Live** orchestration layer, Azure developers now have a complete voice agent pipeline without routing audio through Deepgram, AssemblyAI, or any other third-party vendor.

## The Gap That Existed Until Yesterday

Building a real-time voice agent has always required assembling three components: STT, LLM, and TTS. The LLM part has been well-served for a while. The TTS part got serviceable. The STT piece was the persistent gap for Azure-native developers: OpenAI Whisper does not stream (it is a batch endpoint with a 25 MB file cap), and Azure’s own previous neural speech models did not offer true streaming transcription with partial hypotheses. The standard workaround was to add Deepgram or AssemblyAI to the stack — which works fine but introduces a third-party dependency, a separate pricing agreement, and an additional latency hop.

MAI-Transcribe-2-Streaming is the answer to that problem. It streams. It is first-party. And its accuracy numbers are not typical launch-day marketing figures.

## The Accuracy Numbers Are Independently Verified

Artificial Analysis ran the model against 38 competitors on its AA-WER Streaming benchmark and put MAI-Transcribe-2-Streaming at **number one for both Final Transcript accuracy and First Partial Transcript accuracy**, with a 2.5% word error rate. The final transcript is delivered 0.13 seconds after end of speech. Partial hypotheses start arriving around 100 ms after audio begins — meaning an LLM sitting downstream can start reasoning before the speaker has even finished their sentence.

For context: [Deepgram Nova-3](https://deepgram.com/learn/voice-agent-architecture-stt-llm-tts-pipeline-design), currently the most widely deployed streaming STT in voice agents, posts 5–7% WER in production. That is a meaningful gap. Mishearing “cancel” as “confirm,” or “staging” as “staying,” at 6% WER happens regularly enough to matter in production workflows. The 2.5% figure is not a rounding improvement — it is a substantively different accuracy tier.

## How to Integrate

There are two integration paths. The first is a direct WebSocket connection using Microsoft’s Realtime API:

```
wss://{resource}.services.ai.azure.com/mai/v1/realtime?intent=transcription
Header: api-key: {your-key}
```

Push 100–200 ms audio packets into the socket, and the model returns two event types: `partial` (updated hypothesis as speech continues) and `final` (committed segment after a pause). The second path is the [Azure Speech SDK](https://learn.microsoft.com/en-us/azure/ai-services/speech-service/mai-transcribe-2-streaming), which wraps the WebSocket with retry logic, connection management, and audio buffering — the better starting point for most teams.

If you are on the Vercel AI SDK, support just landed via [PR #21910](https://github.com/vercel/ai/pull/21910):

```
experimental_streamTranscribe({
  model: 'microsoft/mai-transcribe-2-streaming'
})
```

The model is also accessible through [OpenRouter](https://openrouter.ai/microsoft/mai-transcribe-2) and the Vercel AI Gateway, which lowers the barrier for JS/TS teams not already on Azure.

## Where It Fits in the Latency Budget

Voice agents live and die by round-trip latency. Twilio’s engineering team recommends budgeting roughly 350 ms for STT, 375 ms for LLM first token, and 100 ms for TTS first byte — a total target around 1.1 seconds mouth-to-ear. Beyond 800 ms, users start noticing the gap.

At 130 ms to final transcript, MAI-Transcribe-2-Streaming comfortably fits the STT budget. What streaming STT gives you beyond accuracy is the ability to overlap stages: the LLM can begin inference while the speaker is still talking. That time saving is real and measurable — typically 200–400 ms off total round-trip time.

## The Honest Comparison

| Model | WER | Streaming | Latency (final) | Price | SLA | 
|---|---|---|---|---|---|
| MAI-Transcribe-2-Streaming | 2.5% | Yes | 0.13s | $9/1K min | No (preview) | 
| Deepgram Nova-3 | 5–7% | Yes | ~0.3–0.45s | ~$3–4/1K min | Yes | 
| Whisper Large v3 | ~12% multilingual | No | Batch only | $0.36/1K min | Yes | 

MAI-Transcribe-2-Streaming wins on accuracy. Deepgram wins on price, latency (slightly), and production readiness. If you are launching a voice product in the next six months and need an SLA, Deepgram is still the pragmatic choice. If you are building on Azure and willing to stay in preview, MAI-Transcribe-2-Streaming gives you the best transcript quality available right now.

## The Caveats You Need to Know

Microsoft lists all three new models as **public preview** — no SLA, not recommended for production. The $0.54/hour introductory price ($9 per 1,000 minutes) is 5.4× more expensive than the non-streaming batch version and roughly double what Deepgram charges. LiveKit integration — critical for the open-source voice agent ecosystem — is listed as “coming soon.” And while the STT model handles 60 languages, the TTS counterpart (MAI-Voice-2.1) covers 23 — a mismatch that complicates fully multilingual deployments.

None of these are deal-breakers for experimentation. They are all deal-breakers for production commitment right now.

## Bottom Line

Microsoft needed a streaming STT model to complete its voice agent stack, and it shipped one with genuinely best-in-class accuracy. The [documentation is live](https://learn.microsoft.com/en-us/azure/ai-services/speech-service/mai-transcribe-2-streaming), the Vercel AI SDK integration is merged, and the MAI Playground lets you hear the full pipeline working in under five minutes. Build a proof of concept now. Do not pull Deepgram out of a production deployment until Microsoft ships a GA release with an SLA attached.

For what it is — the most accurate streaming STT on the market, now fully integrated into the Azure Foundry voice pipeline — MAI-Transcribe-2-Streaming is worth the preview friction if you are building voice agents and accuracy is your primary concern.
