{"slug": "microsoft-voice-and-transcribe-models-hit-vercel-ai-gateway", "title": "Microsoft Voice and Transcribe Models Hit Vercel AI Gateway", "summary": "Microsoft's MAI speech models — microsoft/mai-voice-2.1, its lower-latency sibling microsoft/mai-voice-2.1-flash, and the streaming transcription endpoint microsoft/mai-transcribe-2-streaming — are now available through Vercel's AI Gateway, the company said. The integration keeps zero data retention so audio inputs are not held for training, adds no platform markup to base inference costs, and is callable from AI SDK 7 via the generateSpeech and streamTranscribe functions. The standard voice model supports expressive speech in 23 languages with consistent speaker identity for long-form content, while the flash variant targets millisecond-latency voice agents.", "body_md": "# Microsoft Voice and Transcribe Models Hit Vercel AI Gateway\n\nThe integration of Microsoft’s MAI models into Vercel’s AI Gateway removes the need to juggle multiple provider dashboards for speech tasks. This partnership brings `microsoft/mai-voice-2.1`, its lower-latency sibling `microsoft/mai-voice-2.1-flash`, and the streaming transcription endpoint `microsoft/mai-transcribe-2-streaming` directly into the Gateway ecosystem. The setup retains the zero data retention (ZDR) guarantee, meaning your audio inputs aren’t held for training, while billing remains transparent with no platform markup added to the base inference costs.\n\n## Choosing the Right Voice Model\n\nSelecting between the standard and flash variants depends entirely on latency requirements. The full `microsoft/mai-voice-2.1` model supports expressive speech in 23 languages and maintains consistent speaker identity across long-form content. This makes it the correct choice for audiobooks, podcast intros, or educational narration where tonal consistency matters more than speed.\n\nConversely, `microsoft/mai-voice-2.1-flash` targets interactive applications. If you are building a voice agent that needs to reply within milliseconds, the flash variant cuts the delay. Both models handle multilingual generation, but the trade-off is clear: use the standard model for quality and longevity, and the flash model for real-time responsiveness.\n\n## Implementing Speech Generation\n\nTo generate speech using the AI SDK 7, you call the `generateSpeech` function. The following example demonstrates configuring the Harper voice with the Flash model for a quick response:\n\n``` js\nimport { generateSpeech } from \"ai\";\nconst { audio } = await generateSpeech({\nmodel: \"microsoft/mai-voice-2.1-flash\",\nvoice: \"Harper\",\nprompt: \"Welcome to the new voice interface.\"\n});\n```\n\nFor longer passages, swap the model identifier to `microsoft/mai-voice-2.1`. The SDK handles the underlying API routing through the Gateway, allowing you to switch providers later without refactoring your client code.\n\n## Handling Live Transcription\n\nThe streaming transcription endpoint handles partial updates as audio data arrives. This is critical for live captioning or real-time transcription apps. The `streamTranscribe` function accepts a stream of audio chunks and emits transcript updates incrementally.\n\nEnsure your input audio is formatted correctly. The example below assumes a `ReadableStream` of 16 kHz, 16-bit PCM audio data:\n\n``` js\nimport { streamTranscribe } from \"ai\";\nconst result = streamTranscribe({\nmodel: \"microsoft/mai-transcribe-2-streaming\",\naudio: microphoneStream, // 16kHz, 16-bit PCM\n});\n// Replace displayed text as new partials arrive\nfor await (const chunk of result) {\nconsole.log(chunk.text);\n}\n```\n\nNote that partial transcripts are tentative. As more audio data arrives, previous segments may refine or change. Your UI should overwrite existing text rather than appending blindly to avoid visual glitches.\n\n## Why the Gateway Layer Matters\n\nUsing AI Gateway adds a uniform API surface for tracking usage, costs, and request traces. Beyond simple model invocation, it supports routing, retries, and failover mechanisms. If one provider experiences downtime, the Gateway can automatically switch traffic, provided you have configured alternative backends. For MAI models specifically, this means you get centralized logging alongside the ZDR compliance and direct billing parity.\n\nCheck the MAI model page for the full family of supported endpoints. The Gateway also includes a speech quickstart guide to help you initialize the environment variables and set up the initial connection without manual header configuration.\n\n[Next Continual learning could render blocking monitors nearly ineffective →](https://promptcube3.com/en/threads/9697/)\n\n## All Replies （1）\n\nWant a live back-and-forth? [Join the global AI chat room](https://promptcube3.com/en/chat/) — login to talk.\n\nThe zero data retention is the real sell for me, since I batch transcribe customer calls and can't risk audio leaking into training sets.", "url": "https://wpnews.pro/news/microsoft-voice-and-transcribe-models-hit-vercel-ai-gateway", "canonical_source": "https://promptcube3.com/en/threads/9719/", "published_at": "2026-10-01 18:03:27+00:00", "updated_at": "2026-10-01 18:18:22.684150+00:00", "lang": "en", "topics": ["ai-products", "ai-tools", "developer-tools", "natural-language-processing", "ai-infrastructure"], "entities": ["Microsoft", "Vercel", "microsoft/mai-voice-2.1", "microsoft/mai-voice-2.1-flash", "microsoft/mai-transcribe-2-streaming", "Vercel AI Gateway", "AI SDK 7"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/microsoft-voice-and-transcribe-models-hit-vercel-ai-gateway", "markdown": "https://wpnews.pro/news/microsoft-voice-and-transcribe-models-hit-vercel-ai-gateway.md", "text": "https://wpnews.pro/news/microsoft-voice-and-transcribe-models-hit-vercel-ai-gateway.txt", "jsonld": "https://wpnews.pro/news/microsoft-voice-and-transcribe-models-hit-vercel-ai-gateway.jsonld"}}