Microsoft Voice and Transcribe Models Hit Vercel AI Gateway Microsoft's MAI speech models — microsoft/mai-voice-2.1, its lower-latency sibling microsoft/mai-voice-2.1-flash, and the streaming transcription endpoint microsoft/mai-transcribe-2-streaming — are now available through Vercel's AI Gateway, the company said. The integration keeps zero data retention so audio inputs are not held for training, adds no platform markup to base inference costs, and is callable from AI SDK 7 via the generateSpeech and streamTranscribe functions. The standard voice model supports expressive speech in 23 languages with consistent speaker identity for long-form content, while the flash variant targets millisecond-latency voice agents. Microsoft Voice and Transcribe Models Hit Vercel AI Gateway The integration of Microsoft’s MAI models into Vercel’s AI Gateway removes the need to juggle multiple provider dashboards for speech tasks. This partnership brings microsoft/mai-voice-2.1 , its lower-latency sibling microsoft/mai-voice-2.1-flash , and the streaming transcription endpoint microsoft/mai-transcribe-2-streaming directly into the Gateway ecosystem. The setup retains the zero data retention ZDR guarantee, meaning your audio inputs aren’t held for training, while billing remains transparent with no platform markup added to the base inference costs. Choosing the Right Voice Model Selecting between the standard and flash variants depends entirely on latency requirements. The full microsoft/mai-voice-2.1 model supports expressive speech in 23 languages and maintains consistent speaker identity across long-form content. This makes it the correct choice for audiobooks, podcast intros, or educational narration where tonal consistency matters more than speed. Conversely, microsoft/mai-voice-2.1-flash targets interactive applications. If you are building a voice agent that needs to reply within milliseconds, the flash variant cuts the delay. Both models handle multilingual generation, but the trade-off is clear: use the standard model for quality and longevity, and the flash model for real-time responsiveness. Implementing Speech Generation To generate speech using the AI SDK 7, you call the generateSpeech function. The following example demonstrates configuring the Harper voice with the Flash model for a quick response: js import { generateSpeech } from "ai"; const { audio } = await generateSpeech { model: "microsoft/mai-voice-2.1-flash", voice: "Harper", prompt: "Welcome to the new voice interface." } ; For longer passages, swap the model identifier to microsoft/mai-voice-2.1 . The SDK handles the underlying API routing through the Gateway, allowing you to switch providers later without refactoring your client code. Handling Live Transcription The streaming transcription endpoint handles partial updates as audio data arrives. This is critical for live captioning or real-time transcription apps. The streamTranscribe function accepts a stream of audio chunks and emits transcript updates incrementally. Ensure your input audio is formatted correctly. The example below assumes a ReadableStream of 16 kHz, 16-bit PCM audio data: js import { streamTranscribe } from "ai"; const result = streamTranscribe { model: "microsoft/mai-transcribe-2-streaming", audio: microphoneStream, // 16kHz, 16-bit PCM } ; // Replace displayed text as new partials arrive for await const chunk of result { console.log chunk.text ; } Note that partial transcripts are tentative. As more audio data arrives, previous segments may refine or change. Your UI should overwrite existing text rather than appending blindly to avoid visual glitches. Why the Gateway Layer Matters Using AI Gateway adds a uniform API surface for tracking usage, costs, and request traces. Beyond simple model invocation, it supports routing, retries, and failover mechanisms. If one provider experiences downtime, the Gateway can automatically switch traffic, provided you have configured alternative backends. For MAI models specifically, this means you get centralized logging alongside the ZDR compliance and direct billing parity. Check the MAI model page for the full family of supported endpoints. The Gateway also includes a speech quickstart guide to help you initialize the environment variables and set up the initial connection without manual header configuration. Next Continual learning could render blocking monitors nearly ineffective → https://promptcube3.com/en/threads/9697/ All Replies (1) Want a live back-and-forth? Join the global AI chat room https://promptcube3.com/en/chat/ — login to talk. The zero data retention is the real sell for me, since I batch transcribe customer calls and can't risk audio leaking into training sets.