Microsoft MAI Voice Models Arrive in LiveKit for Expressive TTS Agents Microsoft AI voice models are now available in LiveKit Agents through an official Microsoft AI text-to-speech plugin, giving developers a documented way to use expressive MAI voices such as MAI-Voice-2-Flash in LiveKit-based voice applications. The plugin is installed via the livekit-agents[microsoft-ai]~=1.8 package extra and requires MICROSOFT_AI_TTS_API_KEY and MICROSOFT_AI_TTS_REGION environment variables, with LiveKit's documented example configuring the en-US-Harper voice at a 24 kHz sample rate. The integration expands the set of supported speech providers for real-time phone, web, and interactive assistants without requiring a custom connection to Microsoft's speech services. Microsoft AI voice models are now available in LiveKit Agents through an official Microsoft AI text-to-speech plugin. The integration gives developers a documented way to use expressive MAI voices, including MAI-Voice-2-Flash , in LiveKit-based voice applications https://scalevise.com/resources/ai-agents/ . For teams building real-time phone, web, or interactive assistants, the change expands the set of supported speech providers without requiring a separate custom connection to Microsoft’s speech services. The key development is a direct TTS integration, not a broad rollout of every MAI modality inside LiveKit. LiveKit’s official Microsoft AI TTS plugin documentation https://docs.livekit.io/agents/models/tts/microsoft-ai/ shows how its Agents framework can synthesize speech using MAI voices. Its example configures the en-US-Harper voice with the MAI-Voice-2-Flash model and a 24 kHz sample rate. That matters because voice agent quality is shaped by more than language-model responses. The text-to-speech layer determines how an agent delivers those responses to callers and users. A supported MAI option inside LiveKit gives developers a clearer integration path when expressive speech is a requirement for an agent experience. The Microsoft AI plugin makes MAI voices available as a TTS provider within LiveKit Agents. In practical terms, a developer can configure a LiveKit agent session to send generated text to Microsoft AI TTS and receive synthesized audio for the conversation. LiveKit’s documented example uses the following configuration elements: microsoft ai.TTS provider in a LiveKit agent session. MAI-Voice-2-Flash model. en-US-Harper voice identifier. The plugin is installed through the LiveKit Agents package extra: livekit-agents microsoft-ai ~=1.8 . Developers must also provide the MICROSOFT AI TTS API KEY and MICROSOFT AI TTS REGION environment variables. Those requirements make the Microsoft service credentials a core part of deployment, rather than an optional enhancement after an agent is built. Microsoft’s broader Foundry communications also list LiveKit among the platforms where MAI models, including voice and speech models https://scalevise.com/resources/microsoft-mai-transcribe-2-mai-voice-2-models/ , are available. The LiveKit-specific documentation is the more useful implementation reference because it describes the plugin, credentials, and code-level configuration needed to use MAI speech in an agent. | Component | Role in a LiveKit voice agent | Verified detail | |---|---|---| | LiveKit Agents | Agent framework and session configuration | Supports Microsoft AI as a TTS provider through its plugin ecosystem. | | Microsoft AI TTS plugin | Connection between LiveKit Agents and MAI voices | Installed with the microsoft-ai LiveKit Agents package extra. | | MAI-Voice-2-Flash | Text-to-speech model used in the documented example | Shown with the en-US-Harper voice at 24 kHz. | | Azure Speech resource | Authentication source | Requires a resource key and region, supplied through environment variables. | A supported provider plugin can reduce integration work for teams already using LiveKit. Instead of designing and maintaining their own adapter between an agent session and Microsoft’s TTS service, they can use LiveKit’s documented provider interface. That is particularly relevant when a team wants to test voice choices as part of an existing agent workflow. Potential applications include customer support bots, interactive assistants, and automated phone menus that need to speak their responses aloud. These are use cases for the integration, not evidence that a particular business outcome is guaranteed. Teams still need to design conversation flows, connect relevant business systems https://scalevise.com/services/api-system-integrations , test escalation paths, and evaluate audio quality in the environments where customers will actually use the service. The integration also makes provider selection a more concrete product decision. Businesses can compare how a voice fits their customer experience, while developers retain a common LiveKit agent architecture. The documented example demonstrates one MAI model, one voice, and one sample rate. It does not establish the full set of supported languages, voices, regions, or performance characteristics for every deployment. LiveKit states that its Inference system supports pay-as-you-go pricing across providers and global concurrency management. However, the supplied documentation does not provide a specific price for MAI voice usage through LiveKit, nor does it define concurrency limits, language coverage, regional availability, or latency targets for MAI deployments at scale. That means planning should begin with validation rather than assumptions. Before committing a customer-facing workflow, a team should confirm its Azure Speech resource configuration, assess the applicable provider and LiveKit pricing, and test the chosen voice under realistic call or session volumes. Real-time agents are sensitive to the combined performance of speech synthesis, agent logic, network conditions, and the rest of the application stack. For companies, the practical opportunity is straightforward: MAI speech can now be evaluated within LiveKit’s established agent workflow. The practical limitation is equally important: the available documentation confirms the integration path, but it does not answer every commercial or performance question that matters in a production rollout. If you are moving from a voice-agent prototype to a useful business workflow, the integration layer deserves as much attention as the model choice. Scalevise can help connect AI agents to the tools, data, and actions that make customer conversations productive through an MCP setup tailored to your systems https://scalevise.com/services/mcp-setup . This can reduce manual handoffs and turn an assistant into a controlled operational workflow. Discuss your AI integration project with Scalevise. What are Microsoft MAI voices in LiveKit? Microsoft MAI voices are available as a text-to-speech provider for LiveKit Agents through LiveKit’s official Microsoft AI plugin. The documented example uses MAI-Voice-2-Flash. What is required to use the Microsoft AI TTS plugin? Developers need the LiveKit Agents Microsoft AI package extra, an Azure Speech resource key, an Azure region, and the MICROSOFT AI TTS API KEY and MICROSOFT AI TTS REGION environment variables. Which MAI voice configuration does LiveKit document? LiveKit documents an example using the MAI-Voice-2-Flash model, the en-US-Harper voice, and a 24 kHz sample rate. Does the documentation specify MAI voice pricing or all supported languages? No. The supplied documentation describes pay-as-you-go pricing across providers through LiveKit Inference, but it does not provide MAI-specific pricing, full language coverage, regional availability, or latency details. Microsoft MAI voices are now a documented TTS option for LiveKit Agents, giving developers a direct route to use MAI-Voice-2-Flash and related Microsoft AI speech capabilities in voice agent workflows. The integration is meaningful because it simplifies the provider connection, but successful deployment still depends on testing credentials, costs, voice suitability, and real-time performance for the intended use case.