{"slug": "microsoft-ai-releases-new-transcription-and-text-to-speech-models-for-voice", "title": "Microsoft AI releases new transcription and text-to-speech models for voice agents", "summary": "Microsoft AI released MAI-Transcribe-2-Streaming, a real-time transcription model that Microsoft says ranks first for accuracy on Artificial Analysis, transcribes 60 languages, and returns first partial results in just over 100 milliseconds at an introductory price of $0.54 per hour of audio through the end of the year. Microsoft also released two text-to-speech models: MAI-Voice-2.1, which speaks 23 languages in the same voice with a native accent in each, and the MAI-Voice-2.1-Flash variant, which Microsoft says hits 150 milliseconds of latency and costs $15 per million characters instead of $22. Both voice models clone a voice from a few seconds of reference audio, include built-in safeguards, and are available through Microsoft Foundry, the MAI Playground, and OpenRouter; in one test, about half of 4,000 participants thought the voices belonged to a real person.", "body_md": "# Microsoft AI releases new transcription and text-to-speech models for voice agents\n\n**Microsoft AI has released MAI-Transcribe-2-Streaming, a new model for real-time transcription.** Microsoft says it ranks first for accuracy on Artificial Analysis. The model transcribes 60 languages and delivers its first partial results in just over 100 milliseconds. Microsoft says this lets voice agents respond while someone is still mid-sentence. Through the end of the year, an hour of audio costs $0.54 at the introductory price.\n\nMicrosoft also released two new text-to-speech models. [MAI-Voice-2.1](https://learn.microsoft.com/en-us/azure/ai-services/speech-service/mai-voices?context=%2Fazure%2Ffoundry%2Fcontext%2Fcontext&tabs=mai-voice-2-1-flash&pivots=ai-foundry) is designed to speak 23 languages in the same voice, with a native accent in each one. According to Microsoft, the MAI-Voice-2.1-Flash variant hits a latency of 150 milliseconds and costs $15 per million characters instead of $22.\n\nBoth voice models can clone a voice from just a few seconds of reference audio, and built-in safeguards are meant to prevent misuse. The models are available through Microsoft Foundry and the [MAI Playground](http://playground.microsoft.ai/), among other platforms, and the two voice models are also on [OpenRouter](https://aka.ms/mai-openrouter). In one test, about half of the 4,000 participants thought the voices belonged to a real person.\n\n```\nAI News Without the Hype – Curated by Humans\n\n\t\t\t\t\tSubscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive \"AI Radar\" frontier report six times a year, full archive access, and access to our comment section.\t\t\t\t\n\n\t\t\t\t\tSubscribe now\n```\n\n[Microsoft](https://microsoft.ai/news/our-first-streaming-transcription-model/)", "url": "https://wpnews.pro/news/microsoft-ai-releases-new-transcription-and-text-to-speech-models-for-voice", "canonical_source": "https://the-decoder.com/microsoft-ai-releases-new-transcription-and-text-to-speech-models-for-voice-agents/", "published_at": "2026-10-02 09:20:39+00:00", "updated_at": "2026-10-02 09:35:59.547482+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-products", "natural-language-processing", "generative-ai"], "entities": ["Microsoft", "Microsoft AI", "MAI-Transcribe-2-Streaming", "MAI-Voice-2.1", "MAI-Voice-2.1-Flash", "Artificial Analysis", "Microsoft Foundry", "OpenRouter"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/microsoft-ai-releases-new-transcription-and-text-to-speech-models-for-voice", "markdown": "https://wpnews.pro/news/microsoft-ai-releases-new-transcription-and-text-to-speech-models-for-voice.md", "text": "https://wpnews.pro/news/microsoft-ai-releases-new-transcription-and-text-to-speech-models-for-voice.txt", "jsonld": "https://wpnews.pro/news/microsoft-ai-releases-new-transcription-and-text-to-speech-models-for-voice.jsonld"}}