# Microsoft previews MAI-Image-2.5-Pro and MAI-Voice-2-Flash

> Source: <https://www.testingcatalog.com/microsoft-previews-mai-image-2-5-pro-and-mai-voice-2-flash/>
> Published: 2026-07-26 07:29:03+00:00

Microsoft has opened public previews of MAI-Image-2.5-Pro and MAI-Voice-2-Flash, expanding its in-house AI model family with one option tuned for maximum visual fidelity and another designed for high-volume speech. The models are available through Microsoft Foundry, with trials also offered in the MAI Playground, and are aimed at creative teams, developers and enterprise contact centers.

MAI-Image-2.5-Pro is Microsoft’s highest-fidelity image model to date. It is built for hero artwork, detailed editing and precise text inside generated images, with natural-language commands supporting rapid revisions. Pricing is $5 per 1 million text input tokens, $8 per 1 million image input tokens and $106 per 1 million image output tokens.

MAI-Voice-2-Flash focuses on speed, scale and lower operating costs. Microsoft says it is twice as fast as MAI-Voice-2 and 32% cheaper while retaining natural prosody and high acoustic quality. It costs $15 per 1 million characters and is intended for responsive voice agents and large call-center workloads.

The previews arrive as Microsoft’s broader MAI stack moves deeper into its products. Bing Image Creator now uses MAI-Image-2.5 end-to-end by default, while PowerPoint uses the model for image-to-image work and OneDrive relies on it for key editing tasks. Microsoft reports up to 84% lower GPU costs in PowerPoint compared with GPT-Image-2, plus a 26% rise in save rates and about 25% lower P95 latency in OneDrive. MAI-Voice-2-Flash now powers Microsoft’s enterprise contact-center platform and is integrated into Azure Voice Live, with Microsoft citing GPU cost reductions of up to 89% in the contact-center deployment.

The launch follows Microsoft’s year-long push to build purpose-built models trained on clean, traceable enterprise-grade data without distillation from third-party systems. Its strategy is to match each product with a different point on the quality, speed and cost curve while controlling the models that serve millions of users. WPP Global Chief Creative Officer Rob Reilly described the image model’s text rendering as a breakthrough and said its natural-language editing makes creative iteration faster and more intuitive.
