Microsoft has added MAI-Transcribe-2 to the Microsoft Foundry model catalog in public preview. Announced as part of Build 2026, the speech-to-text model expands Microsoft's first-party MAI lineup across transcription, reasoning, image generation, and voice. For developers and teams evaluating transcription services, the immediate significance is access to an early version of a new Microsoft model through Foundry, rather than a confirmed new benchmark for price, speed, or accuracy.
Microsoft lists MAI-Transcribe-2 among four MAI models entering public preview: MAI-Thinking-1, MAI-Image-2.5, MAI-Voice-2, and the new transcription model. The company's Build 2026 Microsoft Foundry update directs users to try the new models in the Foundry catalog.
That public-preview designation matters. It makes the model available for early testing, but it is not the same as a detailed production specification, general-availability commitment, or published service-level promise. Businesses considering it for recorded meetings, interviews, calls, video captions, or searchable audio archives should treat the release as an opportunity to evaluate fit against their own audio and workflow requirements.
The confirmed announcement positions MAI-Transcribe-2 as the transcription component of the MAI model family and makes it available through Microsoft Foundry's catalog in public preview. Foundry is Microsoft's platform for accessing and working with AI models, and the Build release broadens its first-party catalog beyond a single modality.
The announcement also places MAI-Transcribe-2 within a larger Foundry update that includes Fireworks AI availability, Managed Compute, fine-tuning, and governance and observability enhancements. Those surrounding platform changes show Microsoft's broader effort to make Foundry a place where developers can evaluate and use models from different providers and across different tasks.
What the supplied Build materials do not establish is equally important. They do not publish a fixed price for MAI-Transcribe-2, a speed benchmark, a quality benchmark, supported languages, audio limits, or a direct comparison with OpenAI's GPT-Transcribe. Claims that the model is the cheapest, fastest, highest quality, or 10 times faster than GPT-Transcribe are not substantiated by the cited Microsoft announcement.
| New MAI model in Foundry public preview | Role confirmed in the supplied research | What the Build recap confirms |
|---|---|---|
| MAI-Transcribe-2 | Speech-to-text transcription | Available to try in the Foundry catalog |
| MAI-Thinking-1 | Listed as a new MAI model | Entering public preview |
| MAI-Image-2.5 | Listed as a new MAI model | Entering public preview |
| MAI-Voice-2 | Listed as a new MAI model | Entering public preview |
A useful evaluation should begin with the real work the transcript must support. Speech-to-text quality is not only about converting spoken words into text. Teams may need accurate speaker names, terminology, punctuation, timestamps, captions, searchable records, or text suitable for follow-up summaries and actions. The Build announcement confirms catalog availability, but it does not define MAI-Transcribe-2 performance on those requirements.
A practical trial can therefore compare a representative set of recordings across the tools already in use. Use clean recordings as well as the audio that creates operational problems, such as calls with multiple speakers, industry-specific vocabulary, varying accents, or background noise. Review the output manually before treating any result as a basis for a customer-facing caption, a compliance-sensitive record, or an automated downstream action.
For video workflows, the key question is not simply whether a model can create text. It is whether the output can reliably fit the existing process: obtaining audio, submitting it for transcription, storing the result, reviewing exceptions, and passing approved text into captions, notes, search, or other applications. The supplied materials do not describe a specific MAI-Transcribe-2 video integration path, so teams should verify the relevant Foundry implementation and commercial details before changing a live workflow. Transcription costs can include more than model usage. A workflow may also involve audio preparation, storage, human review, caption formatting, and integrations with the systems where recordings and transcripts are used. Because Microsoft has not published MAI-Transcribe-2 pricing in the supplied announcement, a cost comparison with GPT-Transcribe or other transcription tools cannot yet be made from this release alone.
The same caution applies to speed. A model's total turnaround time may depend on the recording length, system design, queueing, file handling, and review process, in addition to model inference. The announcement's public-preview status makes testing more valuable than relying on unsupported comparative marketing claims.
For teams that already use Microsoft services, Foundry catalog access may simplify an initial evaluation by putting MAI-Transcribe-2 alongside other models in the same platform. That is a potential workflow advantage, not proof that the model will outperform alternatives. The appropriate decision is evidence-based: test quality on relevant recordings, confirm current commercial terms, and measure end-to-end processing in the intended workflow. As transcription becomes part of operational automation, the value often comes from what happens after the text is produced. A reviewed transcript can support internal search, meeting follow-ups, content preparation, or structured records. But those outcomes depend on the reliability of the transcription and the design of the surrounding process, neither of which is fully specified in the Build announcement.
Businesses exploring AI transcription should avoid treating a model catalog addition as a plug-and-play workflow upgrade. Scalevise's AI workflow automation service can help map recordings to the systems where transcripts create value, build review steps that reduce manual rework, and test automation before it affects customer or operational processes. This creates a clearer basis for comparing tools and deciding whether a new model belongs in the workflow. Discuss an AI automation project with Scalevise.
What is MAI-Transcribe-2? MAI-Transcribe-2 is a Microsoft MAI family speech-to-text model. Microsoft announced it as a new addition to the Microsoft Foundry catalog at Build 2026.
Has Microsoft confirmed MAI-Transcribe-2 pricing?
The supplied Build announcement does not provide a fixed price for MAI-Transcribe-2. Businesses should confirm current pricing and applicable terms in Foundry before estimating costs.
Is MAI-Transcribe-2 10 times faster than GPT-Transcribe?
The supplied Microsoft primary source does not confirm a 10 times speed comparison with GPT-Transcribe. It also does not publish a direct speed, price, or quality comparison between the two models.
MAI-Transcribe-2 gives Microsoft Foundry users an early opportunity to test a new first-party transcription model alongside other newly previewed MAI models. The confirmed news is its public-preview availability, not a verified claim of market-leading price, speed, or accuracy. Teams should use relevant recordings and workflow measures to determine whether it is a practical addition to their transcription process.