cd /news/artificial-intelligence/meet-mai-transcribe-2-a-faster-and-m… · home topics artificial-intelligence article
[ARTICLE · art-120570] src=microsoft.ai ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

Meet MAI-Transcribe-2: A faster and more accurate speech recognition model

MAI-Transcribe-2, a new speech recognition model from MAI, claims to be the fastest, most accurate, and cheapest in the world, ranking first on the FLEURS benchmark across 60 languages with an average Word-Error-Rate of 5.2%. The model is up to 10x faster than OpenAI's GPT-Transcribe, 7x faster than ElevenLabs' Scribe v2, and 5x faster than Gemini 3.5 Transcribe, and is priced at $0.10 per hour as a limited-time offer until the end of the year.

read3 min views9 publishedSep 3, 2026
Meet MAI-Transcribe-2: A faster and more accurate speech recognition model
Image: Microsoft (auto-discovered)

#

MAI-Transcribe-2 is the fastest, most accurate and cheapest speech recognition model in the world

Introducing MAI‑Transcribe‑2. It’s not only our most capable transcription model yet, but the most capable and efficient amongst our competitors.

With new features like diarization, configurable transcription styles, and word-level timestamps, MAI-Transcribe-2 beats other leading models like Gemini 3.5 Transcribe, GPT-Transcribe, Whisper V3-Large, and ScribeV2, while also handling a broader range of real‑world audio.

Our model ranks first on the FLEURS benchmark across 60 languages with an average Word-Error-Rate of 5.2%, defines the Pareto Frontier for accuracy and latency on Artificial Analysis, and ranks second on the Artificial Analysis Word-Error-Rate leaderboard, continuing the hill-climbing from previous versions.

Solve more challenges with a single model

From clinical note-taking to legal documentation, and from accessibility to closed captioning, MAI-Transcribe-2 is designed to take on real-world applications, with: Faster inference with substantially lower latency, especially for long‑form audio, with up to 10× faster processing than leading competitors.** Speaker diarizationdistinguishes between speakers and attributes words to the right person within a recording Word‑level timestampsprovide precise timing for every word, enabling more accurate alignment, search, navigation, and editing Keyword biasinghelps the model recognize domain-specific terminology, abbreviations, names, and other terms that can be difficult to distinguish from context aloneConfigurable transcription styles** give developers control over the output. The “verbatim” setting captures speech as spoken, including filler words and false starts, for compliance and analysis workloads. The “clean” setting removes fillers to produce more readable captions, notes, and published transcriptsCode switching supports conversations that naturally move between languages, including commonly blended language pairs such as Hinglish and SpanglishAutomatic language identification accurately detects the specific language being spoken without users needing to specify in advanceRobust performance in noisy conditions helps maintain transcription quality beyond controlled recording environmentsAccurate across 60 languages to provide quality transcription for developers around the world

Efficiency without sacrificing accuracy

MAI‑Transcribe‑2 is incredibly efficient, leading the Artificial Analysis accuracy-latency Pareto frontier, combining leading transcription quality with market‑leading batch speed. It’s significantly faster than the latest models from major competitors, delivering a clear advantage when low‑latency transcription is critical.

Based on evals run by Artificial Analysis, the model is 10x faster than OpenAI’s GPT‑Transcribe, 7x faster than ElevenLabs’ Scribe v2, and 5x faster than Gemini 3.5 Transcribe while delivering higher accuracy.

Consistent quality across 60 languages

MAI‑Transcribe‑2 is accurate across more languages than any other model.

Our evaluations on the public, multilingual benchmark FLEURS show that it maintains a consistently high accuracy bar across all tested. Developers needing to transcribe across multiple languages can choose one single model, reducing complexity and even potentially saving GPU utilization issues.

High performance at the best price

Highly efficient means highly cost-effective. MAI-Transcribe-2’s speed and throughput allow us to offer the most competitive price in the market, helping developers process more audio without compromising transcription quality. At launch, MAI-Transcribe-2 will be priced at $0.10 per hour as a limited-time offer until the end of the year.

Try MAI-Transcribe-2 today MAI-Transcribe-2’s powerful capabilities are available to demo today, through Microsoft Foundry,__ MAI Playground__ and

.

Open Router

#

Build the Future With Us

We’re a lean, fast-moving lab made up of some of the world’s most talented minds. We have an exciting roadmap of compute at MAI, which is ramping quickly and extensively. And we have an ambitious mission we truly believe in. We’re also fortunate to partner with incredible product teams giving our models the chance to reach billions of users and create immense positive impact. If you’re a brilliant, highly-ambitious and low ego individual, you’ll fit right in—come and join us as we work on our next generation of models!

Explore all jobs

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @mai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/meet-mai-transcribe-…] indexed:0 read:3min 2026-09-03 ·