cd /news/artificial-intelligence/buzzasr-a-swarm-of-100-monolingual-s… · home topics artificial-intelligence article
[ARTICLE · art-125378] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

BuzzASR: A Swarm of 100+ Monolingual Speech Recognition Models

Researchers released BuzzASR, a collection of 102 language-specialized fine-tuned Whisper models for automatic speech recognition, according to an arXiv paper (2609.09554v1). BuzzASR models outperform Whisper-large-v3 on 77 of 102 languages, reducing character error rates by a factor of over 2.8 on average, and achieve state-of-the-art CER among open-source systems on 27 of 102 languages on the combined FLEURS and Common Voice test set. The team's monolingual tokenizer replacement strategy yields an average 3.3x improvement in compression rate over Whisper's multilingual BPE, with gains up to 21.7x, and all models, code, and detailed results are released at lemn-lab.github.io/buzz-asr.

by read1 min views1 publishedSep 10, 2026

arXiv:2609.09554v1 Announce Type: new Abstract: We introduce BuzzASR, a collection of language-specialized fine-tuned Whisper models adapted for automatic speech recognition (ASR) in 102 languages. Large end-to-end Transformer-based ASR models such as Whisper have revolutionized ASR, but most prominent models are highly multilingual. As a result, these models often perform poorly on languages less well-represented in their training set. While it has long been known that effective language adaptation can be achieved through simple fine-tuning on monolingual data, this strategy has only been applied to a small number of languages. We massively scale up this simple approach to 102 languages covered in the FLEURS dataset, while also implementing a more complex language adaptation strategy that integrates monolingual tokenizer replacement and data augmentation using text-only fine-tuning. BuzzASR models outperform Whisper-large-v3 on 77 out of 102 languages, reducing character error rates (CER) by a factor of over 2.8 on average. Our models achieve state-of-the-art CER among open-source systems on 27 of 102 languages on the combined FLEURS and Common Voice test set. Our tokenizer replacement strategy yields an average 3.3x improvement in compression rate (characters per token) over Whisper's multilingual BPE, with gains of up to 21.7x. We release all models, code, and detailed results: https://lemn-lab.github.io/buzz-asr

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @buzzasr 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/buzzasr-a-swarm-of-1…] indexed:0 read:1min 2026-09-10 ·