cd /news/artificial-intelligence/david-ai-s-21-language-benchmark-fin… · home › topics › artificial-intelligence › article
[ARTICLE · art-143057] src=runtimewire.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

David AI's 21-language benchmark finds English hides speech-recognition gaps

David AI's DAI-ASR-I18N benchmark, built by CEO Tomer Cohen and CTO Ben Wiley, compared 14 speech-recognition systems on unscripted conversation across 21 languages and found error rates varied by an average of 3.7 percentage points across five languages with the most consistent results, versus 25.4 points across Bengali, Hindi, Marathi, Tamil and Telugu. Microsoft AI's MAI-Transcribe-2 posted the lowest transcription error rate in 18 of 21 languages, while ElevenLabs' Scribe v2 led in the other three, including English. The San Francisco startup's public evaluation set omits French, Italian and Korean pending privacy review, and the source corpus contains 147 hours of unscripted conversations recorded on separate synchronized channels and transcribed verbatim.

by read4 min views3 publishedOct 1, 2026
David AI's 21-language benchmark finds English hides speech-recognition gaps
Image: Runtimewire (auto-discovered)

Tomer Cohen and Ben Wiley tested 14 systems on natural conversation; the public files omit French, Italian and Korean pending privacy review.

        By [Ryan Merket](https://runtimewire.com/author/ryan-merket)
        · Published 

Primary source: [David AI](https://x.com/withdavidai/status/2105490026108965329)

Why it matters #

A multilingual benchmark built around unscripted speech challenges teams to choose recognition and diarization systems using the language and recording conditions they actually serve, rather than English scores alone.

Tomer Cohen and Ben Wiley built David AI around a practical constraint: voice models need recordings that resemble real conversation, with speakers and their words kept distinct. The San Francisco startup's new DAI-ASR-I18N benchmark puts that premise to a public test, comparing 14 speech systems on conversation across 21 languages.

The clearest result is a warning against using English as a proxy for the rest of the world. On David AI's public evaluation set, systems' error rates varied by an average of 3.7 percentage points across five languages where results were most consistent. Across Bengali, Hindi, Marathi, Tamil and Telugu, the average spread was 25.4 points. The benchmark's authors say Microsoft AI's MAI-Transcribe-2 had the lowest transcription error rate in 18 of 21 languages; ElevenLabs' Scribe v2 led in the other three, including English.

The research report is framed around a real weakness in how speech systems are judged: an English score can make competing models look closer than they are in many other languages. That is useful evidence for teams choosing a system for a specific language, though the results remain tied to this dataset and its recording conditions.

A benchmark built around actual conversation

Cohen, David AI's CEO, previously served as chief of staff at Scale AI and worked at McKinsey, according to Y Combinator's company profile. Wiley, the company's CTO, led engineering for Scale AI's public-sector generative AI platform and earlier worked as a software engineer at Microsoft. Their experience at Scale put them close to the data work behind AI systems; David AI's benchmark extends that work from supplying data to measuring what models do with it.

An account of David AI's origin by YC partner Diana Hu describes Cohen and Wiley asking founders during the Summer 2024 batch about difficult multimodal problems. A robotics company pointed to a shortage of high-quality voice data, and the founders built a phone-call recording app over a weekend. Hu described their thesis as a need for a "Common Crawl for audio" - a broad supply of usable recordings for speech models. The account is Hu's description of the company's early path, rather than a direct founder interview.

DAI-ASR-I18N follows that thesis into evaluation. David AI says its source corpus contains 147 hours of unscripted conversations. Native speakers were paired for open-ended sessions, recorded on separate synchronized channels, and transcribed verbatim. Each transcript was reviewed by two annotators and a language expert. Fillers, repetitions, false starts, interruptions and backchannels remain in the material instead of being cleaned out of the reference transcript.

That design helps test the conditions that polished read-aloud scripts miss. It also gives the benchmark an important boundary: the recordings feature two speakers on separate microphones. David AI notes that this setup can provide diarization models with cues unavailable in single-microphone or multi-speaker recordings. Only 15 recordings fall into the benchmark's highest-overlap tier, so its hardest-overlap results represent a narrow slice of the data.

Strong rankings, with limits

The benchmark also separates transcription from speaker diarization - identifying who spoke when. David AI says every dedicated diarization model it tested outperformed every combined transcription-and-diarization system. NVIDIA Nemotron 3 Diarization posted the lowest overall diarization error rate in the public results, at 18.8%. Every system made its most diarization errors in conversations with more overlapping speech.

The evaluation harness and scoring code are available on GitHub under the MIT License. The dataset on Hugging Face is more restricted: it contains identifiable voices, and access requires accepting a data-use agreement for non-commercial research and evaluation. The dataset card lists 1,035 clips and 42.1 hours of public two-channel audio from 682 speakers. It also says French, Italian and Korean are excluded from the downloadable release pending an internal privacy review, despite appearing in the 21-language benchmark coverage.

The release snapshot is dated September 26th, and David AI's announcement thread appeared on October 1st. The leaderboard covers 21 languages, but the downloadable release omits French, Italian and Korean pending privacy review. A private test set is also held back to reduce the chance that models have already trained on its recordings.

For Cohen and Wiley, making the scoring process reproducible is a way to turn David AI's data expertise into a reference point for model builders, not just a supplier relationship. Their benchmark gives customers and researchers a public method to compare systems on speech that sounds less like a prepared demo. Its strongest immediate finding is also its most practical one: English-only testing can conceal large language-specific performance gaps.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @david ai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/david-ai-s-21-langua…] indexed:0 read:4min 2026-10-01 · —