David AI's 21-language benchmark finds English hides speech-recognition gaps David AI's DAI-ASR-I18N benchmark, built by CEO Tomer Cohen and CTO Ben Wiley, compared 14 speech-recognition systems on unscripted conversation across 21 languages and found error rates varied by an average of 3.7 percentage points across five languages with the most consistent results, versus 25.4 points across Bengali, Hindi, Marathi, Tamil and Telugu. Microsoft AI's MAI-Transcribe-2 posted the lowest transcription error rate in 18 of 21 languages, while ElevenLabs' Scribe v2 led in the other three, including English. The San Francisco startup's public evaluation set omits French, Italian and Korean pending privacy review, and the source corpus contains 147 hours of unscripted conversations recorded on separate synchronized channels and transcribed verbatim. David AI's 21-language benchmark finds English hides speech-recognition gaps Tomer Cohen and Ben Wiley tested 14 systems on natural conversation; the public files omit French, Italian and Korean pending privacy review. By Ryan Merket https://runtimewire.com/author/ryan-merket ยท Published Primary source: David AI https://x.com/withdavidai/status/2105490026108965329 Why it matters A multilingual benchmark built around unscripted speech challenges teams to choose recognition and diarization systems using the language and recording conditions they actually serve, rather than English scores alone. Tomer Cohen https://www.ycombinator.com/companies/david-ai?ref=runtimewire and Ben Wiley https://www.ycombinator.com/companies/david-ai?ref=runtimewire built David AI https://withdavid.ai/?ref=runtimewire around a practical constraint: voice models need recordings that resemble real conversation, with speakers and their words kept distinct. The San Francisco startup's new DAI-ASR-I18N benchmark https://research.withdavid.ai/blog/asr-i18n?ref=runtimewire puts that premise to a public test, comparing 14 speech systems on conversation across 21 languages. The clearest result is a warning against using English as a proxy for the rest of the world. On David AI's public evaluation set https://research.withdavid.ai/blog/asr-i18n?ref=runtimewire leaderboard , systems' error rates varied by an average of 3.7 percentage points across five languages where results were most consistent. Across Bengali, Hindi, Marathi, Tamil and Telugu, the average spread was 25.4 points. The benchmark's authors say Microsoft AI's MAI-Transcribe-2 had the lowest transcription error rate in 18 of 21 languages; ElevenLabs' Scribe v2 led in the other three, including English. The research report is framed around a real weakness in how speech systems are judged: an English score can make competing models look closer than they are in many other languages. That is useful evidence for teams choosing a system for a specific language, though the results remain tied to this dataset and its recording conditions. A benchmark built around actual conversation Cohen, David AI's CEO, previously served as chief of staff at Scale AI and worked at McKinsey, according to Y Combinator's company profile https://www.ycombinator.com/companies/david-ai?ref=runtimewire . Wiley, the company's CTO, led engineering for Scale AI's public-sector generative AI platform and earlier worked as a software engineer at Microsoft. Their experience at Scale put them close to the data work behind AI systems; David AI's benchmark extends that work from supplying data to measuring what models do with it. An account of David AI's origin by YC partner Diana Hu describes Cohen and Wiley asking founders during the Summer 2024 batch about difficult multimodal problems. A robotics company pointed to a shortage of high-quality voice data, and the founders built a phone-call recording app over a weekend. Hu described their thesis as a need for a "Common Crawl for audio" - a broad supply of usable recordings for speech models. The account is Hu's description of the company's early path https://www.linkedin.com/posts/sdianahu david-ai-powering-the-voice-era-of-ai-activity-7333565898642923520-a1At?ref=runtimewire , rather than a direct founder interview. DAI-ASR-I18N follows that thesis into evaluation. David AI says its source corpus contains 147 hours of unscripted conversations. Native speakers were paired for open-ended sessions, recorded on separate synchronized channels, and transcribed verbatim. Each transcript was reviewed by two annotators and a language expert. Fillers, repetitions, false starts, interruptions and backchannels remain in the material instead of being cleaned out of the reference transcript. That design helps test the conditions that polished read-aloud scripts miss. It also gives the benchmark an important boundary: the recordings feature two speakers on separate microphones. David AI notes that this setup can provide diarization models with cues unavailable in single-microphone or multi-speaker recordings. Only 15 recordings fall into the benchmark's highest-overlap tier, so its hardest-overlap results represent a narrow slice of the data. Strong rankings, with limits The benchmark also separates transcription from speaker diarization https://runtimewire.com/models/huggingface/pyannote-speaker-diarization-f455c3ef4f01af08 - identifying who spoke when. David AI says every dedicated diarization model it tested outperformed every combined transcription-and-diarization system. NVIDIA Nemotron 3 Diarization https://runtimewire.com/models/huggingface/nvidia-nemotron-3-diarization-a18533bcdf59005a posted the lowest overall diarization error rate in the public results, at 18.8%. Every system made its most diarization errors in conversations with more overlapping speech. The evaluation harness and scoring code are available on GitHub https://github.com/withdavid-ai/dai-asr-i18n?ref=runtimewire under the MIT License. The dataset on Hugging Face https://huggingface.co/datasets/davidai-labs/DAI-ASR-I18N?ref=runtimewire is more restricted: it contains identifiable voices, and access requires accepting a data-use agreement for non-commercial research and evaluation. The dataset card lists 1,035 clips and 42.1 hours of public two-channel audio from 682 speakers. It also says French, Italian and Korean are excluded from the downloadable release pending an internal privacy review, despite appearing in the 21-language benchmark coverage. The release snapshot is dated September 26th, and David AI's announcement thread https://x.com/withdavidai/status/2105490026108965329?ref=runtimewire appeared on October 1st. The leaderboard covers 21 languages, but the downloadable release omits French, Italian and Korean pending privacy review. A private test set is also held back to reduce the chance that models have already trained on its recordings. For Cohen and Wiley, making the scoring process reproducible is a way to turn David AI's data expertise into a reference point for model builders, not just a supplier relationship. Their benchmark gives customers and researchers a public method to compare systems on speech that sounds less like a prepared demo. Its strongest immediate finding is also its most practical one: English-only testing can conceal large language-specific performance gaps.