{"slug": "david-ai-s-21-language-benchmark-finds-english-hides-speech-recognition-gaps", "title": "David AI's 21-language benchmark finds English hides speech-recognition gaps", "summary": "David AI's DAI-ASR-I18N benchmark, built by CEO Tomer Cohen and CTO Ben Wiley, compared 14 speech-recognition systems on unscripted conversation across 21 languages and found error rates varied by an average of 3.7 percentage points across five languages with the most consistent results, versus 25.4 points across Bengali, Hindi, Marathi, Tamil and Telugu. Microsoft AI's MAI-Transcribe-2 posted the lowest transcription error rate in 18 of 21 languages, while ElevenLabs' Scribe v2 led in the other three, including English. The San Francisco startup's public evaluation set omits French, Italian and Korean pending privacy review, and the source corpus contains 147 hours of unscripted conversations recorded on separate synchronized channels and transcribed verbatim.", "body_md": "# David AI's 21-language benchmark finds English hides speech-recognition gaps\n\n**Tomer Cohen and Ben Wiley tested 14 systems on natural conversation; the public files omit French, Italian and Korean pending privacy review.**\n\n        By [Ryan Merket](https://runtimewire.com/author/ryan-merket)\n        · Published \n\nPrimary source: [David AI](https://x.com/withdavidai/status/2105490026108965329)\n\n## Why it matters\n\nA multilingual benchmark built around unscripted speech challenges teams to choose recognition and diarization systems using the language and recording conditions they actually serve, rather than English scores alone.\n\n[Tomer Cohen](https://www.ycombinator.com/companies/david-ai?ref=runtimewire) and [Ben Wiley](https://www.ycombinator.com/companies/david-ai?ref=runtimewire) built [David AI](https://withdavid.ai/?ref=runtimewire) around a practical constraint: voice models need recordings that resemble real conversation, with speakers and their words kept distinct. The San Francisco startup's new [DAI-ASR-I18N benchmark](https://research.withdavid.ai/blog/asr-i18n?ref=runtimewire) puts that premise to a public test, comparing 14 speech systems on conversation across 21 languages.\n\nThe clearest result is a warning against using English as a proxy for the rest of the world. On David AI's [public evaluation set](https://research.withdavid.ai/blog/asr-i18n?ref=runtimewire#leaderboard), systems' error rates varied by an average of 3.7 percentage points across five languages where results were most consistent. Across Bengali, Hindi, Marathi, Tamil and Telugu, the average spread was 25.4 points. The benchmark's authors say Microsoft AI's MAI-Transcribe-2 had the lowest transcription error rate in 18 of 21 languages; ElevenLabs' Scribe v2 led in the other three, including English.\n\nThe research report is framed around a real weakness in how speech systems are judged: an English score can make competing models look closer than they are in many other languages. That is useful evidence for teams choosing a system for a specific language, though the results remain tied to this dataset and its recording conditions.\n\n### A benchmark built around actual conversation\n\nCohen, David AI's CEO, previously served as chief of staff at Scale AI and worked at McKinsey, according to [Y Combinator's company profile](https://www.ycombinator.com/companies/david-ai?ref=runtimewire). Wiley, the company's CTO, led engineering for Scale AI's public-sector generative AI platform and earlier worked as a software engineer at Microsoft. Their experience at Scale put them close to the data work behind AI systems; David AI's benchmark extends that work from supplying data to measuring what models do with it.\n\nAn account of David AI's origin by YC partner Diana Hu describes Cohen and Wiley asking founders during the Summer 2024 batch about difficult multimodal problems. A robotics company pointed to a shortage of high-quality voice data, and the founders built a phone-call recording app over a weekend. Hu described their thesis as a need for a \"Common Crawl for audio\" - a broad supply of usable recordings for speech models. The account is [Hu's description of the company's early path](https://www.linkedin.com/posts/sdianahu_david-ai-powering-the-voice-era-of-ai-activity-7333565898642923520-a1At?ref=runtimewire), rather than a direct founder interview.\n\nDAI-ASR-I18N follows that thesis into evaluation. David AI says its source corpus contains 147 hours of unscripted conversations. Native speakers were paired for open-ended sessions, recorded on separate synchronized channels, and transcribed verbatim. Each transcript was reviewed by two annotators and a language expert. Fillers, repetitions, false starts, interruptions and backchannels remain in the material instead of being cleaned out of the reference transcript.\n\nThat design helps test the conditions that polished read-aloud scripts miss. It also gives the benchmark an important boundary: the recordings feature two speakers on separate microphones. David AI notes that this setup can provide diarization models with cues unavailable in single-microphone or multi-speaker recordings. Only 15 recordings fall into the benchmark's highest-overlap tier, so its hardest-overlap results represent a narrow slice of the data.\n\n### Strong rankings, with limits\n\nThe benchmark also separates transcription from [speaker diarization](https://runtimewire.com/models/huggingface/pyannote-speaker-diarization-f455c3ef4f01af08) - identifying who spoke when. David AI says every dedicated diarization model it tested outperformed every combined transcription-and-diarization system. NVIDIA [Nemotron 3 Diarization](https://runtimewire.com/models/huggingface/nvidia-nemotron-3-diarization-a18533bcdf59005a) posted the lowest overall diarization error rate in the public results, at 18.8%. Every system made its most diarization errors in conversations with more overlapping speech.\n\nThe evaluation harness and scoring code are [available on GitHub](https://github.com/withdavid-ai/dai-asr-i18n?ref=runtimewire) under the MIT License. The [dataset on Hugging Face](https://huggingface.co/datasets/davidai-labs/DAI-ASR-I18N?ref=runtimewire) is more restricted: it contains identifiable voices, and access requires accepting a data-use agreement for non-commercial research and evaluation. The dataset card lists 1,035 clips and 42.1 hours of public two-channel audio from 682 speakers. It also says French, Italian and Korean are excluded from the downloadable release pending an internal privacy review, despite appearing in the 21-language benchmark coverage.\n\nThe release snapshot is dated September 26th, and David AI's [announcement thread](https://x.com/withdavidai/status/2105490026108965329?ref=runtimewire) appeared on October 1st. The leaderboard covers 21 languages, but the downloadable release omits French, Italian and Korean pending privacy review. A private test set is also held back to reduce the chance that models have already trained on its recordings.\n\nFor Cohen and Wiley, making the scoring process reproducible is a way to turn David AI's data expertise into a reference point for model builders, not just a supplier relationship. Their benchmark gives customers and researchers a public method to compare systems on speech that sounds less like a prepared demo. Its strongest immediate finding is also its most practical one: English-only testing can conceal large language-specific performance gaps.", "url": "https://wpnews.pro/news/david-ai-s-21-language-benchmark-finds-english-hides-speech-recognition-gaps", "canonical_source": "https://runtimewire.com/article/david-ai-multilingual-conversation-speech-benchmark", "published_at": "2026-10-01 06:46:59+00:00", "updated_at": "2026-10-01 07:16:55.275860+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "natural-language-processing", "ai-research"], "entities": ["David AI", "Tomer Cohen", "Ben Wiley", "DAI-ASR-I18N", "Microsoft AI", "MAI-Transcribe-2", "ElevenLabs", "Scribe v2"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/david-ai-s-21-language-benchmark-finds-english-hides-speech-recognition-gaps", "markdown": "https://wpnews.pro/news/david-ai-s-21-language-benchmark-finds-english-hides-speech-recognition-gaps.md", "text": "https://wpnews.pro/news/david-ai-s-21-language-benchmark-finds-english-hides-speech-recognition-gaps.txt", "jsonld": "https://wpnews.pro/news/david-ai-s-21-language-benchmark-finds-english-hides-speech-recognition-gaps.jsonld"}}