Bangalore-based Sarvam AI outperforms OpenAI and ElevenLabs on Indian language voice recognition Sarvam AI's Saaras V3 speech recognition model achieved a word error rate of approximately 19.3% on the IndicVoices top-10 language benchmarks, outperforming OpenAI's GPT-4o Transcribe, ElevenLabs' Scribe v2, Google's Gemini 3 Pro, and Deepgram's Nova-3. The Bengaluru-based startup trained Saaras V3 on over one million hours of multilingual Indian audio, and the model also leads the Svarah benchmark for Indian-accented English. Sarvam AI is one of twelve startups collaborating with the Indian government under the IndiaAI mission. Bangalore-based Sarvam AI outperforms OpenAI and ElevenLabs on Indian language voice recognition Sarvam AI's Saaras V3 model achieves lower error rates than GPT-4o Transcribe and ElevenLabs Scribe on Indian language benchmarks, trained on over one million hours of multilingual audio A Bengaluru startup just quietly outscored some of the biggest names in AI on a task that matters enormously in a country with 22 official languages and roughly 1.4 billion people: understanding what they’re actually saying. Sarvam AI’s latest speech recognition model, Saaras V3, posted a word error rate of approximately 19.3% on the IndicVoices top-10 language benchmarks, beating OpenAI’s GPT-4o Transcribe, ElevenLabs’ Scribe v2, Google’s Gemini 3 Pro, and Deepgram’s Nova-3. In speech recognition, a lower word error rate means fewer mistakes. Why Indian languages are an AI stress test India’s 22 scheduled languages span multiple script families, phonetic systems, and grammatical structures. Add in the widespread practice of code-mixing, where speakers blend Hindi and English mid-sentence, or Tamil and English, or any number of combinations, and you get audio data that would make most Western-trained models break out in digital hives. Sarvam trained Saaras V3 on over one million hours of multilingual Indian audio, with specific attention to noisy speech environments and code-mixed conversations. The performance gap between Saaras V3 and its global competitors widened further on lower-resource Indian languages, suggesting its architecture and data pipeline were purpose-built for this exact challenge rather than retrofitted from an English-first approach. Saaras V3 also earned a leading position on the Svarah benchmark, which specifically measures how well models handle Indian-accented English. More than just transcription Saaras V3 sits within Sarvam’s broader India-first technology stack, which includes a text-to-speech system called Bulbul, translation services, and a conversational agent platform named Sarvam Samvaad. The model supports native real-time streaming, speaker diarization the ability to distinguish between different speakers in a conversation , and automatic language detection. These capabilities are suited for call centers, media companies, and enterprise customer service operations that need to function at scale. Sarvam AI builds models covering all 22 scheduled Indian languages plus English. The IndiaAI mission and government backing Sarvam AI is one of twelve startups collaborating with the Indian government under the IndiaAI mission, a national initiative aimed at developing indigenous multilingual and multimodal AI technologies. Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy https://cryptobriefing.com/editorial-policy/ .