Bengaluru-based Sarvam AI says the latest version of its speech recognition model, Saaras V3, has posted lower word error rates than a set of widely used global speech-to-text systems, including Google’s Gemini 3 Pro and OpenAI’s GPT-4o Transcribe, on benchmarks designed around Indian languages and Indian-accented English.
The claim is based on benchmark charts shared publicly by Sarvam AI co-founder Pratyush Kumar on X, where the company compared Saaras V3 against Gemini 3 Pro, GPT-4o Transcribe, Deepgram Nova-3, and ElevenLabs Scribe v2 using the IndicVoices and Svarah datasets.
Sarvam says Saaras V3 led across the most widely used Indian languages in IndicVoices and also topped Svarah, a benchmark built around Indian-accented English speech from speakers across multiple states.
The key number Sarvam is highlighting
On a subset covering the 10 most popular languages in the IndicVoices dataset, Sarvam reports Saaras V3 delivered a word error rate of about 19.3%, lower than the competing systems it tested against. The company also says the gap becomes larger on the remaining IndicVoices languages, which include lower-resource Indian languages.
On Svarah, Sarvam says Saaras V3 again recorded the lowest word error rate among the compared systems.
What’s new in Saaras V3?
Sarvam is positioning Saaras V3 as a meaningful step beyond a routine model refresh. The company says the model is built on a new architecture and expands support to all 22 scheduled Indian languages, in addition to English.
One of the biggest changes: native real-time streaming speech recognition. In practical terms, the model can start producing text while the audio is still playing, instead of waiting for the entire clip to end . Sarvam says the streaming version is designed to keep accuracy close to batch processing while reducing latency, a combination aimed at scenarios like live captions, voice assistants, call-centre tools, and real-time transcription.
Sarvam’s technical blog (as referenced in the report) says Saaras V3 was trained on more than one million hours of multilingual audio, spanning Indian languages, accents and varied recording conditions, with a specific focus on code-mixed and noisy speech.
Training, the company says, included large-scale pre-training, supervised fine-tuning, reinforcement learning, and post-training steps intended to reduce “long-tail” errors and improve consistency across languages.
Beyond transcription: language detection, timestamps, and speaker separation
Sarvam is also pitching Saaras V3 as more than a plain speech-to-text engine. The company says the model supports automatic language detection, word-level timestamps, and speaker diarisation (separating and labelling different speakers in a conversation), features that matter for structured outputs like call analytics, meeting transcripts, media subtitling, and customer support workflows.
The company also says it offers multiple operating modes that trade off latency for accuracy, from a “fast” setting optimised for low time-to-first-token to more accuracy-focused options where transcript quality is the priority.
Speech recognition for Indian languages is notoriously unforgiving: multiple scripts, heavy code-mixing, regional accents, and noisy real-world audio. Sarvam’s pitch is that models trained primarily on Western and English-language data tend to stumble on these conditions and that targeted training for Indian formats and language patterns can produce measurable gains.
That argument mirrors Sarvam’s earlier benchmark claims in document AI. The company has previously said its document-focused model, Sarvam Vision, scored higher than several general-purpose systems on tasks such as document OCR, layout understanding, reading order detection, and table parsing, especially on multi-script Indian documents.
Sarvam says Saaras V3 extends the same thesis into speech recognition, especially for Indian languages, code-mixed inputs and Indian-accented English.
A quick snapshot of Sarvam AI’s broader stack
Sarvam AI describes itself as building speech, language and multimodal AI systems for Indian use cases, focusing on task-specific models rather than a single general-purpose chatbot.
Alongside Saaras, Sarvam’s portfolio includes Bulbul (text-to-speech for Indian languages), Saarika (speech-to-text transcription), Mayura (text translation), and Sarvam-M (a multilingual reasoning language model) . It also lists Samvaad, a voice-based conversational application built on top of its speech and language models.
Sarvam is also among the 12 startups working with the Indian government under the IndiaAI mission to develop indigenous multilingual and multimodal large language models.
Sarvam’s numbers and charts are company-reported and shared via a co-founder’s post on X. The benchmarks cited, IndicVoices and Svarah, are clearly relevant to Indian speech, but the claims in this update are best read as performance results disclosed by the company, not as independently audited rankings.
Also Read: Deedy Das Praises Sarvam’s Indic AI; Ashwini Vaishnav Calls It Sovereign AI Success
















