Sarvam AI 发布 Sarvam Audio 语音识别模型
Sarvam Audio: Speech Recognition beyond Transcription
Sarvam AI 发布 Sarvam Audio,这是 Sarvam 3B 的音频扩展,覆盖 22 种印度语言和印度英语,支持五种转写格式控制、最长 60 分钟音频的说话人分离,以及基于对话上下文的语音识别。
Sarvam Audio 把语音识别从转写扩展到格式控制、说话人分离和上下文理解,并给出与 GPT-4o-Transcribe、Gemini-3-Flash 的对比结果。
Research
February 2, 2026 · 6 min read

India is a voice-first country. The rigid structure of a keyboard struggles to capture the fluidity of Indian languages, and for most people, speaking is simply more natural than typing. From farmers checking crop prices, to gig workers receiving navigation instructions, to elderly users navigating WhatsApp and smart TVs, voice is the default mode of interaction.
This reality presents both opportunity and challenge. Traditional automatic speech recognition (ASR) systems perform well on clean, read-speech benchmarks, but they often break down in real-world Indian settings. In practice, accuracy alone is an insufficient lens. Speech recognition in India must go beyond transcription.
Three core challenges stand out:
First, script control Indian speakers freely mix English into their speech. In some applications, English words must be preserved in Roman script while in others they must be transliterated into the native script. A single fixed output format does not work.
Second, multi-speaker separation Real-world audio often involves multiple speakers talking simultaneously. Accurate recognition requires not just transcription, but reliable speaker identification and attribution.
Third, contextual awareness Speech recognition systems must improve using context derived from prior turns in a conversation or from accumulated context in long-form audio. Without this, short utterances, ambiguous phrases, and noisy segments are routinely misinterpreted.
Sarvam Audio addresses these challenges.
Sarvam Audio is an audio extension of Sarvam 3B, a 3-billion-parameter language model pre-trained from scratch on English and 22 Indian languages. Sarvam Audio moves beyond traditional ASR systems by producing multiple, user-controlled transcription formats and modeling speech as a contextual signal to support robust recognition across long durations, conversational settings, multi-speaker audio, and action-oriented voice interactions.
Finely Controlled Transcription Format
Real-world speech applications require control not just over what is transcribed, but how it is rendered. Indian speech is inherently multilingual and code-mixed, and different downstream use cases demand different transcription formats.
Sarvam Audio allows applications to explicitly specify the desired transcription style at inference time. Sarvam Audio supports five transcription modes:
Audio LM Voice Features
Evaluation
To quantify the impact of transcription format control, Sarvam Audio is evaluated against GPT-4o-Transcribe and Gemini-3-Flash using the Word Error Rate (WER) metric across three transcription styles unnormalised, normalised, and code-mixed.
Across all transcription styles, Sarvam Audio consistently outperforms baseline models, demonstrating that format control need not come at the cost of accuracy.
State-of-the-Art Diarized Speech Recognition
Real-world audio is rarely single-speaker. Meetings, interviews, and conversations often involve multiple speakers, overlapping speech, and rapid turn-taking. Accurate speech recognition in these settings requires both high-quality transcription and reliable speaker attribution.
Sarvam Audio achieves state-of-the-art performance on diarized speech recognition for audio up to60 minutes, accurately transcribing speech while identifying who spoke what.
Evaluation
- Word Diarization Error Rate (WDER) (lower is better) Measures the percentage of aligned words (including matches and substitutions) that are attributed to the wrong speaker.
- Diarization Error Rate (DER) (lower is better) Measures the sum of speaker confusion, false alarms, and missed speech, normalized by total speech time.
Evaluation setup
- In-house benchmark built from real-world meeting recordings
- Expert human annotations
- Audio length: 1-60 minutes
- Up to 8 speakers
- Significant overlapping speech
Contextual Speech Recognition
Context is essential for decoding real-world audio. LLM backbone allows Sarvam Audio to leverage context given via textual description or conversational history, to significantly improve transcription quality in tricky scenarios.
- Resolving linguistic ambiguity in short utterances When a user responds with “नौ (Nau)” to a quantity prompt, Sarvam Audio uses conversational context to correctly interpret it as the Hindi number nine, rather than the English word no.
- Recovering meaning from noisy audio In degraded acoustic environments, if a user says “Bhaiya, loc son bhejo”, Sarvam Audio uses delivery-domain context to reconstruct the intended phrase “Bhaiya, location bhejo”.
- Domain-aware transcription In a stock-market discussion, Sarvam Audio correctly transcribes “M&M” as Mahindra & Mahindra, rather than the literal phrase “M and M”.
Examples
Evaluation
On a benchmark that mirrors real-world conversational speech across Indian languages,Sarvam Audio outperforms Gemini-3-Flash.
Rather than relying solely on WER or CER because they capture only word-level similarity, an LLM-based intent and entity preservation score is used that better reflects real conversational and command-based use cases.
Evaluation measures:
- Intent preservation-Whether the core action and intent of the utterance are correctly understood.
- Entity preservation-Whether critical entities such as names, numbers, locations, and organizations are retained.
Sarvam Audio demonstrates a consistent performance advantage across languages. The evaluation framework is open-sourced here, and the Synthetic Contextual ASR Benchmark (Indic) is publicly released on
Hugging Face.
Benchmark overview.This benchmark evaluates context-aware ASR for voice-bot interactions across the top 10 Indian languages. Each sample represents a single dialog turn and includes audio, ground-truth transcription, language, and rich conversational context (bot persona, history, and prompt). The dataset is synthetically generated across domains such as Banking, E-commerce, and Healthcare, and is designed to test contextual biasing, intent preservation, and end-to-end spoken language understanding.
Speech to Command
Voice agents are now everywhere. Most existing systems rely on a two-stage architecture: audio is first transcribed by an ASR model and then interpreted by a separate LLM. While effective, this approach introduces additional latency and often breaks contextual continuity particularly for short or noisy utterances.
Sarvam Audio demonstrates that high-precision function calling and parameter extraction can be performed directly on the audio modality.
By operating end-to-end on speech:
- Intent and context are preserved
- Latency is significantly reduced
- System complexity is simplified
Production-grade understanding can be achieved with small, domain-specific fine-tuning datasets, enabling fast and reliable deployment of voice agents without large-model overhead.
In the example above, Sarvam Audio infers the appropriate function name and its arguments directly from the speech input, conditioned on prior conversational context, enabling precise function invocation.
Conclusion
Sarvam Audio rethinks speech recognition for India from the ground up. It delivers state-of-the-art ASR across 22 Indian languages and Indian English, while addressing the realities of code-mixing, script variation, long-form audio, overlapping speakers, and conversational context.
More importantly, it goes beyond traditional ASR. With built-in context awareness, diarization, format control, and direct speech-to-command capabilities, Sarvam Audio forms the foundation for a new generation of voice-first applications and agents built for real Indian users.
Sarvam Audio will be available soon on the Sarvam Dashboard.
Voice is the interface. Sarvam Audio makes it work for India.
来源:Sarvam AI · sarvam.ai