CalQuity covers earnings calls in real time, so we routinely test whether transcription models can handle Indian names, financial vocabulary, mixed-language speech and imperfect conference-call audio. This is our field evaluation of Sarvam Saaras v3.
What we tested
We ran Saaras v3 across four Indian earnings calls: IRCTC, Anand Rathi Wealth, Narayana Hrudayalaya and Angel One. The recordings included clean management introductions, Indian proper nouns, Hindi-to-English translation, code-switching, noisy analyst lines, static, muted microphones and overlapping speech.
This was a practical evaluation rather than a formal word-error-rate benchmark. We reviewed whether the output was useful for real-time research, then returned to the supplied audio clips whenever a name, number, speaker turn or unusual phrase looked questionable.
Where Sarvam impressed us
Indian context feels natural
The transcripts retained “sir”, “ji”, crore and lakh amounts, Indian institutions, cities and company-specific vocabulary. That context matters when the output is meant for investors.
Names are often caught correctly
Sarvam picked up management names and proper nouns such as Sanjay Kumar Jain, Feroz Aziz, Jugal Mantri, Agra, Mathura and Prayagraj. Proper nouns still need review, but the baseline is encouraging.
Long-form structure is usable
Time markers and speaker turns make a long call easy to navigate, search and convert into first-pass notes.
Financial speech is mostly intelligible
Terms such as AUM, EBITDA, PAT, GST, CSR, ECL and margins appeared in the transcripts. Individual numbers remain high-risk, but the model is operating in the right domain.
Mixed-language speech is a real strength
Hindi-heavy sections can become readable English without losing the broad structure of the conversation, which is valuable for downstream note-taking.
This opening of the Narayana Hrudayalaya call introduces the management team and company-specific names.
IRCTC Q4 FY26
The closing remarks switch heavily into Hindi and test whether the output remains useful to an English-language research workflow.
Where the model needs review
A conference call contains more than speech: line hiss, static, clipped microphones, room echo, muted participants and overlapping voices. Saaras does not always mark that uncertainty. Its most important failure mode is producing a plausible sentence from audio that is incomplete or not speech at all.
Speaker IDs are not identities
“SPEAKER_13” separates turns but does not identify a person. Crosstalk can also split one person into multiple IDs or merge two voices.
Numbers need special attention
Rupee amounts, crores, percentages and fiscal years are business-critical. Every number used in research should be checked against the recording or source filing.
Time markers are approximate
Timestamps help locate a passage, but they do not establish the precise moment each individual word was spoken.
Silence and overlap need review
A gap may be silence; a fragment may be cross-talk. Both should trigger review rather than confident reconstruction.
This IRCTC section contains a question about the internet-ticketing EBIT margin on a badly degraded line. Parts of the question repeat or smear together, so the precise percentages are not reliable from the recording alone.
Confirmed failure cases
The following examples have enough audible evidence to identify a specific failure. Each card separates the supported ground truth from Sarvam’s output and names the failure mode.
IRCTC Q4 FY26
Background chatter and cross-talk enter the line before the internet-ticketing margin question.
IRCTC Q4 FY26
The clip moves from the end of a passenger-traffic question through “Hello” and into management’s reply.
“SPEAKER_8: Perfect”
IRCTC Q4 FY26
Madhu Chanda Dey is finishing her question about the April–June quarter before management replies.
IRCTC Q4 FY26
A short side comment leaks through when management’s microphone is unmuted after a question about CSR.
IRCTC Q4 FY26
Sonal Minhas introduces herself and checks whether she is audible on a noisy line.
Unresolved stress cases
Two clips surfaced suspicious proper nouns, but the source audio is not clear enough to establish a defensible verbatim ground truth. They remain useful review triggers, not confirmed errors.
IRCTC Q4 FY26
“SPEAKER_9: Okay sir, Gopal.”
IRCTC Q4 FY26
“SPEAKER_10: Ah, is it Titan now?”
“SPEAKER_0: Is it better now?”
The conclusion: Sarvam matters
Sarvam is one of India’s most consequential home-grown AI labs, and Saaras shows why. It is not at the global state of the art yet: the model still needs stronger uncertainty handling when audio becomes noisy or incomplete. But it excels where Indian context matters—names, code-switching, business vocabulary and the natural rhythm of local conversations.
For now, Saaras is best used as a fast first-draft engine, not a source of record. Use it to create the searchable transcript, then verify proper nouns, financial figures, guidance and every passage affected by static or overlapping speech. Better uncertainty handling could move it from a useful assistant to a dependable production layer.
That gap should not obscure the larger point. The brief withdrawal of Anthropic’s Fable 5 from international availability showed how quickly access to a frontier model can change because of decisions made outside India. India needs strong AI systems of its own—not as a patriotic substitute for performance, but as durable infrastructure. Sarvam is not the finished answer, but it is one of the most credible attempts to build it.
