CalQuity
Insight·14 min read·22 July 2026
FinanceAI

Sarvam Saaras v3 Review: Transcribing Indian Earnings Calls in Real Time

A thorough evaluation of Sarvam Saaras v3 across four Indian earnings calls, with audio examples covering names, mixed-language speech, static, cross-talk, and transcript failures.

INDIAN EARNINGS CALL · AUDIO → TRANSCRIPTAUDIOTRANSCRIPTNAMES · NUMBERS · NOISE
CalQuity field notes · Speech AI evaluation

CalQuity covers earnings calls in real time, so we routinely test whether transcription models can handle Indian names, financial vocabulary, mixed-language speech and imperfect conference-call audio. This is our field evaluation of Sarvam Saaras v3.

The short verdict Saaras v3 is a strong first-draft engine for Indian business audio. It understands local names, places, financial language and Hindi-heavy conversation unusually well. Its weakness is uncertainty: when static, muffling or cross-talk enters the recording, it can turn ambiguous sound into confident text. Use it to make long calls searchable quickly, but keep a human review layer between the transcript and any published research.

What we tested

We ran Saaras v3 across four Indian earnings calls: IRCTC, Anand Rathi Wealth, Narayana Hrudayalaya and Angel One. The recordings included clean management introductions, Indian proper nouns, Hindi-to-English translation, code-switching, noisy analyst lines, static, muted microphones and overlapping speech.

This was a practical evaluation rather than a formal word-error-rate benchmark. We reviewed whether the output was useful for real-time research, then returned to the supplied audio clips whenever a name, number, speaker turn or unusual phrase looked questionable.

How to read the examples: “Ground truth” means the words or acoustic condition we could support by listening to the attached clip. Where the 8 kHz recording was too degraded to be definitive, we mark the case as unresolved instead of forcing a conclusion.

Where Sarvam impressed us

Indian context feels natural

The transcripts retained “sir”, “ji”, crore and lakh amounts, Indian institutions, cities and company-specific vocabulary. That context matters when the output is meant for investors.

Names are often caught correctly

Sarvam picked up management names and proper nouns such as Sanjay Kumar Jain, Feroz Aziz, Jugal Mantri, Agra, Mathura and Prayagraj. Proper nouns still need review, but the baseline is encouraging.

Long-form structure is usable

Time markers and speaker turns make a long call easy to navigate, search and convert into first-pass notes.

Financial speech is mostly intelligible

Terms such as AUM, EBITDA, PAT, GST, CSR, ECL and margins appeared in the transcripts. Individual numbers remain high-risk, but the model is operating in the right domain.

Mixed-language speech is a real strength

Hindi-heavy sections can become readable English without losing the broad structure of the conversation, which is valuable for downstream note-taking.

Strength case 1 · Indian names and context · 0:00–0:36

This opening of the Narayana Hrudayalaya call introduces the management team and company-specific names.

Ground truth: The company is Narayana Hrudayalaya Limited. The speakers introduced include Dr. Emmanuel Rupert and Sandhya Jayaraman.
Sarvam output: “Good afternoon everyone and welcome to the quarter four FY26 earnings call of Narayana Hrudayalaya Limited … Dr. Emmanuel Rupert, CEO and MD … Sandhya Jayaraman, Group CFO …”
Assessment: Sarvam preserves the company name, management identities and the structure of the opening cleanly.
Strength case 2 · Hindi-heavy close to readable English · 00:45:47–00:47:03

IRCTC Q4 FY26

The closing remarks switch heavily into Hindi and test whether the output remains useful to an English-language research workflow.

Ground truth: Manoj Kumar Sharma introduces himself as Director of Catering Services, thanks investors and discusses FY26 revenue growth, PAT and continued support.
Sarvam output: “Many greetings to all of you. I am Manoj Kumar Sharma, Director of Catering Services. First of all, I congratulate all investors on behalf of IRCTC. Because of your support and trust, IRCTC has delivered good performance. Financially, in the year 25-26, revenue grew by 12% through operations, and our PAT is around 6% ... your continued support motivates us to move forward and build long-term, sustainable growth. Thank you very much.”
Assessment: The result is readable English that retains the broad meaning and structure of the Hindi-heavy close. The financial figures should still be checked before publication.

Where the model needs review

A conference call contains more than speech: line hiss, static, clipped microphones, room echo, muted participants and overlapping voices. Saaras does not always mark that uncertainty. Its most important failure mode is producing a plausible sentence from audio that is incomplete or not speech at all.

Speaker IDs are not identities

“SPEAKER_13” separates turns but does not identify a person. Crosstalk can also split one person into multiple IDs or merge two voices.

Numbers need special attention

Rupee amounts, crores, percentages and fiscal years are business-critical. Every number used in research should be checked against the recording or source filing.

Time markers are approximate

Timestamps help locate a passage, but they do not establish the precise moment each individual word was spoken.

Silence and overlap need review

A gap may be silence; a fragment may be cross-talk. Both should trigger review rather than confident reconstruction.

Source-quality stress test · Degraded margin question · 12:15–13:25

This IRCTC section contains a question about the internet-ticketing EBIT margin on a badly degraded line. Parts of the question repeat or smear together, so the precise percentages are not reliable from the recording alone.

Ground truth: “My question is on the internet-ticketing EBIT margin …” The remainder is too distorted to quote as a clean numerical statement.
Assessment: Source quality alone can make a transcript unsafe. A polished sentence should not be treated as validation when the underlying line is this degraded.

Confirmed failure cases

The following examples have enough audible evidence to identify a specific failure. Each card separates the supported ground truth from Sarvam’s output and names the failure mode.

Failure case 1 · Noise becomes words · 12:11–12:46

IRCTC Q4 FY26

Background chatter and cross-talk enter the line before the internet-ticketing margin question.

Ground truth: The line is affected by background noise and overlapping speech; the inserted statement about importing and a 9% duty is not supported by the surrounding question.
Sarvam output: “Right now, one day, I'm sorry, we are importing, our duty is 9% correct, we are importing. Sir, can you please repeat your question?”
Failure mode: Background interference is converted into a fluent but contextually unsupported sentence.
Failure case 2 · Unsupported insertion: “Perfect” · ~00:29:30

IRCTC Q4 FY26

The clip moves from the end of a passenger-traffic question through “Hello” and into management’s reply.

Ground truth: “…overall increase in passenger traffic? Hello. Yeah, in our tourism, our this quarter…” No one says “Perfect”.
Sarvam output:
“SPEAKER_8: Perfect”
Failure mode: An unsupported word is inserted as a separate speaker turn.
Failure case 3 · Speaker-turn error: “already” · ~00:30:52

IRCTC Q4 FY26

Madhu Chanda Dey is finishing her question about the April–June quarter before management replies.

Ground truth: “…because we are already one and a half months into it. You want me to comment on a thing which I should not because it’s a listed company, but yes, we are very hopeful.”
Sarvam output: Breaks “already” out of the question as a new “SPEAKER_8” turn before management’s reply.
Failure mode: A word belonging to the question is detached and assigned to a different speaker turn, changing the conversational structure.
Failure case 4 · Ambient speech promoted to dialogue · 00:38:06

IRCTC Q4 FY26

A short side comment leaks through when management’s microphone is unmuted after a question about CSR.

Ground truth: The question asks why CSR rose sharply. “Previous profit” is audible as an isolated side comment, not as part of the Q&A response.
Sarvam output: Assigns “Previous profit” to a separate speaker turn between the question and management’s answer.
Failure mode: Real but irrelevant ambient speech is promoted into the main dialogue, where it can mislead a reader.
Failure case 5 · Context-changing substitution: “Ooty” · 00:31:18

IRCTC Q4 FY26

Sonal Minhas introduces herself and checks whether she is audible on a noisy line.

Ground truth: “Hi, this is Sonal Minhas. I hope I am audible.”
Sarvam output: “My name is Sonal Minas. I am from Ooty.”
Failure mode: “I hope I am audible” becomes an Indian place name, changing both the sentence and the apparent context of the speaker’s introduction.

Unresolved stress cases

Two clips surfaced suspicious proper nouns, but the source audio is not clear enough to establish a defensible verbatim ground truth. They remain useful review triggers, not confirmed errors.

Unresolved case 1 · “Gopal” or “got it”? · 00:16:38–00:16:45

IRCTC Q4 FY26

Ground truth: “You got my point. Thank you.” is clear. The handoff immediately after it is acoustically ambiguous in this 8 kHz clip.
Sarvam output:
“SPEAKER_9: Okay sir, Gopal.”
Assessment: “Gopal” is a business-sensitive proper noun and should be flagged for review, but this clip alone cannot prove that it was hallucinated.
Unresolved case 2 · “Titan” on a degraded line · 00:41:02

IRCTC Q4 FY26

Ground truth: The line is too degraded to establish “Titan” versus “better now” confidently from this clip alone.
Sarvam output:
“SPEAKER_10: Ah, is it Titan now?”
“SPEAKER_0: Is it better now?”
Assessment: The output contains a potentially consequential brand name, but a trusted reference transcript would be needed to call the rendering definitively wrong.

The conclusion: Sarvam matters

Sarvam is one of India’s most consequential home-grown AI labs, and Saaras shows why. It is not at the global state of the art yet: the model still needs stronger uncertainty handling when audio becomes noisy or incomplete. But it excels where Indian context matters—names, code-switching, business vocabulary and the natural rhythm of local conversations.

For now, Saaras is best used as a fast first-draft engine, not a source of record. Use it to create the searchable transcript, then verify proper nouns, financial figures, guidance and every passage affected by static or overlapping speech. Better uncertainty handling could move it from a useful assistant to a dependable production layer.

That gap should not obscure the larger point. The brief withdrawal of Anthropic’s Fable 5 from international availability showed how quickly access to a frontier model can change because of decisions made outside India. India needs strong AI systems of its own—not as a patriotic substitute for performance, but as durable infrastructure. Sarvam is not the finished answer, but it is one of the most credible attempts to build it.