CalQuity covers earnings calls in real time, so we routinely evaluate transcription models. In this review, we compare a global benchmark with India’s local specialist to see which one is more dependable when the audio and conversation are less than perfect.
Why we keep testing transcription models
At CalQuity, earnings calls are not something we can wait to process later. We need to follow what management and analysts are saying as the call unfolds, turn it into usable research and make sure important details do not disappear in the noise. A transcript is the foundation for that work.
The challenge: earnings calls are messy
A typical call is not a studio recording. There may be several speakers, an operator managing the queue, people joining from different locations, muffled microphones, background chatter, static, overlapping voices and participants asking one another to repeat a question. Speakers may also switch between English, Hindi and other Indian languages in the middle of a sentence. The conversation moves quickly through company names, Indian places, rupee amounts, crores, lakhs, percentages and financial jargon.
That is why a transcript can look polished and still be wrong. A missed name is inconvenient; a missed number or guidance comment can change the meaning of an earnings call. The best model is the one that knows the difference between speech and noise—and knows when it is uncertain.
Against that backdrop, we compare AssemblyAI, a global benchmark, with Sarvam, one of India’s strongest home-grown speech models. The question is not simply which one produces the most fluent English. It is which one is more dependable for the real conditions of Indian earnings calls.
Where Sarvam impressed us
Indian context feels natural
The transcripts retained the texture of Indian business conversation: “sir”, “ji”, crore and lakh amounts, Indian institutions, cities and company-specific names. That matters more than a generic word-accuracy score when the final document is meant for investors.
Names are often caught correctly
Across the calls, Sarvam picked up management names and Indian proper nouns such as Sanjay Kumar Jain, Feroz Aziz, Jugal Mantri, Agra, Mathura and Prayagraj. Proper nouns still need review, but the baseline is encouraging.
Long-form structure is usable
The transcripts are easy to navigate, with clear time markers and speaker turns that help locate a question or revisit a disputed passage quickly.
Financial speech is mostly intelligible
Terms and quantities such as AUM, EBITDA, PAT, GST, CSR, ECL, margins and rupee-denominated figures appeared in the transcripts. The model is clearly operating in the right domain, even though individual numbers and terms remain high-risk.
Hindi-to-English conversion in mixed-language closes
In the IRCTC closeout, Sarvam translated a Hindi-heavy segment at ~00:45:47–00:47:03 into readable English, keeping meaning and structure intact. This is exactly the edge where multilingual transcript quality matters most in Indian calls.
This opening of the Narayana Hrudayalaya call introduces the management team and company-specific names.
IRCTC Q4 FY26
This section is at the end of the call and is useful for testing transliteration and translation behavior.
The uncomfortable part: noise becomes language
Conference calls contain a lot of audio that is not speech: line hiss, static, clipped microphones, hold music, room echo and overlapping voices. In our tests, static was not always treated as uncertainty. On noisy stretches, Sarvam could interpret the interference as plausible words. That is the most important limitation because the resulting sentence may look polished while being wrong.
There were also audible line-quality problems in the calls themselves. The transcripts include moments such as “voice was muffled”, requests to repeat a question, and participants being told they were not audible. These are exactly the moments where a reviewer should slow down. A clean-looking transcript is not proof that the underlying audio was clear.
Speaker IDs are not identities
“SPEAKER_13” is useful for separating turns, but it is not a person's name. Diarization can also split one person into multiple IDs or merge voices during crosstalk.
Numbers need special attention
₹, crores, percentages and fiscal-year references are business-critical. A small recognition error can change the meaning of a result, so every number used in a published article or research note should be checked against the recording or source filings.
Time markers are approximate
The time markers are useful for finding a passage, but they do not identify the exact moment each individual word was spoken.
Silence and overlap need review
Gaps may be silence, and overlaps may be crosstalk or a speaker-separation error. They help identify passages to inspect; they do not measure accuracy.
This IRCTC section contains a question about the internet-ticketing EBIT margin on a badly degraded line. Parts of the question repeat or smear together, so the precise percentages are not reliable from the recording alone.
IRCTC Q4 FY26 · timestamp: 00:12:20
This clip isolates the most revealing moment. You can hear the background chatter and cross-talk in the source audio. Sarvam then turns that interference into a sentence that sounds coherent but does not fit the surrounding question.
A case study in transcript disagreements: the IRCTC call
The IRCTC Q4 call surfaced several failure patterns worth reviewing: disputed proper nouns, floating single-word insertions, ambient speech leaking from muted microphones, and severe cross-talk confusions. Here are six concrete examples anchored to the supplied audio. Where the recording is too degraded to establish ground truth, we describe a model disagreement rather than declaring a hallucination.
IRCTC Q4 FY26
Kanishka Gupta has just finished a thought about the August Kranti Rajdhani route. A likely microphone bump or burst of static follows. Sarvam turns the noise into a person’s name; AssemblyAI treats the moment as a clean handoff into the next question.
“SPEAKER_8: You got my point. Thank you.”
“SPEAKER_9: Okay sir, Gopal.”
“SPEAKER_9: And sir the second question would be on the convenience fees across AC and non-AC ticket categories…”
“So the route is not there. You got my point?”
“Thank you. Okay, sir. Got it. And sir, the second question would be on the convenience fees across AC and non-AC ticket categories have largely remained unchanged for several years…”
IRCTC Q4 FY26
The clip begins at the end of a question about passenger traffic, followed by a short pause, “Hello”, and the start of management’s reply. Sarvam inserts “Perfect” as a separate turn at this transition even though that word is not audible in the clip.
“SPEAKER_8: Perfect”
“…overall increase in passenger traffic?”
“Hello.”
“Yeah, in our tourism, our this quarter…”
IRCTC Q4 FY26
Madhu Chanda Dey is wrapping up her follow-up about the April–June quarter. Sarvam detaches the word “already.” from her sentence and emits it on its own line before Sanjay’s cautious reply.
IRCTC Q4 FY26
Karthik has just asked why CSR jumped from ₹7 crore to ₹31 crore. Management’s mic is on mute at that instant, and the moment it un-mutes, a burst of background chatter on the management side is briefly audible. Someone on management’s end actually says the words “previous profit” in passing — it is real speech leaking through, not static. The two engines handle that leak very differently.
IRCTC Q4 FY26
Sonal Minhas introduces herself on a noisy line. Acoustic distortion causes Sarvam to hear an Indian hill station that was never mentioned.
“SPEAKER_0: Thank you. Next question is from the line of Sonal from Presian Capital. Please go ahead.”
“SPEAKER_10: My name is Sonal Minas. I am from Ooty.”
“SPEAKER_0: So your voice is not clear.”
“SPEAKER_10: Am I audible or should I speak again?”
“SPEAKER_8: Yeah, please, please, go ahead.”
“Thank you. Next question is from the line of Sonal from Precian Capital. Please go ahead.”
“Hi, this is Sonal Minas. I hope I am audible.”
“So your voice?”
“Am I audible or should I speak again?”
“Yeah, please, please go ahead.”
IRCTC Q4 FY26
Harsh Yadav is mid-question about the RBI payment-aggregator deadline. After a short audio hiccup, the operator asks if the line is better. Sarvam turns the operator’s question into a brand name.
“SPEAKER_10: Okay, one more question. So, the RBI deadline for”
“SPEAKER_0: RBI deadline”
“SPEAKER_10: Ah, is it Titan now?”
“SPEAKER_0: Is it better now?”
“SPEAKER_10: Am I audible now?”
“SPEAKER_0: Yes sir, this is better.”
“Okay, one more question. So RBI deadline for submitting final application for payment aggregator license has been extended to August 26th…”
“Harsh sir, you’re not audible if you’re speaking anything.”
“Yes, I finished as a startup, but it’s okay, I got it.”
Six moments, several different failure modes. None of them is a measured error rate. Some are clear insertions, while others are disagreements that cannot be resolved confidently from an 8 kHz call recording alone. That distinction matters when deciding what the transcript can safely support.
A snapshot of the results
Word totals are directional rather than an accuracy benchmark. The Narayana Sarvam range reflects three separate runs.
The conclusion: Sarvam matters
Sarvam is one of India’s most consequential home-grown AI labs, and Saaras shows why. It is not at the global state of the art yet: on noisy calls, AssemblyAI is still better at recognising when interference should not become text. But Sarvam excels where Indian context matters—names, code-switching, business vocabulary and the natural rhythm of local conversations.
For now, Saaras is best used as a fast first-draft engine, not a source of record. Use it to create the searchable transcript, then verify proper nouns, financial figures, guidance and every passage affected by static or overlapping speech. Better uncertainty handling could move it from a useful assistant to a dependable production layer.
That gap should not obscure the larger point. The brief withdrawal of Anthropic’s Fable 5 from international availability showed how quickly access to a frontier model can change because of decisions made outside India. India needs strong AI systems of its own—not as a patriotic substitute for performance, but as durable infrastructure. Sarvam is not the finished answer, but it is one of the most credible attempts to build it.
