CalQuity
Insight·16 min read·22 July 2026
FinanceAI

Sarvam vs AssemblyAI: Transcribing Indian Earnings Calls in Real Time

A practical comparison of Sarvam Saaras v3 and AssemblyAI across four Indian earnings calls, with audio examples of names, mixed-language speech, static, and cross-talk.

INDIAN EARNINGS CALL · AUDIO → TRANSCRIPTAUDIOTRANSCRIPTNAMES · NUMBERS · NOISE
CalQuity field notes · Earnings calls

CalQuity covers earnings calls in real time, so we routinely evaluate transcription models. In this review, we compare a global benchmark with India’s local specialist to see which one is more dependable when the audio and conversation are less than perfect.

The short verdict Sarvam Saaras v3 is a promising choice for Indian business audio. It handled the general shape of long earnings calls well, produced usable speaker-labelled transcripts, and showed a useful feel for Indian names, places and financial vocabulary. Its weak point is audio quality: when static, muffling or crosstalk enters the recording, the model can confidently turn noise into words. That makes it a strong first-pass transcription engine—not a hands-off source of record.

Why we keep testing transcription models

At CalQuity, earnings calls are not something we can wait to process later. We need to follow what management and analysts are saying as the call unfolds, turn it into usable research and make sure important details do not disappear in the noise. A transcript is the foundation for that work.

The challenge: earnings calls are messy

A typical call is not a studio recording. There may be several speakers, an operator managing the queue, people joining from different locations, muffled microphones, background chatter, static, overlapping voices and participants asking one another to repeat a question. Speakers may also switch between English, Hindi and other Indian languages in the middle of a sentence. The conversation moves quickly through company names, Indian places, rupee amounts, crores, lakhs, percentages and financial jargon.

That is why a transcript can look polished and still be wrong. A missed name is inconvenient; a missed number or guidance comment can change the meaning of an earnings call. The best model is the one that knows the difference between speech and noise—and knows when it is uncertain.

Against that backdrop, we compare AssemblyAI, a global benchmark, with Sarvam, one of India’s strongest home-grown speech models. The question is not simply which one produces the most fluent English. It is which one is more dependable for the real conditions of Indian earnings calls.

Where Sarvam impressed us

Indian context feels natural

The transcripts retained the texture of Indian business conversation: “sir”, “ji”, crore and lakh amounts, Indian institutions, cities and company-specific names. That matters more than a generic word-accuracy score when the final document is meant for investors.

Names are often caught correctly

Across the calls, Sarvam picked up management names and Indian proper nouns such as Sanjay Kumar Jain, Feroz Aziz, Jugal Mantri, Agra, Mathura and Prayagraj. Proper nouns still need review, but the baseline is encouraging.

Long-form structure is usable

The transcripts are easy to navigate, with clear time markers and speaker turns that help locate a question or revisit a disputed passage quickly.

Financial speech is mostly intelligible

Terms and quantities such as AUM, EBITDA, PAT, GST, CSR, ECL, margins and rupee-denominated figures appeared in the transcripts. The model is clearly operating in the right domain, even though individual numbers and terms remain high-risk.

Hindi-to-English conversion in mixed-language closes

In the IRCTC closeout, Sarvam translated a Hindi-heavy segment at ~00:45:47–00:47:03 into readable English, keeping meaning and structure intact. This is exactly the edge where multilingual transcript quality matters most in Indian calls.

Listen: Indian names and context · 0:00–0:36

This opening of the Narayana Hrudayalaya call introduces the management team and company-specific names.

Sarvam: “Good afternoon everyone and welcome to the quarter four FY26 earnings call of Narayana Hrudayalaya Limited … Dr. Emmanuel Rupert, CEO and MD … Sandhya Jayaraman, Group CFO …”
AssemblyAI Universal 3.5 pro: “Good afternoon everyone and welcome to the Q4 FY26 earnings call of Narayana Rudrayalaya Limited … Dr. Emmanuel Rupert, CEO and MD … Sandhya Jayaraman, Group CFO …”
Listen: Hindi Q&A to English conversion · IRCTC closeout · 00:45:47–00:47:03

IRCTC Q4 FY26

This section is at the end of the call and is useful for testing transliteration and translation behavior.

Sarvam Saaras v3: “Many greetings to all of you. I am Manoj Kumar Sharma, Director of Catering Services. First of all, I congratulate all investors on behalf of IRCTC. Because of your support and trust, IRCTC has delivered good performance. Financially, in the year 25-26, revenue grew by 12% through operations, and our PAT is around 6% ... your continued support motivates us to move forward and build long-term, sustainable growth. Thank you very much.”
AssemblyAI Universal 3.5 pro (auto-translate enabled to English): “जी नमस्कार आप सभों को बहुत-बहुत मस्ताक है मनोज सुमा शर्मा बनाऊं डायरेक्टर कैटिंग सर्विसेज सबसे पहले मैं आरसीटीसी की तरफ से आप सभी इन्वेस्टर्स को बधाई देना चाहूँगा क्योंकि आप लोगों के सहयोग और समर्थन से और आपके विश्वास से आरसीटीसी ने काफी अच्छा परफॉर्मेंस दिया है फाइनेंश पचीस छब्बीस में जो हमारा रेवेन्यू है ऑपरेशन के माध्यम से वो बारह परसेंट ग्रो किया है और हमारा पीएटी जो है वो है आपको सिक्स परसेंट करोड़ सौ किया है और मैं आप लोग के लगातार सतत सहयोग और विश्वास जो कंपनी में आप लोग का है उसके लिए मैं पुनः धन्यवाद करता हूँ आपका आपका द्वारा दिया गया पुरस्कार, मोटिवेशन हमलोग को आगे बढ़ने के लिए प्रेरित करता है ...”
Takeaway: In this segment, AssemblyAI produced Hindi text despite English auto-translate, while Sarvam delivered an English-readable output suitable for downstream note-taking.

The uncomfortable part: noise becomes language

Conference calls contain a lot of audio that is not speech: line hiss, static, clipped microphones, hold music, room echo and overlapping voices. In our tests, static was not always treated as uncertainty. On noisy stretches, Sarvam could interpret the interference as plausible words. That is the most important limitation because the resulting sentence may look polished while being wrong.

There were also audible line-quality problems in the calls themselves. The transcripts include moments such as “voice was muffled”, requests to repeat a question, and participants being told they were not audible. These are exactly the moments where a reviewer should slow down. A clean-looking transcript is not proof that the underlying audio was clear.

Speaker IDs are not identities

“SPEAKER_13” is useful for separating turns, but it is not a person's name. Diarization can also split one person into multiple IDs or merge voices during crosstalk.

Numbers need special attention

₹, crores, percentages and fiscal-year references are business-critical. A small recognition error can change the meaning of a result, so every number used in a published article or research note should be checked against the recording or source filings.

Time markers are approximate

The time markers are useful for finding a passage, but they do not identify the exact moment each individual word was spoken.

Silence and overlap need review

Gaps may be silence, and overlaps may be crosstalk or a speaker-separation error. They help identify passages to inspect; they do not measure accuracy.

Listen: a noisy exchange · 12:15–13:25

This IRCTC section contains a question about the internet-ticketing EBIT margin on a badly degraded line. Parts of the question repeat or smear together, so the precise percentages are not reliable from the recording alone.

Audible reference: “My question is on the internet-ticketing EBIT margin …” The remainder is heavily distorted and should not be quoted as a clean numerical statement.
Why this clip matters: This is a source-quality example rather than a clean model-to-model quotation. The following, shorter clip isolates the point where background interference is converted into plausible words.
Listen: when noise becomes words · 12:11–12:46

IRCTC Q4 FY26 · timestamp: 00:12:20

This clip isolates the most revealing moment. You can hear the background chatter and cross-talk in the source audio. Sarvam then turns that interference into a sentence that sounds coherent but does not fit the surrounding question.

Sarvam transcript: “Right now, one day, I'm sorry, we are importing, our duty is 9% correct, we are importing. Sir, can you please repeat your question?”
AssemblyAI transcript: “We cannot fully hear you because of noise from background.”
Comparison: AssemblyAI filters this irrelevant background noise out of the main transcript at this point and focuses on the fact that the speaker is not audible. Sarvam instead attempts to transcribe the cross-talk, producing the confusing “we are importing” passage.
What this shows: This is strong evidence that Sarvam can convert background interference or badly overlapped speech into plausible words. It is not a formal acoustic annotation, so the fairest claim is “observed noise-related hallucination,” not a measured error rate.

A case study in transcript disagreements: the IRCTC call

The IRCTC Q4 call surfaced several failure patterns worth reviewing: disputed proper nouns, floating single-word insertions, ambient speech leaking from muted microphones, and severe cross-talk confusions. Here are six concrete examples anchored to the supplied audio. Where the recording is too degraded to establish ground truth, we describe a model disagreement rather than declaring a hallucination.

How to read these examples: The “audible reference” is limited to speech that can be supported by the attached clip. Sarvam and AssemblyAI sections show model output, so a word that is absent from the recording may be the evidence being discussed—not a caption of what you should expect to hear.
1. “Gopal” or “got it”? · 00:16:38–00:16:45

IRCTC Q4 FY26

Kanishka Gupta has just finished a thought about the August Kranti Rajdhani route. A likely microphone bump or burst of static follows. Sarvam turns the noise into a person’s name; AssemblyAI treats the moment as a clean handoff into the next question.

Sarvam transcript (lines around 00:16:38):
“SPEAKER_8: You got my point. Thank you.”
“SPEAKER_9: Okay sir, Gopal.”
“SPEAKER_9: And sir the second question would be on the convenience fees across AC and non-AC ticket categories…”
AssemblyAI transcript (lines around 00:16:40):
“So the route is not there. You got my point?”
“Thank you. Okay, sir. Got it. And sir, the second question would be on the convenience fees across AC and non-AC ticket categories have largely remained unchanged for several years…”
Audio-audit note: “You got my point, thank you” is clear. The handoff immediately after it is acoustically ambiguous in this 8 kHz recording: Sarvam resolves it as “Gopal”, while AssemblyAI resolves it as “got it”. Without a trusted reference transcript, this is evidence of a proper-noun disagreement—not definitive proof that “Gopal” was hallucinated.
2a. The floating “Perfect.” · ~00:29:30

IRCTC Q4 FY26

The clip begins at the end of a question about passenger traffic, followed by a short pause, “Hello”, and the start of management’s reply. Sarvam inserts “Perfect” as a separate turn at this transition even though that word is not audible in the clip.

Audible reference: “…number of tickets booked or holidays booked or, you know, overall increase in passenger traffic. Hello. Yeah, in our tourism, our this quarter…”
Sarvam model output (exact line):
“SPEAKER_8: Perfect”
AssemblyAI model output (exact lines):
“…overall increase in passenger traffic?”
“Hello.”
“Yeah, in our tourism, our this quarter…”
What this shows: The audio moves directly from the passenger-traffic question into “Hello” and management’s reply. Only Sarvam inserts “Perfect” as a separate, unsupported speaker turn.
2b. The floating “already.” · ~00:30:52

IRCTC Q4 FY26

Madhu Chanda Dey is wrapping up her follow-up about the April–June quarter. Sarvam detaches the word “already.” from her sentence and emits it on its own line before Sanjay’s cautious reply.

Audible reference: “…the April–June quarter, because we are already one and a half months into it. You want me to comment on a thing which I should not because it’s a listed company, but yes, we are very hopeful. Thank you. Okay, thank you.”
Sarvam model output: Breaks “already” out of the question as a new “SPEAKER_8” turn before management’s reply.
AssemblyAI model output: Keeps “already” inside the question and continues directly into management’s reply.
What this shows: The word “already” is part of Madhu’s question, not a separate speaker’s reply. Sarvam breaks it out as its own turn, then Sanjay’s answer begins. AssemblyAI keeps the sentence intact, so the relationship between Madhu’s question and Sanjay’s reply stays readable.
3. The “Previous profit.” glitch · 00:38:06

IRCTC Q4 FY26

Karthik has just asked why CSR jumped from ₹7 crore to ₹31 crore. Management’s mic is on mute at that instant, and the moment it un-mutes, a burst of background chatter on the management side is briefly audible. Someone on management’s end actually says the words “previous profit” in passing — it is real speech leaking through, not static. The two engines handle that leak very differently.

Audible reference: “…what is the rationale for such a sharp increase in CSR?” A short side comment follows, then: “You see, CSR, as you must be knowing, it depends on the profitability of the last three years…”
Sarvam model output: Assigns the side comment “Previous profit” to a separate speaker turn.
AssemblyAI model output: Drops the side comment and moves directly from the question into management’s answer.
What this shows: “Previous profit” is not a hallucination in the strict sense — it really was said, by someone on the management side whose mic was supposed to be muted. Sarvam captures the words as a separate speaker turn and prints them between Karthik’s question and the management answer. AssemblyAI recognises the leak as ambient noise from a side conversation and drops it cleanly, so the Q&A reads as a continuous exchange. Both behaviours are defensible in isolation, but only AssemblyAI’s is useful in a research transcript: the “previous profit” fragment is meaningless without the surrounding context, and a reader is more likely to be misled by it than informed.
4. The “Ooty” hallucination · 00:31:18

IRCTC Q4 FY26

Sonal Minhas introduces herself on a noisy line. Acoustic distortion causes Sarvam to hear an Indian hill station that was never mentioned.

Sarvam transcript (lines around 00:31:09–00:31:32):
“SPEAKER_0: Thank you. Next question is from the line of Sonal from Presian Capital. Please go ahead.”
“SPEAKER_10: My name is Sonal Minas. I am from Ooty.”
“SPEAKER_0: So your voice is not clear.”
“SPEAKER_10: Am I audible or should I speak again?”
“SPEAKER_8: Yeah, please, please, go ahead.”
AssemblyAI transcript (lines around 00:31:07–00:31:32):
“Thank you. Next question is from the line of Sonal from Precian Capital. Please go ahead.”
“Hi, this is Sonal Minas. I hope I am audible.”
“So your voice?”
“Am I audible or should I speak again?”
“Yeah, please, please go ahead.”
What this shows: The questioner is not “from Ooty”; she is asking whether she is audible. Under accent pressure and poor line quality, Sarvam aggressively guesses English words and completely changes the context of the conversation. A single hallucinated location in a transcript is enough to send a reader to a completely wrong conclusion about who is speaking.
5. The “Titan” hallucination · 00:41:02

IRCTC Q4 FY26

Harsh Yadav is mid-question about the RBI payment-aggregator deadline. After a short audio hiccup, the operator asks if the line is better. Sarvam turns the operator’s question into a brand name.

Sarvam transcript (lines around 00:40:55–00:41:11):
“SPEAKER_10: Okay, one more question. So, the RBI deadline for”
“SPEAKER_0: RBI deadline”
“SPEAKER_10: Ah, is it Titan now?”
“SPEAKER_0: Is it better now?”
“SPEAKER_10: Am I audible now?”
“SPEAKER_0: Yes sir, this is better.”
AssemblyAI transcript (lines around 00:40:55–00:41:11):
“Okay, one more question. So RBI deadline for submitting final application for payment aggregator license has been extended to August 26th…”
“Harsh sir, you’re not audible if you’re speaking anything.”
“Yes, I finished as a startup, but it’s okay, I got it.”
Audio-audit note: The line is too degraded to establish “Titan” versus “better now” confidently from this clip alone. The defensible finding is that the systems disagree at a business-sensitive proper noun; a trusted reference transcript would be needed to call either rendering definitive.

Six moments, several different failure modes. None of them is a measured error rate. Some are clear insertions, while others are disagreements that cannot be resolved confidently from an 8 kHz call recording alone. That distinction matters when deciding what the transcript can safely support.

A snapshot of the results

Word totals are directional rather than an accuracy benchmark. The Narayana Sarvam range reflects three separate runs.

The conclusion: Sarvam matters

Sarvam is one of India’s most consequential home-grown AI labs, and Saaras shows why. It is not at the global state of the art yet: on noisy calls, AssemblyAI is still better at recognising when interference should not become text. But Sarvam excels where Indian context matters—names, code-switching, business vocabulary and the natural rhythm of local conversations.

For now, Saaras is best used as a fast first-draft engine, not a source of record. Use it to create the searchable transcript, then verify proper nouns, financial figures, guidance and every passage affected by static or overlapping speech. Better uncertainty handling could move it from a useful assistant to a dependable production layer.

That gap should not obscure the larger point. The brief withdrawal of Anthropic’s Fable 5 from international availability showed how quickly access to a frontier model can change because of decisions made outside India. India needs strong AI systems of its own—not as a patriotic substitute for performance, but as durable infrastructure. Sarvam is not the finished answer, but it is one of the most credible attempts to build it.