CalQuity
Insight·9 min read·22 July 2026
FinanceAI

Sarvam Saaras v4 Review: A Better Listener for Indian Earnings Calls?

A hands-on review of Sarvam Saaras v4 across earnings calls, with audio-verified examples covering names, code-mixed speech, diarization, crosstalk, and hallucinations.

INDIAN EARNINGS CALL · AUDIO → TRANSCRIPTAUDIOTRANSCRIPTNAMES · NUMBERS · NOISE
CalQuity field notes · Earnings calls

CalQuity covers thousands of earnings calls in real time. Therefore, we routinely evaluate speech-to-text models. This field review tests Sarvam’s Saaras v4 multi-speaker model for transcription quality, speaker structure, Indian language context, and difficult conference-call audio.

The short verdict Saaras v4 is a capable first-draft engine for Indian earnings calls. It handles mixed Hindi and English naturally in transcribe mode, preserves useful speaker-labelled structure across long calls, and avoids several errors found when benchmarked against the Saaras v3 model it replaces. But during severe crosstalk and bad connections, it can still turn uncertain audio into confident words.

The challenge: earnings calls are messy

A typical call has an operator, management, analysts joining from different locations, muffled microphones, static, background conversations and overlapping voices. Speakers switch between English, Hindi and other Indian languages while discussing company names, Indian places, crores, percentages and financial jargon. A transcript can look polished and still be wrong.

That makes this a test of judgment as much as recognition. A useful model must identify speech, separate speakers and preserve Indian context—but it must also resist inventing certainty when the line is unclear.

How Saaras v4 performed

Fewer phantom words

Two conspicuous name-like errors from v3 disappeared in v4. Both examples are explained in detail below.

Strong multi-speaker structure

Across the four calls, v4 produced 11 to 19 synthetic speaker IDs and timestamped turns that remained practical for navigation and review.

Useful financial context

AUM, PAT, CSR, RBI deadlines, margins, crores and company-specific language generally remained intelligible across the long recordings.

Noise errors still survive

V4 retained a detached “Already” and misrecognised Sonal Minhas’s audibility check. Both cases are explained in detail below.

How to read the examples: “V3 API output” and “V4 API output” reproduce verbatim transcript excerpts supplied with the API runs. An ellipsis marks omitted surrounding text; it is not model output. Each example then separates the audio-verified reference from our inference.
Proper-noun accuracy · 00:00–00:36
V4 API output — verbatim excerpt: “Good afternoon everyone and welcome to the Quarter 4 FY26 earnings call of Narayana Hrudayalaya Limited … Mr. Virendra Shetty, Vice Chairman …”
Audio-verified reference: “Narayana Hrudayalaya Limited … Mr. Viren Shetty, Vice Chairman …”
Inference: V4 recognises the difficult company name but changes “Viren” to “Virendra.” Management names still require verification.
Hindi-English closeout · 00:45:47–00:47:03
V4 API output — transcribe mode, verbatim excerpt: “आप सबों को बहुत-बहुत नमस्कार। मैं मनोज कुमार शर्मा बोल रहा हूँ Director Catering Services … Financial year 25-26 में जो हमारा Revenue है Operation के माध्यम से वो 12% Grow किया है …”
V4 API output — translate mode, verbatim excerpt: “A very warm welcome to all of you. I am Manoj Kumar Sharma, Director of Catering Services … In the financial year 25-26, our revenue through operations has grown by 12% …”
Audio-verified reference: The speaker identifies himself as Manoj Kumar Sharma, states his catering-services role, and reports 12% revenue growth through operations.
Inference: These are two deliberate API modes, not competing outputs from one run. Transcribe preserves the code-mixed speech; translate produces readable English without losing the key facts in this excerpt.

The hardest test: noisy IRCTC moments

IRCTC’s call contains a concentrated stretch of static, background conversation and participants asking whether they are audible. These are not cherry-picked clean sentences; they are the places where transcription engines reveal whether they can separate speech from interference.

Conversational phrase: “got it” · 00:16:38–00:16:45
V3 API output — verbatim excerpt: “You got my point. Thank you.” / “Okay sir, Gopal.”
V4 API output — verbatim excerpt: “You got my point. Thank you.” / “Okay sir, got it.”
Audio-verified reference: “You got my point. Thank you.” / “Okay, sir. Got it.”
Inference: V4 correctly recognises “got it” as a conversational acknowledgement rather than inventing a person named “Gopal,” as v3 did.
Speaker attribution: detached “Already” · ~00:30:52
V4 API output — complete standalone turn: “Already.”
Audio-verified reference: Analyst: “…because we are already one and a half months into it.” Management then begins: “You want me to comment on a thing which I should not because it’s a listed company…”
Inference: The API duplicates “already” and assigns the duplicate to a separate speaker turn. The word belongs only inside the analyst’s question, so this is a diarization error that changes the apparent exchange.
Audibility check misrecognised · 00:31:18
V4 API output — verbatim excerpt: “I am Sonal Minhas, I am a modeling.”
Audio-verified reference: “Hi, this is Sonal Minhas. I hope I am audible.”
Inference: V4 gets the speaker’s name but misrecognises the audibility check. Poor line quality can still turn a routine introduction into a plausible-looking false phrase.
Connection-quality exchange · 00:41:02
V3 API output — verbatim excerpt: “Ah, is it Titan now?”
V4 API output — verbatim excerpt:
“Is it better now?”
“It’s better now.”
“Am I audible now?”
Audio-verified reference: The same three-line exchange: one speaker asks if the connection is better, another confirms, and the first checks audibility.
Inference: These are three speaker turns from one V4 run—not three output modes. V4 correctly captures the connection check and avoids v3’s false “Titan” reference.
Background speech treated as main conversation · 00:38:06
V3 API output — complete standalone turn: “Previous profit.”
V4 API output — complete standalone turn: “No previous profits.”
Audio-verified reference: A brief side comment about previous profit is audible, but it is ambient speech rather than part of the analyst-management exchange.
Inference: Even though this is not part of the main exchange, both v3 and v4 successfully pick up a faint background remark that passes through the recording. That sensitivity is useful. The remaining limitation is presentation: a background-speech label would help readers distinguish it from the analyst-management discussion, and the supplied reference does not verify v4’s exact wording.

A snapshot of the runs

CallV3 segmentsV3 wordsV3 speakersV4 segmentsV4 wordsV4 speakers
IRCTC Q4 & FY262615,968112585,95611
Anand Rathi Wealth Q1 FY272219,696132209,59613
Narayana Hrudayalaya Q4 FY26370–37611,528–11,6801435911,37914
Angel One Q1 FY2729213,1701928713,04219

These counts describe how the models structured their outputs; they do not measure accuracy. The v3 Narayana range reflects three separate runs. V4 produced slightly fewer segments and words on every call while retaining the same number of detected speakers.

So, is Saaras v4 good enough?

For a live first draft, yes—and more convincingly than v3. The multi-speaker output is navigable, Indian financial language generally survives, and several obvious word-level errors from v3 have disappeared.

For publication, investment research or automatic extraction of financial facts, it still needs a review layer. Proper nouns, financial figures and every passage affected by crosstalk or an audibility check should be verified against the recording or source filings.

Method note: We tested Sarvam’s Saaras v4 multi-speaker API with diarization and timestamps enabled. The calls used transcribe mode, with a separate translate-mode run for the Hindi closeout.