Change language to
0:00

NVIDIA says fine-tuning Nemotron 3.5 ASR on 133.7 hours of Saudi Arabic speech cut its word error rate on Najdi and Hijazi dialects from 55.05% to 29.96%. The result was measured on a target-dialect test split from SDAIA’s SADA dataset, according to the Saudi Press Agency.

The figures come from a fine-tuning recipe, not a new consumer app or a claim that every Arabic dialect is now solved. In its technical post, NVIDIA details how it adapted the multilingual speech-recognition model while checking that English and broader Arabic performance did not regress.

Subscribe to our Newsletters for more Tech Stories

What changed in NVIDIA’s Saudi Arabic speech test

The target split’s word error rate fell by 25.09 percentage points, a relative reduction of about 45.6%. NVIDIA also reported lower error rates across the full SADA test set and on the FLEURS English and Arabic checks.

Test set and metricBefore fine-tuningAfter fine-tuning
SADA Najdi + Hijazi — word error rate55.05%29.96%
Full SADA — word error rate58.84%35.61%
SADA Najdi + Hijazi — character error rate31.63%12.18%
Full SADA — character error rate35.40%15.97%
FLEURS English — word error rate11.04%10.42%
FLEURS Arabic — word error rate12.67%11.41%

The comparison is against NVIDIA’s own before-and-after tests, not a head-to-head evaluation against every speech-recognition service. SADA is a 668-hour Saudi Arabic corpus drawn from 57 Saudi Broadcasting Authority shows, covering read and spontaneous speech; the ICASSP 2024 paper introducing SADA describes its dialect and recording mix.

NVIDIA Nemotron 3.5 ASR: fine-tuning for Najdi and Hijazi

NVIDIA retained 103,559 of 125,490 utterances after filtering, giving it 133.7 hours of target speech. Its training mix was 90% Saudi Arabic, 7% English and 3% Arabic from FLEURS. Keeping a small replay of languages the model already knew is intended to reduce the risk that specialising it for Saudi speech weakens its other capabilities.

NVIDIA's Nemotron 3.5 ASR fine-tuning workflow for Najdi and Hijazi speech

The base model is described in NVIDIA’s Nemotron 3.5 ASR model card as a 600-million-parameter streaming recogniser. It lists 40 language-locales, but separates them into transcription-ready, broad-coverage and adaptation-ready tiers; the headline number is not a promise of equal accuracy across every locale.

NVIDIA also tested inference settings after training. A wider attention context and beam-8 MALSD decoding lowered the target word error rate by a further 2.71 points without another training run, at the cost of roughly 800 milliseconds of extra latency. The company says that trade-off is better suited to batch transcription than to applications where immediate responses matter.

Automatic speech recognition turns spoken audio into text; conversational systems add further steps such as generating a reply. That makes this work distinct from NVIDIA’s PersonaPlex voice model, which is built for real-time listening and speaking rather than transcription alone.

What the Saudi Arabic result does not establish

The test covers Najdi and Hijazi speech in SADA. It does not show equivalent performance for every Saudi dialect, Modern Standard Arabic in every setting, Emirati Arabic or noisy telephone conversations. A language appearing in a model card is not the same as a measured result on a company’s own callers.

For UAE organisations, the work is a method to evaluate, not evidence of a ready-made Emirati speech product. A deployment would need suitable labelled audio and testing against the dialects, accents, microphones and background noise it expects to encounter. Tbreak has previously covered du’s Arabic telecom language model, a separate project for internal telecom operations; NVIDIA’s post does not say that du uses Nemotron or this ASR workflow.

The useful point is narrower than “Arabic support”: a multilingual speech model can improve on a specific dialect when its training and test data represent that dialect. The remaining question for any operator is whether the model works on its own recordings, at the latency and error rate its service can tolerate.

Is word error rate the same as transcription accuracy?

No. Word error rate counts substitutions, deletions and insertions against a reference transcript, then divides by the reference word count. It is not simply the inverse of accuracy, and insertions can make the score exceed 100%.

Would a 29.96% word error rate be good enough for customer support?

The score alone cannot answer that. Word error rate treats every word equally, while a missed negation, account number or customer name may carry more risk than a harmless filler word. A support team should test its own calls and review errors that affect the task, not rely on a universal cutoff.

What does a team need before fine-tuning a model for another dialect?

It needs accurately transcribed audio from the target speakers and recording conditions, plus a separate held-out set for evaluation. The data must represent the dialect and domain the system will face; filtering should remove unusable recordings without discarding difficult but valid speech.

When is an extra 800ms of transcription latency worth accepting?

It may be reasonable for batch jobs such as transcribing recorded meetings or archived calls, where a delay after recording is acceptable. A live voice assistant or real-time agent would need to test whether the accuracy gain justifies the slower response.