cd /news/artificial-intelligence/fine-tuning-nvidia-nemotron-for-saud… · home › topics › artificial-intelligence › article
[ARTICLE · art-143024] src=developer.nvidia.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Fine-Tuning NVIDIA Nemotron for Saudi Arabic Dialects, with a Path to Other Languages

NVIDIA published a fine-tuning workflow for its Nemotron 3.5 ASR streaming 0.6B model that adapts the 40-language-locale speech recognizer to Saudi Najdi and Hijazi dialects using the NeMo framework, the SADA 2022 and FLEURS datasets, and a 12,000-step baseline run on two NVIDIA RTX PRO 6000 Blackwell Workstation Edition GPUs. The recipe combines minimal corpus curation, weighted replay mixing to reduce catastrophic forgetting, partial encoder unfreezing, and length bucketing, and NVIDIA notes the approach does not generalize into evidence for every Arabic dialect or deployment environment.

by read13 min views1 publishedOct 1, 2026
Fine-Tuning NVIDIA Nemotron for Saudi Arabic Dialects, with a Path to Other Languages
Image: NVIDIA Developer Blog

Automatic speech recognition must handle how people actually speak, not only the languages and styles that dominate pretraining data. Regional dialects and local recording conditions are often underrepresented, so a multilingual model that performs well on broad benchmarks may still fall short in deployment.

Saudi Arabic makes that concrete. A model may recognize Modern Standard Arabic or English yet struggle with Najdi and Hijazi speech, or local recording conditions. Fine-tuning only on the target dialect can improve it while weakening other languages.

NVIDIA Nemotron 3.5 ASR supports multilingual streaming transcription across 40 language-locales, including transcription-ready Arabic, but deployment-specific dialects and recording conditions still benefit from fine-tuning.

This post shows how to adapt it with the NVIDIA NeMo framework and the ASR fine-tuning recipe: curate a low-resource corpus, build a weighted replay mix, fine-tune with efficient batching, and evaluate transcription quality on an independent set.

When to use this workflow #

This pipeline is useful when you have enough labeled speech to specialize an ASR model, but not enough to train one from scratch: dialect adaptation, domain-specific transcription, deployments that must retain existing languages.

Its techniques solve different problems:

  • Minimal curation removes unusable labels and obvious alignment failures without discarding scarce, difficult speech.
  • Replay mixing interleaves a small amount of previously learned data to reduce catastrophic forgetting.
  • Partial encoder unfreezing limits how many parameters change — cheaper and faster than a full fine-tune, at some cost to accuracy.
  • Length bucketing reduces padding and makes training practical for streaming encoders.
  • Beam search and a larger attention context can improve offline accuracy without retraining, at the cost of latency and compute.

These are not universal defaults. Replay protects only what its data represents; partial unfreezing needs re-tuning when the mix changes. And this workflow doesn’t generalize into evidence for every Arabic dialect or deployment environment.

Fine-tuning walkthrough #

Prerequisites

  • NVIDIA NeMo , PyTorch, Python, OmegaConf
  • This tutorial uses SADA 2022 andFLEURS
  • Working knowledge of Python, model fine-tuning, WER and CER
  • Two GPUs were used for the 12,000-step baseline experiment; exact GPU model and memory: NVIDIA RTX PRO 6000 Blackwell Workstation Edition GPUs

1. Curate the target corpus without filtering away the problem

Start by selecting the dialects you intend to deploy. For the initial SADA experiment, that meant Najdi and Hijazi:

SAUDI_DIALECTS = {"najdi", "hijazi"}
seen = set()

for item in manifest:
    dialect = str(item.get("speaker_dialect", "")).lower().strip()
    seen.add(dialect)
    if dialect not in SAUDI_DIALECTS:
        continue
    keep(item)

missing = SAUDI_DIALECTS - seen
assert not missing, f"never matched: {missing}; observed {sorted(seen)}"

Next, remove references the model cannot learn and clips that are probably misaligned.

SADA marks inaudible speech with غيرواضح; because the model cannot emit that annotation, every occurrence creates an unavoidable error. In the code below, the markers appear as Unicode escapes so it renders left-to-right; they spell غيرواضح and غير واضح. The code applies a few other checks as well.

MIN_DURATION, MAX_DURATION   = 0.5, 30.0     # drop clips too short or too long
MIN_CHAR_RATE, MAX_CHAR_RATE = 1.5, 35.0     # drop misaligned transcripts

ANNOTATION_MARKERS = ['\u063a\u064a\u0631\u0648\u0627\u0636\u062d',
                      '\u063a\u064a\u0631 \u0648\u0627\u0636\u062d']

for row in manifest:
    txt  = normalize_arabic(row['text'])
    rate = len(txt) / row['duration']

    if not txt or txt.lower() == 'nan':                      continue  # stringified NaN
    if any(m in row['text'] for m in ANNOTATION_MARKERS):    continue  # annotation, not speech
    if not (MIN_DURATION  <= row['duration'] <= MAX_DURATION):   continue
    if not (MIN_CHAR_RATE <= rate <= MAX_CHAR_RATE):             continue

    keep(row)
python
def normalize_arabic(t):                    # target dialect + FLEURS Arabic
    t = re.sub(r'[\u064b-\u0670]', '', t)                        # diacritics
    t = re.sub(r'[\u0623\u0625\u0622\u0627]', '\u0627', t)       # alef variants
    t = t.replace('\u0649', '\u064a').replace('\u0629', '\u0647') # alef maqsura, taa marbuta
    t = re.sub(r'[^\u0600-\u06ff\w\s]', '', t)                   # punctuation
    return ' '.join(t.split())

These checks retained 103,559 of 125,490 utterances: 133.7 hours, or 82.5% of the starting set. The goal is to remove structurally bad examples, not hard accents or noisy speech merely because the base model performs poorly on them. If you use automated quality scores, first inspect their distribution; the SADA run found that a default UTMOS threshold of 3.0 would have rejected almost everything.

As a second pass, the broader SADA curation pipeline used the NVIDIA NeMo Curator to standardize audio and remove severely degraded clips. MonoConversionStage converted inputs to mono, while UTMOSFilterStage and SIGMOSFilterStage scored perceived quality and background noise. Instead of applying their default thresholds, we first ran the stages in score-only mode, inspected the score distributions, then set corpus-specific cutoffs: UTMOS ≥ 1.25, SIGMOS noise ≥ 1.5, SIGMOS overall ≥ 1.5. They passed about 85% of duration-valid samples. This preserved challenging but usable dialect speech a generic threshold would have discarded.

2. Start with one dataset and monitor model behavior

For a new language or domain, begin with the simplest experiment you can interpret: one representative target dataset, a conventional full fine-tune, and a fixed evaluation set. The goal is to observe how the model responds, verify that the training pipeline works, and create a baseline against which later changes can be measured.

In our case, SADA served as that first dataset. The model is Cache-Aware FastConformer-RNNT with prompted multilingual streaming (strip_lang_tags, target_lang: ar-AR), which behaves differently from plain English streaming models and requires explicit language conditioning in the manifest.

Your starting corpus, training duration, and hardware settings will be different. The configuration changes below document how this particular experiment evolved as we encountered optimization, memory, and distributed-training issues; they are troubleshooting examples, not recommended defaults. Hyperparameters not listed, including weight decay, gradient clipping, mixed precision, and effective batch size in utterances, were left at NeMo framework defaults and weren’t tuned in this experiment.

Area Configuration changes
Optimization Learning rate: default → 1e-4 →2e-5 ; warmup: 10,000 → 100 → 50 steps; optimizer: AdamW (NeMo default); scheduler: Noam;d_model=1024
Data Batch duration: 400 → 300 → 200 seconds to prevent out-of-memory errors; is_tarred=false ; workers: 4 training and 2 validation; bucket and shuffle buffers: 1,000
Validation Batch size 8; validation every epoch
Checkpoint and decoding Retain the best three checkpoints; use greedy decoding

Table 1. Configuration changes made during the SADA-only baseline experiment

In the initial SADA-only experiment, we used the validation split to monitor fine-tuning progress. The pretrained model produced 49.5% WER on the validation split and 59% WER on the full training corpus. The gap reflects noisier audio and greater dialectal variation in the training data.

The validation baseline (49.5%) served as our primary metric throughout. WER improved to 47.8% after the first 10-epoch run, remained at 47.8% after another 10 epochs, and reached 46.7% during the longer v4 continuation. After epoch 45, validation WER stopped improving. This modest gain and clear plateau suggested that continuing the same full fine-tuning setup was unlikely to deliver a substantial improvement.

3. What worked: A narrower target, a replay stream, and bucketed batches

A narrower target with minimal curation

Instead of asking the model to improve on 11dialects at once, we trained only on Najdi and Hijazi, the two we intended to use and we used minimal curation (see number 1, above), which kept 103,559 of 125,490 utterances, or 82.5%:

A replay stream

Fine-tuning on Saudi speech alone overwrites what the model learned in pretraining. The defence is replay: mix a small stream of previously-learned data back in so the model keeps being asked to do the old job while it learns the new one.

We used 10% FLEURS, split 7% English and 3% Arabic against 90% Saudi speech. Declare those proportions rather than concatenating the files because a sliding-window shuffle may not reach rows appended to the end of a large manifest until late in training, so a concatenated replay set effectively doesn’t exist for most of the run.

from omegaconf import OmegaConf

mix = OmegaConf.create([
    {"type": "nemo", "manifest_filepath": "sada_train.jsonl", "weight": 0.90},
    {"type": "nemo", "manifest_filepath": "fleurs_en.jsonl",  "weight": 0.07},
    {"type": "nemo", "manifest_filepath": "fleurs_ar.jsonl",  "weight": 0.03},
])
OmegaConf.save(mix, "input_cfg.yaml")
cfg.train_ds.manifest_filepath = None
cfg.train_ds.input_cfg = "input_cfg.yaml"

Seven percent of English was enough. FLEURS English hold with a slight improvement from 11.04% to 10.42% (see Table 2, below), while the model was specialising in Arabic dialect speech.

Bucketed batches

The second change was less obvious and mattered just as much.

Duration-based bucketing groups similar-length utterances into the same batch:

cfg.train_ds.use_bucketing        = True
cfg.train_ds.num_buckets          = 30
cfg.train_ds.batch_size           = None
cfg.train_ds.batch_duration       = 400.0
cfg.train_ds.quadratic_duration   = 15.0

Note that num_buckets alone does nothing; use_bucketing defaults to False, so setting the bucket count without the flag is a silent no-op that looks configured.

The result

Those three changes together, over 12,000 steps and about 4.5 hours on two GPUs:

Test split Before After
SADA Najdi + Hijazi WER 55.05% 29.96%
SADA Najdi + Hijazi CER 31.63% 12.18%
Full SADA WER 58.84% 35.61%
Full SADA CER 35.40% 15.97%
FLEURS English WER 11.04% 10.42%
FLEURS English CER 6.47% 4.53%
FLEURS Arabic WER 12.67% 11.41%
FLEURS Arabic CER 5.55% 3.97%

Table 2. WER and CER before and after fine-tuning on the target dialects, the full SADA test set, and English retention checks

Specializing didn’t cost us the dialects we dropped. The model improved by 25 points where we aimed it and 23 points across everything, and English improved at the same time. That looked like a finished result. All numbers were measured with NeMo evaluation script.

4. Adaptation depth: How much of the encoder to update

The model has 24 encoder layers. A full fine-tune updates all of them; freezing the encoder preserves it but restricts acoustic adaptation. In between, you can unfreeze the top N layers and leave the rest alone, while always training the decoder, joint network and prompt embeddings.

for parameter in model.parameters():
    parameter.requires_grad = False

for name, parameter in model.named_parameters():
    if any(part in name for part in ("decoder", "joint", "prompt")):
        parameter.requires_grad = True

for layer in model.encoder.layers[-8:]:
    for parameter in layer.parameters():
        parameter.requires_grad = True

trainable = sum(p.numel() for p in model.parameters() if p.requires_grad)
total = sum(p.numel() for p in model.parameters())
print(f"trainable: {trainable/1e6:.1f}M / {total/1e6:.1f}M")

In the recorded top-eight recipe, 230.4 million parameters were trainable and 407.6 million frozen. We tested two depths against each other, changing nothing else:

Encoder layers updated SADA WER SADA CER
pretrained baseline 55.05% 31.63%
top 6 33.42% 14.10%
top 8 32.32% 13.53%
all 24 29.96% 12.18%

Table 3. WER and CER by adaptation depth, Najdi and Hijazi test split

More trainable capacity was better at every step. With 134 hours of target speech, the data supports updating the whole encoder, and partial freezing costs 2.4 points against the full fine-tune.

That is a finding about this data volume, not a general rule. Freezing is a way of spending less so it becomes the right choice when your corpus is thinner than ours or memory is the binding constraint.

5. Improving without retraining: Larger context and beam search decoding

Training is only half the system. Before paying for another training run, check what you can get at inference time.

This checkpoint exposes several attention context sizes, selectable without retraining. The second number is how many future frames the encoder may attend to before committing to an output: [56, 3] is the streaming default, [56, 13] the widest it supports. Switching to the highest-lookahead context reduced WER by 1.31 absolute points without retraining. Its main trade-off is approximately 800 ms of additional latency.

Configuration WER CER vs. Greedy
Greedy, [56, 3] 29.96% 12.18% +0.00
Greedy, [56, 13] 28.65% 11.39% -1.31
MALSD beam-4, [56, 3] 28.81% 11.48% -1.15
MALSD beam-8, [56, 3] 28.62% 11.40% -1.34
MALSD beam-8, [56, 13] 27.25% 10.63% -2.71

Table 4. Decoding configurations on the fine-tuned checkpoint, SADA Najdi and Hijazi test split

The primary trade-off is additional latency rather than retraining cost. Thirteen lookahead frames instead of three is roughly 800 ms more buffering before each output. For batch transcription, such as call archives, media, meeting recordings, that provides an accuracy gain without retraining, where the additional buffering latency is acceptable. For live captioning, it may be unsuitable. The right setting follows from the deployment, not from the accuracy number.

Beam-search performance depended strongly on the decoder configuration. In a later sweep, NeMo’s batched malsd_batch strategy produced useful gains at modest cost: beam 4 reduced SADA WER from 29.96 % to 28.81 % at 0.59× the greedy runtime, while beam 8 reached 28.62% at 0.64×. Beam 4 therefore offered the best speed–accuracy balance in our test. We did not evaluate MAES or NGPU-LM fusion, so our conclusion is limited to MALSD.

One detail worth setting regardless of strategy: strip_lang_tags=True. Otherwise the locale tag is emitted as literal text and scored as an insertion on every utterance.

Attention Context Chunk Size (Latency) Use Case
[56, 0] 80ms (Ultra-Low) Ultra-low-latency Voice Agents
[56, 1] 160ms (Low) Interactive Voice Agents, Conversational AI
[56, 3] 320ms (Balanced) Conversational AI, Live Captioning
[56, 6] 560ms (Medium) High accuracy with reasonable latency
[56, 13] 1.12s (High) Highest accuracy with high latency

Table 5. Attention context sizes with their output latency and typical use

Applying the workflow to other languages #

The loop stays the same: curate, mix replay, choose an adaptation depth, bucket by duration, evaluate on independent target and regression sets. What changes per language is the corpus metadata, transcription conventions, normalization, tokenizer coverage, script handling, and capability metrics.

For another language, start with a small audited manifest and test the base tokenizer on names, numerals, borrowed words, and mixed-script sentences. Replace the Arabic normalizer with one that preserves meaningful distinctions in the target script, and validate unicode normalization and encoding end to end.

For languages without whitespace word boundaries, WER may be misleading, Use a character-, token-, or morpheme-level measure instead, and state its limits. Replay data should represent the capabilities you need to retain, not just convenient high-resource speech, and should be real rather than synthetic wherever possible.

Extend the workflow with speaker diarization #

Fine-tuning ASR improves what is transcribed; speaker diarization adds who spoke when. In multi-speaker audio, a diarization model identifies speech segments and assigns consistent speaker labels. Combining those labels and timestamps with the fine-tuned Nemotron 3.5 ASR output creates a speaker-attributed transcript that separates each participant’s words, useful for meetings, interviews, contact centers, classrooms, and other multi-speaker environments.

This workflow can also help speech-data providers prepare multi-speaker corpora by automatically generating speaker-turn timestamps and anonymous speaker labels for human review before the data is used for ASR fine-tuning or evaluation.

NVIDIA Nemotron 3 Diarization, just released, extends this workflow from dialect-aware transcription to speaker-attributed transcription for up to 8 speakers. It is an open model that can be added to any ASR system you’re already using. Diarization doesn’t transcribe speech itself; its speaker boundaries and labels are aligned with the ASR timestamps to produce the final structured transcript. Learn more about the architecture, benchmarks, and how to get started in the Hugging Face blog.

Next steps #

The conversation that prompted this tutorial raised a more important question than capability alone: How to make dialect adaptation useful, improving the target dialect, retaining existing skills, and spending training effort where it changes the deployed system. The experiments point to a practical answer. Curate lightly, mix replay by weight, add examples of the exact behavior you need, and update as much of the encoder as your data supports. Then optimize decoding, and evaluate each capability on an independent set.

Get started #

Finetuning notebook:https://github.com/nvidia-riva/tutorials/blob/main/asr-finetune-nemotron-3.5-asr-streaming-prompt.ipynb

Finetuning skill:https://github.com/NVIDIA/skills/tree/main/skills/nemotron-asr-finetune

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @nvidia 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/fine-tuning-nvidia-n…] indexed:0 read:13min 2026-10-01 · —