{"slug": "fine-tuning-nvidia-nemotron-for-saudi-arabic-dialects-with-a-path-to-other", "title": "Fine-Tuning NVIDIA Nemotron for Saudi Arabic Dialects, with a Path to Other Languages", "summary": "NVIDIA published a fine-tuning workflow for its Nemotron 3.5 ASR streaming 0.6B model that adapts the 40-language-locale speech recognizer to Saudi Najdi and Hijazi dialects using the NeMo framework, the SADA 2022 and FLEURS datasets, and a 12,000-step baseline run on two NVIDIA RTX PRO 6000 Blackwell Workstation Edition GPUs. The recipe combines minimal corpus curation, weighted replay mixing to reduce catastrophic forgetting, partial encoder unfreezing, and length bucketing, and NVIDIA notes the approach does not generalize into evidence for every Arabic dialect or deployment environment.", "body_md": "Automatic speech recognition must handle how people actually speak, not only the languages and styles that dominate pretraining data. Regional dialects and local recording conditions are often underrepresented, so a multilingual model that performs well on broad benchmarks may still fall short in deployment.\n\nSaudi Arabic makes that concrete. A model may recognize Modern Standard Arabic or English yet struggle with Najdi and Hijazi speech, or local recording conditions. Fine-tuning only on the target dialect can improve it while weakening other languages.\n\n[NVIDIA Nemotron 3.5 ASR](https://huggingface.co/nvidia/nemotron-3.5-asr-streaming-0.6b) supports multilingual streaming transcription across 40 language-locales, including transcription-ready Arabic, but deployment-specific dialects and recording conditions still benefit from fine-tuning.\n\nThis post shows how to adapt it with the [NVIDIA NeMo framework](https://github.com/NVIDIA/NeMo) and the [ASR fine-tuning recipe](https://github.com/nvidia-riva/tutorials/blob/main/asr-finetune-nemotron-3.5-asr-streaming-prompt.ipynb): curate a low-resource corpus, build a weighted replay mix, fine-tune with efficient batching, and evaluate transcription quality on an independent set.\n\n## When to use this workflow\n\nThis pipeline is useful when you have enough labeled speech to specialize an ASR model, but not enough to train one from scratch: dialect adaptation, domain-specific transcription, deployments that must retain existing languages.\n\nIts techniques solve different problems:\n\n- **Minimal curation** removes unusable labels and obvious alignment failures without discarding scarce, difficult speech.\n- **Replay mixing** interleaves a small amount of previously learned data to reduce catastrophic forgetting.\n- **Partial encoder unfreezing** limits how many parameters change — cheaper and faster than a full fine-tune, at some cost to accuracy.\n- **Length bucketing** reduces padding and makes training practical for streaming encoders.\n- **Beam search and a larger attention context** can improve offline accuracy without retraining, at the cost of latency and compute.\n\nThese are not universal defaults. Replay protects only what its data represents; partial unfreezing needs re-tuning when the mix changes. And this workflow doesn’t generalize into evidence for every Arabic dialect or deployment environment.\n\n## Fine-tuning walkthrough\n\n### Prerequisites\n\n- [NVIDIA NeMo](https://docs.nvidia.com/nemo-framework/user-guide/latest/nemotoolkit/asr/intro.html) , PyTorch, Python, OmegaConf\n- This tutorial uses [SADA 2022](https://www.kaggle.com/datasets/sdaiancai/sada2022) and[FLEURS](https://huggingface.co/datasets/google/fleurs)\n- Working knowledge of Python, model fine-tuning, WER and CER\n- Two GPUs were used for the 12,000-step baseline experiment; exact GPU model and memory: NVIDIA RTX PRO 6000 Blackwell Workstation Edition GPUs\n\n### 1. Curate the target corpus without filtering away the problem\n\nStart by selecting the dialects you intend to deploy. For the initial SADA experiment, that meant Najdi and Hijazi:\n\n```\nSAUDI_DIALECTS = {\"najdi\", \"hijazi\"}\nseen = set()\n\nfor item in manifest:\n    dialect = str(item.get(\"speaker_dialect\", \"\")).lower().strip()\n    seen.add(dialect)\n    if dialect not in SAUDI_DIALECTS:\n        continue\n    keep(item)\n\nmissing = SAUDI_DIALECTS - seen\nassert not missing, f\"never matched: {missing}; observed {sorted(seen)}\"\n```\n\nNext, remove references the model cannot learn and clips that are probably misaligned.\n\nSADA marks inaudible speech with غيرواضح; because the model cannot emit that annotation, every occurrence creates an unavoidable error.\nIn the code below, the markers appear as Unicode escapes so it renders left-to-right; they spell غيرواضح and غير واضح. The code applies a few other checks as well.\n\n```\nMIN_DURATION, MAX_DURATION   = 0.5, 30.0     # drop clips too short or too long\nMIN_CHAR_RATE, MAX_CHAR_RATE = 1.5, 35.0     # drop misaligned transcripts\n\nANNOTATION_MARKERS = ['\\u063a\\u064a\\u0631\\u0648\\u0627\\u0636\\u062d',\n                      '\\u063a\\u064a\\u0631 \\u0648\\u0627\\u0636\\u062d']\n\nfor row in manifest:\n    txt  = normalize_arabic(row['text'])\n    rate = len(txt) / row['duration']\n\n    if not txt or txt.lower() == 'nan':                      continue  # stringified NaN\n    if any(m in row['text'] for m in ANNOTATION_MARKERS):    continue  # annotation, not speech\n    if not (MIN_DURATION  <= row['duration'] <= MAX_DURATION):   continue\n    if not (MIN_CHAR_RATE <= rate <= MAX_CHAR_RATE):             continue\n\n    keep(row)\npython\ndef normalize_arabic(t):                    # target dialect + FLEURS Arabic\n    t = re.sub(r'[\\u064b-\\u0670]', '', t)                        # diacritics\n    t = re.sub(r'[\\u0623\\u0625\\u0622\\u0627]', '\\u0627', t)       # alef variants\n    t = t.replace('\\u0649', '\\u064a').replace('\\u0629', '\\u0647') # alef maqsura, taa marbuta\n    t = re.sub(r'[^\\u0600-\\u06ff\\w\\s]', '', t)                   # punctuation\n    return ' '.join(t.split())\n```\n\nThese checks retained 103,559 of 125,490 utterances: 133.7 hours, or 82.5% of the starting set. The goal is to remove structurally bad examples, not hard accents or noisy speech merely because the base model performs poorly on them. If you use automated quality scores, first inspect their distribution; the SADA run found that a default UTMOS threshold of 3.0 would have rejected almost everything.\n\nAs a second pass, the broader SADA curation pipeline used the [NVIDIA NeMo Curator](https://docs.nvidia.com/nemo/curator/curate-audio/process-data/quality-filtering) to standardize audio and remove severely degraded clips. `MonoConversionStage` converted inputs to mono, while `UTMOSFilterStage` and `SIGMOSFilterStage` scored perceived quality and background noise. Instead of applying their default thresholds, we first ran the stages in score-only mode, inspected the score distributions, then set corpus-specific cutoffs: UTMOS ≥ 1.25, SIGMOS noise ≥ 1.5, SIGMOS overall ≥ 1.5. They passed about 85% of duration-valid samples. This preserved challenging but usable dialect speech a generic threshold would have discarded.\n\n### 2. Start with one dataset and monitor model behavior\n\nFor a new language or domain, begin with the simplest experiment you can interpret: one representative target dataset, a conventional full fine-tune, and a fixed evaluation set. The goal is to observe how the model responds, verify that the training pipeline works, and create a baseline against which later changes can be measured.\n\nIn our case, SADA served as that first dataset. The model is Cache-Aware FastConformer-RNNT with prompted multilingual streaming (`strip_lang_tags`, `target_lang: ar-AR`), which behaves differently from plain English streaming models and requires explicit language conditioning in the manifest.\n\nYour starting corpus, training duration, and hardware settings will be different. The configuration changes below document how this particular experiment evolved as we encountered optimization, memory, and distributed-training issues; they are troubleshooting examples, not recommended defaults. Hyperparameters not listed, including weight decay, gradient clipping, mixed precision, and effective batch size in utterances, were left at NeMo framework defaults and weren’t tuned in this experiment.\n\n| Area | Configuration changes | \n|---|---|\n| Optimization | Learning rate: default → `1e-4` →`2e-5` ; warmup: 10,000 → 100 → 50 steps; optimizer: AdamW (NeMo default); scheduler: Noam;`d_model=1024` | \n| Data loading | Batch duration: 400 → 300 → 200 seconds to prevent out-of-memory errors; `is_tarred=false` ; workers: 4 training and 2 validation; bucket and shuffle buffers: 1,000 | \n| Validation | Batch size 8; validation every epoch | \n| Checkpoint and decoding | Retain the best three checkpoints; use greedy decoding | \n\n*Table 1. Configuration changes made during the SADA-only baseline experiment*\n\nIn the initial SADA-only experiment, we used the validation split to monitor fine-tuning progress. The pretrained model produced 49.5% WER on the validation split and 59% WER on the full training corpus. The gap reflects noisier audio and greater dialectal variation in the training data.\n\nThe validation baseline (49.5%) served as our primary metric throughout. WER improved to 47.8% after the first 10-epoch run, remained at 47.8% after another 10 epochs, and reached 46.7% during the longer v4 continuation. After epoch 45, validation WER stopped improving. This modest gain and clear plateau suggested that continuing the same full fine-tuning setup was unlikely to deliver a substantial improvement.\n\n### 3. What worked: A narrower target, a replay stream, and bucketed batches\n\n#### A narrower target with minimal curation\n\nInstead of asking the model to improve on 11dialects at once, we trained only on Najdi and Hijazi, the two we intended to use and we used minimal curation (see number 1, above), which kept 103,559 of 125,490 utterances, or 82.5%:\n\n#### A replay stream\n\nFine-tuning on Saudi speech alone overwrites what the model learned in pretraining. The defence is replay: mix a small stream of previously-learned data back in so the model keeps being asked to do the old job while it learns the new one.\n\nWe used 10% FLEURS, split 7% English and 3% Arabic against 90% Saudi speech. Declare those proportions rather than concatenating the files because a sliding-window shuffle may not reach rows appended to the end of a large manifest until late in training, so a concatenated replay set effectively doesn’t exist for most of the run.\n\n``` python\nfrom omegaconf import OmegaConf\n\nmix = OmegaConf.create([\n    {\"type\": \"nemo\", \"manifest_filepath\": \"sada_train.jsonl\", \"weight\": 0.90},\n    {\"type\": \"nemo\", \"manifest_filepath\": \"fleurs_en.jsonl\",  \"weight\": 0.07},\n    {\"type\": \"nemo\", \"manifest_filepath\": \"fleurs_ar.jsonl\",  \"weight\": 0.03},\n])\nOmegaConf.save(mix, \"input_cfg.yaml\")\ncfg.train_ds.manifest_filepath = None\ncfg.train_ds.input_cfg = \"input_cfg.yaml\"\n```\n\nSeven percent of English was enough. FLEURS English hold with a slight improvement from 11.04% to 10.42% (see Table 2, below), while the model was specialising in Arabic dialect speech.\n\n#### Bucketed batches\n\nThe second change was less obvious and mattered just as much.\n\nDuration-based bucketing groups similar-length utterances into the same batch:\n\n```\ncfg.train_ds.use_bucketing        = True\ncfg.train_ds.num_buckets          = 30\ncfg.train_ds.batch_size           = None\ncfg.train_ds.batch_duration       = 400.0\ncfg.train_ds.quadratic_duration   = 15.0\n```\n\nNote that `num_buckets` alone does nothing;  `use_bucketing` defaults to `False`, so setting the bucket count without the flag is a silent no-op that looks configured.\n\n### The result\n\nThose three changes together, over 12,000 steps and about 4.5 hours on two GPUs:\n\n| Test split | Before | After | \n|---|---|---|\n| SADA Najdi + Hijazi WER | 55.05% | 29.96% | \n| SADA Najdi + Hijazi CER | 31.63% | 12.18% | \n| Full SADA WER | 58.84% | 35.61% | \n| Full SADA CER | 35.40% | 15.97% | \n| FLEURS English WER | 11.04% | 10.42% | \n| FLEURS English CER | 6.47% | 4.53% | \n| FLEURS Arabic WER | 12.67% | 11.41% | \n| FLEURS Arabic CER | 5.55% | 3.97% | \n\n*Table 2. WER and CER before and after fine-tuning on the target dialects, the full SADA test set, and English retention checks*\n\nSpecializing didn’t cost us the dialects we dropped. The model improved by 25 points where we aimed it and 23 points across everything, and English improved at the same time. That looked like a finished result. All numbers were measured with [NeMo evaluation script](https://github.com/NVIDIA/NeMo/blob/main/examples/asr/speech_to_text_eval.py).\n\n### 4. Adaptation depth: How much of the encoder to update\n\nThe model has 24 encoder layers. A full fine-tune updates all of them; freezing the encoder preserves it but restricts acoustic adaptation. In between, you can unfreeze the top N layers and leave the rest alone, while always training the decoder, joint network and prompt embeddings.\n\n```\nfor parameter in model.parameters():\n    parameter.requires_grad = False\n\nfor name, parameter in model.named_parameters():\n    if any(part in name for part in (\"decoder\", \"joint\", \"prompt\")):\n        parameter.requires_grad = True\n\nfor layer in model.encoder.layers[-8:]:\n    for parameter in layer.parameters():\n        parameter.requires_grad = True\n\ntrainable = sum(p.numel() for p in model.parameters() if p.requires_grad)\ntotal = sum(p.numel() for p in model.parameters())\nprint(f\"trainable: {trainable/1e6:.1f}M / {total/1e6:.1f}M\")\n```\n\nIn the recorded top-eight recipe, 230.4 million parameters were trainable and 407.6 million frozen. We tested two depths against each other, changing nothing else:\n\n| Encoder layers updated | SADA WER | SADA CER | \n|---|---|---|\n| pretrained baseline | 55.05% | 31.63% | \n| top 6 | 33.42% | 14.10% | \n| top 8 | 32.32% | 13.53% | \n| all 24 | 29.96% | 12.18% | \n\n*Table 3. WER and CER by adaptation depth, Najdi and Hijazi test split*\n\nMore trainable capacity was better at every step. With 134 hours of target speech, the data supports updating the whole encoder, and partial freezing costs 2.4 points against the full fine-tune.\n\nThat is a finding about this data volume, not a general rule. Freezing is a way of spending less so it becomes the right choice when your corpus is thinner than ours or memory is the binding constraint.\n\n### 5. Improving without retraining: Larger context and beam search decoding\n\nTraining is only half the system. Before paying for another training run, check what you can get at inference time.\n\nThis checkpoint exposes several attention context sizes, selectable without retraining. The second number is how many future frames the encoder may attend to before committing to an output: `[56, 3]` is the streaming default, `[56, 13]` the widest it supports. Switching to the highest-lookahead context reduced WER by **1.31 absolute points without retraining**. Its main trade-off is approximately 800 ms of additional latency.\n\n| Configuration | WER | CER | vs. Greedy | \n|---|---|---|---|\n| Greedy, `[56, 3]` | 29.96% | 12.18% | +0.00 | \n| Greedy, `[56, 13]` | 28.65% | 11.39% | -1.31 | \n| MALSD beam-4, `[56, 3]` | 28.81% | 11.48% | -1.15 | \n| MALSD beam-8, `[56, 3]` | 28.62% | 11.40% | -1.34 | \n| MALSD beam-8, `[56, 13]` | 27.25% | 10.63% | -2.71 | \n\n*Table 4. Decoding configurations on the fine-tuned checkpoint, SADA Najdi and Hijazi test split*\n\nThe primary trade-off is additional latency rather than retraining cost. Thirteen lookahead frames instead of three is roughly 800 ms more buffering before each output. For batch transcription, such as call archives, media, meeting recordings, that provides an accuracy gain without retraining, where the additional buffering latency is acceptable. For live captioning, it may be unsuitable. The right setting follows from the deployment, not from the accuracy number.\n\nBeam-search performance depended strongly on the decoder configuration. In a later sweep, NeMo’s batched `malsd_batch` strategy produced useful gains at modest cost: beam 4 reduced SADA WER from 29.96 % to 28.81 % at 0.59× the greedy runtime, while beam 8 reached 28.62% at 0.64×. Beam 4 therefore offered the best speed–accuracy balance in our test. We did not evaluate MAES or NGPU-LM fusion, so our conclusion is limited to MALSD.\n\nOne detail worth setting regardless of strategy: `strip_lang_tags=True`. Otherwise the locale tag is emitted as literal text and scored as an insertion on every utterance.\n\n| Attention Context | Chunk Size (Latency) | Use Case | \n|---|---|---|\n| `[56, 0]` | 80ms (Ultra-Low) | Ultra-low-latency Voice Agents | \n| `[56, 1]` | 160ms (Low) | Interactive Voice Agents, Conversational AI | \n| `[56, 3]` | 320ms (Balanced) | Conversational AI, Live Captioning | \n| `[56, 6]` | 560ms (Medium) | High accuracy with reasonable latency | \n| `[56, 13]` | 1.12s (High) | Highest accuracy with high latency | \n\n*Table 5. Attention context sizes with their output latency and typical use*\n\n## Applying the workflow to other languages\n\nThe loop stays the same: curate, mix replay, choose an adaptation depth, bucket by duration, evaluate on independent target and regression sets. What changes per language is the corpus metadata, transcription conventions, normalization, tokenizer coverage, script handling, and capability metrics.\n\nFor another language, start with a small audited manifest and test the base tokenizer on names, numerals, borrowed words, and mixed-script sentences. Replace the Arabic normalizer with one that preserves meaningful distinctions in the target script, and validate unicode normalization and encoding end to end.\n\nFor languages without whitespace word boundaries, WER may be misleading, Use a character-, token-, or morpheme-level measure instead, and state its limits. Replay data should represent the capabilities you need to retain, not just convenient high-resource speech, and should be real rather than synthetic wherever possible.\n\n## Extend the workflow with speaker diarization\n\nFine-tuning ASR improves what is transcribed; speaker diarization adds who spoke when. In multi-speaker audio, a diarization model identifies speech segments and assigns consistent speaker labels. Combining those labels and timestamps with the fine-tuned Nemotron 3.5 ASR output creates a speaker-attributed transcript that separates each participant’s words, useful for meetings, interviews, contact centers, classrooms, and other multi-speaker environments.\n\nThis workflow can also help speech-data providers prepare multi-speaker corpora by automatically generating speaker-turn timestamps and anonymous speaker labels for human review before the data is used for ASR fine-tuning or evaluation.\n\n[NVIDIA Nemotron 3 Diarization](https://huggingface.co/nvidia/Nemotron-3-Diarization-preview), just released, extends this workflow from dialect-aware transcription to speaker-attributed transcription for up to 8 speakers. It is an open model that can be added to any ASR system you’re already using. Diarization doesn’t transcribe speech itself; its speaker boundaries and labels are aligned with the ASR timestamps to produce the final structured transcript. Learn more about the architecture, benchmarks, and how to get started in the [Hugging Face blog](https://huggingface.co/blog/nvidia/nemotron-diarization).\n\n## Next steps\n\nThe conversation that prompted this tutorial raised a more important question than capability alone: How to make dialect adaptation useful, improving the target dialect, retaining existing skills, and spending training effort where it changes the deployed system. The experiments point to a practical answer. Curate lightly, mix replay by weight, add examples of the exact behavior you need, and update as much of the encoder as your data supports. Then optimize decoding, and evaluate each capability on an independent set.\n\n## Get started\n\n**Finetuning notebook:**[https://github.com/nvidia-riva/tutorials/blob/main/asr-finetune-nemotron-3.5-asr-streaming-prompt.ipynb](https://github.com/nvidia-riva/tutorials/blob/main/asr-finetune-nemotron-3.5-asr-streaming-prompt.ipynb)\n\n**Finetuning skill:**[https://github.com/NVIDIA/skills/tree/main/skills/nemotron-asr-finetune](https://github.com/NVIDIA/skills/tree/main/skills/nemotron-asr-finetune)", "url": "https://wpnews.pro/news/fine-tuning-nvidia-nemotron-for-saudi-arabic-dialects-with-a-path-to-other", "canonical_source": "https://developer.nvidia.com/blog/fine-tuning-nvidia-nemotron-for-saudi-arabic-dialects-with-a-path-to-other-languages/", "published_at": "2026-10-01 05:00:00+00:00", "updated_at": "2026-10-01 05:17:10.132681+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "natural-language-processing", "ai-research", "ai-tools"], "entities": ["NVIDIA", "NVIDIA Nemotron 3.5 ASR", "NVIDIA NeMo", "SADA 2022", "FLEURS", "NVIDIA RTX PRO 6000 Blackwell Workstation Edition", "PyTorch", "OmegaConf"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/fine-tuning-nvidia-nemotron-for-saudi-arabic-dialects-with-a-path-to-other", "markdown": "https://wpnews.pro/news/fine-tuning-nvidia-nemotron-for-saudi-arabic-dialects-with-a-path-to-other.md", "text": "https://wpnews.pro/news/fine-tuning-nvidia-nemotron-for-saudi-arabic-dialects-with-a-path-to-other.txt", "jsonld": "https://wpnews.pro/news/fine-tuning-nvidia-nemotron-for-saudi-arabic-dialects-with-a-path-to-other.jsonld"}}