NVIDIA Announces Fine-Tuning of Nemotron 3.5 ASR for Saudi Arabic Dialects
NVIDIA Nemotron 3.5 ASR supports multilingual streaming transcription and is now fine-tuned for Saudi Arabic dialects, including Najdi and Hijazi.
NVIDIA announced the fine-tuning of their Nemotron 3.5 ASR model for Saudi Arabic dialects, specifically Najdi and Hijazi. This fine-tuning process uses the NVIDIA NeMo framework and reduces the word error rate (WER) from 55.05% to 29.96% on the target test split.
Key facts
| Fact | Detail | The source says |
|---|---|---|
| Word Error Rate (WER) reduction | from 55.05% to 29.96% | “Fine-tuning on 133.7 hours of Najdi and Hijazi speech reduced word error rate from 55.05% to 29.96% on the target test split” |
| English WER improvement | from 11.04% to 10.42% | “improved English performance from 11.04% to 10.42% WER without degrading other Arabic dialects” |
| Trainable parameters | 230.4 million | “Updating all 24 encoder layers achieved the lowest error rates at the cost of 230.4 million trainable parameters” |
What happened
NVIDIA has announced the fine-tuning of their Nemotron 3.5 ASR model for Saudi Arabic dialects, specifically Najdi and Hijazi. The fine-tuning process uses the NVIDIA NeMo framework and reduces the word error rate (WER) from 55.05% to 29.96% on the target test split. This workflow involves curating a low-resource corpus, building a weighted replay mix, fine-tuning with efficient batching, and evaluating transcription quality on an independent set. The process also improves English performance from 11.04% to 10.42% WER without degrading other Arabic dialects. Updating all 24 encoder layers achieved the lowest error rates but at the cost of 230.4 million trainable parameters. Freezing layers reduces compute needs when data or memory is limited.
Source: NVIDIA
