Empirical Analysis of Transfer Learning Efficiency for Low-Resource Nigerian Language Automatic Speech Recognition

📖 ABSTRACT/OVERVIEW

Automatic speech recognition for Nigerian languages including Hausa, Yoruba, and Igbo remains severely limited due to the low availability of labelled speech training data, and transfer learning from multilingual pre-trained models offers a potential solution whose empirical efficiency has not been characterised for these specific language contexts. This study empirically assessed transfer learning approaches for developing Hausa and Yoruba ASR systems using minimal labelled data, addressing a documented gap in Nigerian language technology research. Wav2Vec 2.0 (base model trained on LibriSpeech 960h) and Whisper (medium multilingual model) were fine-tuned on Hausa and Yoruba speech datasets of varying sizes: 1-hour, 5-hour, and 20-hour subsets sampled from the Mozilla Common Voice dataset (Hausa: 26 hours total, Yoruba: 18 hours total). Word error rate was measured on identical held-out test sets for all fine-tuning conditions. Whisper fine-tuned on 5 hours of Hausa data achieved a WER of 34.2 percent, compared to Wav2Vec 2.0 at 41.7 percent under the same conditions. With 20 hours of training data, Whisper WER improved to 21.8 percent, approaching practical usability for simple command and query applications. Yoruba WER was consistently 8 to 12 percentage points higher than Hausa under equivalent conditions, attributed to tonal complexity. Performance on rural-accented speech degraded by 18 to 24 percent compared to urban speaker test sets, identifying accent diversity as a critical data collection priority. The study recommends a coordinated Nigerian language speech corpus collection initiative funded by TETFUND to enable higher-accuracy ASR development.

Need Complete Chapters of the Above Topic?

Get high-quality, Zero-AI research materials with current citations.

Request via WhatsApp 💬