📖 ABSTRACT/OVERVIEW
Automatic speech recognition systems for Nigerian languages are severely limited by small labelled audio datasets, and data augmentation techniques that artificially expand training corpora offer a potential solution whose effectiveness requires empirical evaluation. This study empirically tested the effectiveness of four data augmentation strategies (speed perturbation, pitch shifting, room impulse response simulation, and SpecAugment) for improving Hausa language speech recognition models trained on the Common Voice Mozilla Hausa dataset comprising 9.2 hours of validated audio. A wav2vec 2.0 architecture was fine-tuned under each augmentation condition and evaluated on a held-out Kano and Katsina regional dialect test set. Word error rates were used as the primary evaluation metric. Baseline (no augmentation) achieved a WER of 28.4 percent. SpecAugment alone reduced WER to 23.1 percent. Combined speed perturbation and SpecAugment achieved the lowest WER of 21.7 percent. Room impulse response simulation increased WER slightly, suggesting mismatch with the clean recording conditions of the test set. Augmentation benefits were larger for female speakers than male speakers due to initial training set gender imbalance. Cross-dialect generalisation improved by 4.2 percentage points under the best augmentation condition. The study fills an empirical gap in data augmentation effectiveness for Nigerian speech recognition and recommends combined speed perturbation and SpecAugment as the standard augmentation protocol, along with urgent expansion of the Common Voice Hausa corpus through community contribution campaigns.
Keywords: speech recognition, data augmentation, Hausa language, wav2vec 2.0, low-resource NLP
Need Complete Chapters of the Above Topic?
Get high-quality, Zero-AI research materials with current citations.
Request via WhatsApp 💬