📖 ABSTRACT/OVERVIEW
Deep learning natural language processing models have achieved state-of-the-art performance on major world languages but remain significantly underdeveloped for Nigerian indigenous languages, with Igbo and Yoruba representing the most important priority targets given their speaker populations. This study examined the research gap in deep learning NLP for Igbo and Yoruba through systematic literature review and original model experimentation. A systematic scoping review of ten databases from 2018 to 2024 identified 34 publications on NLP for Nigerian languages, of which 19 addressed Yoruba and 12 addressed Igbo. Pre-trained language model fine-tuning (mBERT, XLM-RoBERTa) was available for Yoruba but not Igbo. Text classification benchmarks existed for Yoruba sentiment but not for named entity recognition or question answering. Original experiments fine-tuned XLM-RoBERTa on a compiled 12,000-sentence Igbo news corpus assembled from publicly available sources. The fine-tuned model achieved 73.4 percent accuracy on a text classification task, significantly below comparable Yoruba (88.1 percent) and Swahili (91.2 percent) baselines. The performance gap was attributed primarily to training corpus size. The study identifies three priority research needs: large-scale curated Igbo and Yoruba text corpora, standardised benchmark evaluation tasks, and morphology-aware tokenisation for tonal language features. Recommendations include a TETFUND-funded Nigerian Indigenous Languages NLP Consortium and collaboration with Imo State University and Obafemi Awolowo University to develop annotated corpora.
Keywords: NLP, deep learning, Igbo, Yoruba, Nigerian languages
Need Complete Chapters of the Above Topic?
Get high-quality, Zero-AI research materials with current citations.
Request via WhatsApp 💬