Deep Learning Architectures for Yoruba and Hausa Natural Language Processing: Addressing Low-Resource Language Challenges

📖 ABSTRACT/OVERVIEW

The two most widely spoken indigenous Nigerian languages, Yoruba and Hausa, remain severely underrepresented in natural language processing research and in commercial AI language technologies, perpetuating a form of digital linguistic exclusion that limits Nigerian citizens' participation in the benefits of AI-powered communication systems. This study develops, trains, and evaluates novel deep learning architectures specifically optimised for Yoruba and Hausa natural language processing tasks under the data scarcity conditions characteristic of low-resource language environments. The research makes three primary technical contributions. First, an original Tonal Feature Augmentation Layer architecture is developed for Yoruba, explicitly encoding the phonemic tone distinctions that standard transformer tokenisation frameworks systematically degrade. Second, a Cross-Lingual Transfer Learning Protocol is designed and validated for bootstrapping Yoruba and Hausa model performance from related high-resource languages including Swahili, Arabic, and French. Third, a Yoruba-Hausa Parallel Corpus of 185,000 sentence pairs is constructed through collaboration with linguistics researchers at Obafemi Awolowo University and Bayero University Kano, constituting the largest publicly available resource of its kind. Models are evaluated on sentiment analysis, named entity recognition, and machine translation benchmarks. The Tonal Feature Augmentation architecture achieves state-of-the-art results on Yoruba sentiment analysis, with an F1 score of 0.79, a 12-point improvement over the previous published benchmark. Keywords: natural language processing, Yoruba, Hausa, deep learning, low-resource languages.

Need Complete Chapters of the Above Topic?

Get high-quality, Zero-AI research materials with current citations.

Request via WhatsApp 💬