📖 ABSTRACT/OVERVIEW
Natural language processing research has predominantly focused on high-resource languages, leaving the majority of Nigeria's over 500 languages severely underserved, with minimal corpora, pre-trained models, and task benchmarks available for research or application development. This study addresses this gap by developing and evaluating NLP datasets and baseline models for three low-resource Nigerian languages: Hausa, Yoruba, and Tiv. Corpora were compiled through web scraping, community crowdsourcing, and licensed media transcripts. The Hausa corpus comprised 180,000 sentences, Yoruba 120,000, and Tiv 45,000. Datasets were cleaned, tokenised, and annotated for three tasks: named entity recognition, sentiment classification, and news topic classification. Baseline models using fine-tuned multilingual BERT (mBERT) and AfriBERTa were trained and evaluated using train-validation-test splits. mBERT fine-tuned on Hausa NER achieved an F1 score of 0.74, compared to 0.68 for the multilingual model without fine-tuning. AfriBERTa outperformed mBERT on all Yoruba tasks by an average margin of 3.2 F1 points. Tiv model performance was limited by corpus size, with NER F1 reaching only 0.51. Data augmentation using back-translation improved Tiv sentiment classification by 4.1 F1 points. The study contributes the first benchmark datasets for Tiv NLP tasks and provides reproducible baseline models for all three languages, with datasets publicly released to facilitate further research.
Keywords: natural language processing, low-resource languages, Hausa, Yoruba, Tiv NLP
Need Complete Chapters of the Above Topic?
Get high-quality, Zero-AI research materials with current citations.
Request via WhatsApp 💬