📖 ABSTRACT/OVERVIEW
Named entity recognition in Nigerian English text, which contains a distinctive mix of local personal names, place names, institutional names, and cultural terminology not well represented in standard NER training corpora, presents a documented but underresearched performance challenge. This study analyses the performance of state-of-the-art NER models on Nigerian English text and develops a Nigerian English NER benchmark dataset. A corpus of 35,000 sentences of Nigerian English was assembled from five domains: news articles, court judgments, government gazettes, social media, and company annual reports. Five entity categories were annotated: person names, organisations, locations, dates, and public offices. Inter-annotator agreement was 0.86 Kappa. Four NER systems were evaluated: spaCy (en_core_web_lg), BERT (fine-tuned on CoNLL-2003), a BERT model fine-tuned on the new Nigerian corpus, and GLiNER. Performance was measured using entity-level F1 score. Standard spaCy achieved an F1 of 0.71 on Nigerian text, compared to 0.91 on standard English benchmarks, confirming significant performance degradation. Fine-tuning BERT on the Nigerian corpus improved F1 to 0.87. GLiNER achieved 0.83 F1 without corpus-specific fine-tuning. Person name recognition was the most challenging category (F1 0.68 for standard models), attributable to the diversity of Nigerian naming conventions not represented in training data. The study contributes the first publicly available Nigerian English NER benchmark dataset and provides baselines for future research.
Keywords: named entity recognition, Nigerian English, BERT, NLP benchmark, text analytics
Need Complete Chapters of the Above Topic?
Get high-quality, Zero-AI research materials with current citations.
Request via WhatsApp 💬