📖 ABSTRACT/OVERVIEW
Nigerian government documents and administrative records span English, Nigerian Pidgin, Yoruba, Hausa, and Igbo, and developing original deep learning architectures for cross-lingual document understanding in this specific multi-language context addresses both an NLP research frontier and a practical e-government digitisation need. This study developed an original deep learning architecture for multi-language document understanding in Nigerian administrative contexts. A model development methodology was employed, combining systematic review of multilingual document AI literature (57 publications from 2019 to 2024), data compilation and annotation of a 28,000-document Nigerian administrative corpus including government notices, court judgements, and civic registration forms in five languages, and iterative architecture design and evaluation. The study identified two key architectural requirements not satisfied by existing multilingual models: handling code-switching between English and vernacular languages within single documents, and processing low-quality scanned document inputs from under-resourced government offices. An original architecture combining an adapted XLM-RoBERTa encoder with document layout attention modules and OCR noise correction layers was designed and trained. The model achieved 86.4 percent accuracy on a multi-label document classification task across five languages, outperforming XLM-RoBERTa baseline by 11.7 percentage points. Performance on code-switched documents improved by 18.2 percentage points over baseline. Expert review by 14 multilingual NLP and e-government specialists confirmed the architecture's original contribution. The study recommends deployment for digitisation of Nigeria's court records as a pilot application.
Keywords: multilingual NLP, document understanding, Nigerian languages, deep learning architecture, e-government
Need Complete Chapters of the Above Topic?
Get high-quality, Zero-AI research materials with current citations.
Request via WhatsApp 💬