Examining the Research Gap in Swahili-to-Nigerian Language Machine Translation Using Transformer Models

📖 ABSTRACT/OVERVIEW

Machine translation between African languages has received limited research attention, and translation between Swahili, East Africa's most widely spoken language, and major Nigerian languages, particularly Hausa, Igbo, and Yoruba, has essentially no published NLP research despite potential value for Pan-African communication and knowledge exchange. This study addresses this research gap by developing parallel corpora and evaluating transformer-based machine translation for Swahili-to-Hausa and Swahili-to-Yoruba language pairs. Parallel sentence pairs were compiled from multilingual Bible corpora (2,800 sentences per pair), translated United Nations documents, and new community-sourced translations (1,200 sentence pairs per pair). Helsinki-NLP's MarianMT, fine-tuned for each language pair, and a custom transformer model with subword tokenisation adapted for the morphologically complex target languages were evaluated. Models were benchmarked using BLEU, chrF, and TER scores on held-out test sets. MarianMT fine-tuned on Swahili-Hausa achieved a BLEU score of 18.4, while the Swahili-Yoruba pair achieved 15.2, reflecting the greater morphological divergence between Swahili and Yoruba. Back-translation augmentation improved BLEU scores by an average of 2.8 points across both pairs. The study constitutes the first published machine translation benchmarks for Swahili-to-Nigerian language pairs and contributes parallel corpora publicly to the Masakhane NLP community repository, establishing a foundation for future African cross-language NLP research.

Keywords: machine translation, Swahili, Hausa, Yoruba, African NLP

Need Complete Chapters of the Above Topic?

Get high-quality, Zero-AI research materials with current citations.

Request via WhatsApp 💬
Departments# Computer Science