📖 ABSTRACT/OVERVIEW
Phishing detection systems are predominantly trained on English-language datasets, and their performance on email content including Nigerian Pidgin, Yoruba, Igbo, and Hausa elements represents a significant unaddressed performance gap. This study empirically evaluated five machine learning phishing detection models (Naive Bayes, SVM, Random Forest, LSTM, and BERT-based) on a Nigerian multilingual email dataset. A dataset of 4,200 emails (2,100 phishing, 2,100 legitimate) was compiled from institutional email logs, VirusTotal submissions, and researcher-collected phishing campaigns. The dataset contained emails with English (62 percent), Nigerian Pidgin (18 percent), Yoruba (8 percent), Igbo (7 percent), and Hausa (5 percent) elements. Models were trained on the ENRON benchmark dataset and evaluated on the Nigerian dataset. Results showed significant performance degradation across all models on the Nigerian dataset. BERT achieved the highest F1 on the Nigerian dataset (0.81) but showed 0.14 F1 reduction from benchmark performance. Performance on Hausa and Igbo sub-corpora was lowest across all models (F1 below 0.70). The study provides original empirical evidence for the Nigerian language phishing detection gap and recommends development of a Nigerian multilingual phishing dataset as a community resource, and fine-tuning of transformer models on Nigerian language phishing samples to improve detection accuracy for locally targeted attacks.
Keywords: phishing detection, machine learning, Nigerian language, multilingual email, detection performance
Need Complete Chapters of the Above Topic?
Get high-quality, Zero-AI research materials with current citations.
Request via WhatsApp 💬