A Systematic Review on Machine Translation for Low Resource Nigerian Languages

Authors

  • Tijani M. Abdulmusawir Department of Computer Science, School of Technology, Federal Polytechnic Idah, Kogi state, Nigeria
  • A. F. Donfack Kana Department of Computer Science, Faculty of Physical Sciences, Ahmadu Bello University, Zaria, Kaduna state, Nigeria
  • Amina H. Abubakar Department of Computer Science, Faculty of Physical Sciences, Ahmadu Bello University, Zaria, Kaduna state, Nigeria

DOI:

https://doi.org/10.25299/itjrd.2025.21277

Keywords:

Artificial Inteligence, Domain Adaption, Machine Translation, Low-Resource

Abstract

Nigeria ranks among Africa's most linguistically diverse countries with over 500 indigenous languages, yet machine translation (MT) research remains severely limited for these low-resource languages. This systematic review examines the current state of MT research for Nigerian languages, identifies persistent challenges, and analyzes methodological trends. A systematic literature search was conducted across eleven databases including PubMed, Web of Science, and Scopus from January 2010 to August 2025. Search terms combined machine translation approaches with Nigerian language terms. Studies were screened using PRISMA guidelines requiring original research with evaluation metrics. From 51 papers, 25 duplicates were removed, 7 excluded for selection criteria, and 3 for lack of contribution, resulting in 16 studies. Only 11 Nigerian languages (2.2% of over 500 languages) were covered, creating a 97.8% research gap. Yoruba led with 4 studies, followed by Igala (3), Igbo and Nigerian Pidgin (2 each). Methods evolved from rule-based (4 studies, 2014 to 2021) through SMT (2 studies, 2016 to 2019) to NMT dominance (10 studies, 2018 to 2025). Idiomatic expression handling was the most persistent challenge (16.7%), followed by complex sentences, data scarcity, and domain specificity (each 9.5%). Nigerian MT research shows severe underrepresentation with persistent challenges in idiomatic expressions and data scarcity across all approaches. Neural method adoption reflects global trends but doesn't address resource constraints. Coordinated national approaches prioritizing parallel corpora creation and institutional partnerships are needed to prevent digital divides and support language preservation.

Downloads

Download data is not yet available.

References

[1] S. F. Ayegba, O. E. Osuagwu, and N. D. Okechukwu, (2014) “Machine translation of noun phrases from English to igala using the rule-based approach” West African Journal of Industrial and Academic Research, vol. 11, pp. 18-28

[2] D. Bahdanau, K. Cho, and Y. Bengio, “Neural machine translation by jointly learning to align and translate.” In Proceedings of the 3rd International Conference on Learning Representations (ICLR). San Diego, CA. 2015; https://arxiv.org/abs/1409.0473

[3] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need.” In Proceedings of the 31st International Conference on Neural Information Processing Systems, 2017, pp. 5998-600, https://doi.org/10.5555/3295222.3295349

[4] X. Wang, Y. Tsvetkov, G. Neubig “Balancing Training for Multilingual Neural Machine Translation” Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 2020, pp. 8526-8537. DOI: 10.18653/v1/2020.acl-main.754F. Stahlberg, “Neural machine translation: A review”. Journal of Artificial Intelligence Research, vol. 70, pp.1-30. 2020. https://doi.org/10.1613/jair.1.12007

[5] D. S. Doris, “Languages spoken in Nigeria as of 2021, by number of speakers” statista. Available P. Koehn, “Statistical Machine Translation” Cambridge University Press, ISBN-13 978-0-511-69132-4. 2010.

[6] W. Yonghui, M. Schuster, Z. Chen, Q. V. Le, and N. Mohammad, “Google’s Neural Machine Translation System: Bridging the Gap between Human and Machine Translation.” In: arXiv:1609.08144v2 [cs.CL], pp 1–23, Oct 2016. https://www.statista.com/statistics/1285383/population-in-nigeria-by-languages-spoken.2024

[7] F. Stahlberg “Neural machine translation: A review. Journal of Artificial Intelligence Research, 70, 1-30.2020. https://doi.org/10.1613/jair.1.12007

[8] O. I. Akinwale, A. O. Adetunmbi, O. O. Obe, A. T. Adesuyi. “ Web-Based English to Yoruba Machine Translation”. International Journal of Language and Linguistics. Vol. 3, No. 3, 2015, pp. 154-159. doi: 10.11648/j.ijll.20150303.17

[9] S. I. Eludiora, and B. A. Ajibade. “Design and Implementation of English To Yorùbá Verb Phrase Machine Translation System.” Koozakar Festschrift, 2021. Available: https:/api.semanticscholar.org/CorpusID:233204633.

[10] T. M. Abdulmusawir, S. F. Ayegba, Y. M. Kayode, and E. C. Christian, “A system for machine translation from English to Ebira using the rule-based approach” Journal of Scientific Research and Reports, 2021; vol. 27 no. 11, pp. 137-148. https://doi.org/10.9734/jsrr/2021/v27i1130465

[11] S. F. Ayegba, “Development of English-to-Igala machine translation system using neural machine translation with attention mechanism,” International Journal of Research and Development. Vol . 8, no. 1, pp. 1-10, 2023.

[12] B. U. Umar, “Nupe-English Neural Machine Translation Using Sequence to Sequence Model With Attention Mechanism.”Natural Language Processing Journal.”Available:https://www.researchgate.net/publication/38227542, 2024.

[13] E. Makoji; F. Sani. “Development and Evaluation of an English-to Igala Neural Machine Translation System using Deep Learning.” International Journal of Innovative Science and Research Technology, 10(5), 914-919. 2025 https://doi.org/10.38124/ijisrt/25may556

[14] I. Ezeani, P. Rayson, I. Onyenwe, C. Uchechukwu, M. Hepple, “Automatic Restoration of Diacritics for Igbo Language”. Transactions of the Association for Computational Linguistics. https://eprints.whiterose.ac.uk/id/eprint/117833/1/TSD_2016_Ezeani.pdf

[15] I. E. Onyenwe., M. Hepple., U. Chinedu, & E. Barnard. “Toward an Effective Igbo Part-of-Speech Tagger”. ACM Transactions on Asian and Low-Resource Language Information Processing (TALLIP). https://aclanthology.org/2019.nsurl-1.18.pdf

[16] T. Nguyen, & D. Chiang, “Improving Rare Word Translation with Dictionaries and Attention Masking.” 2018. arXiv preprint. https://arxiv.org/pdf/2408.09075.pdf

[17] O. E. Ahia, and K. E. Ogueji, “Toward Supervised and Unsupervised Neural Machine translation Baseline for Nigerian Pidgin” arXiv:2003.12660v1[cs.CL]. 2020 DOI: https://doi.org/10.48550/arXiv.2003. 12660

[18] A. Gutkin, I. Demirashin, O. Kjartansson, C. E. Rivera, K. Tubosun , “ Developing an Open-Source Corpus of Yoruba Speech” Proceedings Interspeech , International Speech Communication Association. pp. 404-408, 2020.

[19] I. Orife, “Neural machine translation for Edoid languages.” Journal of Intelligent Information Systems, vol. 56. no. 2, pp. 299-315, 2020.

[20] D. I. Adelani, D. Ruiter, J. O. Alabi, D. Adebonojo, A. Ayeni, M. Adeyemi, A. Awokoya, and C. Espana-Bonet, “MENYO 20K: A multilingual parallel corpus for under-resourced languages.” In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. pp. 3450-3463, 2021.

[21] Butryna, A., et al “Google Crowdsourced Speech Corpora and Related Open-Source Resources for Low-Resource Languages and Dialect.” Proceedings of the 12th Language Resources and Evaluation Conference, pp.3424-3433, 2020. https://doi.org/10.48550/arXiv.2010.06778

[22] M. A. Hedderich, L. Lange, H. Adel, J. Strötgen, and D. Klakow. “A Survey on Recent Approaches for Natural Language Processing in Low-Resource Scenarios.” In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp 2545–2568, Online. Association for Computational Linguistics.

[23] W. Nekoto. Participatory Research for Low-resourced Machine Translation: A Case Study in African Languages. 2020. arXiv:2010.02353v2, https://doi.org/10.48550/arXiv.2010.02353.

[24] S. Abate, M. Woldeyohannis, M. Tachbelie, M. Meshesha, S. Atnafu, W. Gewe, Y. Assabie, H. Abera, B. Seyoum, T. Abebe, W. Tsegaye, A. Lemma, T. Andargie, S. Shifaw (2018) Parallel corpora for bi-lingual English–Ethiopian languages statistical machine translation. In: Proceedings of the 27th international conference on computational linguistics, Santa Fe, New Mexico, USA, pp 3102–3111

[25] J. T. Sefara, V. Marivate, S. G. Zwane, N. Gama, H. Sibisi, P. N. Senoamadi. Transformer-based Machine Translation for Low-resourced Languages embedded with Language Identification.2021 Conference on Information Communications Technology and Society. IEEE, 2021. DOI: 10.1109/ICTAS50802.2021.9394996

[26] Y. Moukafih, N. Sbihi, M. Ghogho, K. Smaili. (2022). Improving Machine Translation of Arabic Dialects Through Multi-task Learning. In: Bandini, S., Gasparini, F., Mascardi, V., Palmonari, M., Vizzari, G. (eds) AIxIA 2021 – Advances in Artificial Intelligence. AIxIA 2021. Lecture Notes in Computer Science(), vol 13196. Springer, Cham. https://doi.org/10.1007/978-3-031-08421-8_40

Downloads

Published

2025-11-25

How to Cite

M. Abdulmusawir, T., Kana, A. F. D., & Abubakar, A. H. (2025). A Systematic Review on Machine Translation for Low Resource Nigerian Languages. IT Journal Research and Development, 10(2), 46–55. https://doi.org/10.25299/itjrd.2025.21277

Issue

Section

Review Article