Shamsuddeen Hassan Muhammad
5 أوراق في مجموعة PaperMetrix
أوراق هذا المؤلف
-
A Few Thousand Translations Go a Long Way! Leveraging Pre-trained Models for African News Translation
2022 · Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies
David Adelani, Jesujoba Alabi, Angela Fan, Julia Kreutzer, Xiaoyu Shen, Machel Reid, Dana Ruiter, Dietrich Klakow, Peter Nabende, Ernie Chang, Tajuddeen Gwadabe, Freshia Sackey, Bonaventure F. P. Dossou, Chris Emezue, Colin Leong, Michael Beukman, Shamsuddeen …
-
AfriHate: A Multilingual Collection of Hate Speech and Abusive Language Datasets for African Languages
2025
Shamsuddeen Hassan Muhammad, Idris Abdulmumin, Abinew Ali Ayele, David Ifeoluwa Adelani, Ibrahim Said Ahmad, Saminu Mohammad Aliyu, Paul Röttger, Abigail Oppong, Andiswa Bukula, Chiamaka Ijeoma Chukwuneke, Ebrahim Chekol Jibril, Elyas Abdi Ismail, Esubalew Alemneh, Hagos …
-
The State of Large Language Models for African Languages: Progress and Challenges
2025 · arXiv (Cornell University)
Large Language Models (LLMs) are transforming Natural Language Processing (NLP), but their benefits are largely absent for Africa's 2,000 low-resource languages. This paper comparatively analyzes African language coverage across six LLMs, eight Small Language Models …
-
Quality at a Glance: An Audit of Web-Crawled Multilingual Datasets
2022 · Transactions of the Association for Computational Linguistics
Abstract With the success of large-scale pre-training and multilingual modeling in Natural Language Processing (NLP), recent years have seen a proliferation of large, Web-mined text datasets covering hundreds of languages. We manually audit the quality …
-
MasakhaNER: Named Entity Recognition for African Languages
2021 · Transactions of the Association for Computational Linguistics
Abstract We take a step towards addressing the under- representation of the African continent in NLP research by bringing together different stakeholders to create the first large, publicly available, high-quality dataset for named entity recognition …