Seid Muhie Yimam
8 papers in the PaperMetrix corpus
Papers by this author
-
Narrowing the Loop: Integration of Resources and Linguistic Dataset Development with Interactive Machine Learning
2015
This thesis proposal sheds light on the role of interactive machine learning and implicit user feedback for manual annotation tasks and se-mantic writing aid applications. First we fo-cus on the cost-effective annotation of train-ing data …
-
Word Complexity is in the Eye of the Beholder
2021
Sian Gooding, Ekaterina Kochmar, Seid Muhie Yimam, Chris Biemann. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
-
Exploring Amharic Hate Speech Data Collection and Classification Approaches
2023
In this paper, we present a study of efficient data selection and annotation strategies for Amharic hate speech.We also build various classification models and investigate the challenges of hate speech data selection, annotation, and classification …
-
EthioLLM: Multilingual Large Language Models for Ethiopian Languages with Task Evaluation
2024 · arXiv (Cornell University)
Large language models (LLMs) have gained popularity recently due to their outstanding performance in various downstream Natural Language Processing (NLP) tasks. However, low-resource languages are still lagging behind current state-of-the-art (SOTA) developments in the field …
-
AfriHate: A Multilingual Collection of Hate Speech and Abusive Language Datasets for African Languages
2025
Shamsuddeen Hassan Muhammad, Idris Abdulmumin, Abinew Ali Ayele, David Ifeoluwa Adelani, Ibrahim Said Ahmad, Saminu Mohammad Aliyu, Paul Röttger, Abigail Oppong, Andiswa Bukula, Chiamaka Ijeoma Chukwuneke, Ebrahim Chekol Jibril, Elyas Abdi Ismail, Esubalew Alemneh, Hagos …
-
The State of Large Language Models for African Languages: Progress and Challenges
2025 · arXiv (Cornell University)
Large Language Models (LLMs) are transforming Natural Language Processing (NLP), but their benefits are largely absent for Africa's 2,000 low-resource languages. This paper comparatively analyzes African language coverage across six LLMs, eight Small Language Models …
-
A Web-based Tool for the Integrated Annotation of Semantic and Syntactic Structures
2016 · International Conference on Computational Linguistics
We introduce the third major release of WebAnno, a generic web-based annotation tool for distributed teams. New features in this release focus on semantic annotation tasks (e.g. semantic role labelling or event annotation) and allow …
-
MasakhaNER: Named Entity Recognition for African Languages
2021 · Transactions of the Association for Computational Linguistics
Abstract We take a step towards addressing the under- representation of the African continent in NLP research by bringing together different stakeholders to create the first large, publicly available, high-quality dataset for named entity recognition …