conference-paper Open access

Exploring Amharic Hate Speech Data Collection and Classification Approaches

Research footprint

At a glance

Citations
3
References
34
Comments
0
Paper overview

Abstract

In this paper, we present a study of efficient data selection and annotation strategies for Amharic hate speech.We also build various classification models and investigate the challenges of hate speech data selection, annotation, and classification for the Amharic language.From a total of over 18 million tweets in our Twitter corpus, 15.1k tweets are annotated by two independent native speakers, and a Cohen's kappa score of 0.48 is achieved.A third annotator, a curator, is also employed to decide on the final gold labels.We employ both classical machine learning and deep learning approaches, which include fine-tuning AmFLAIR and AmRoBERTa contextual embedding models.Among all the models, AmFLAIR achieves the best performance with an F1-score of 72%.We publicly release the annotation guidelines, keywords/lexicon entries, datasets, models, and associated scripts with a permissive license 1 .

Record transparency

Publication details

DOI
10.26615/978-954-452-092-2_006
OpenAlex
W4388608092
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.