conference-paper

Leveraging BERT Models for Toxicity Detection in Moroccan Dialect

Research footprint

At a glance

Citations
0
References
44
Comments
0
Paper overview

Abstract

The widespread use of social media and increased internet accessibility have led to the generation of vast amounts of data. Much of the online textual content consists of informal expressions of opinions by individuals, which often include hateful language. The advent of transformer-based models has greatly enhanced addressing such challenges in textual data. With their ability to process context bidirectionally and handle large-scale data efficiently, transformers have revolutionized natural language processing, particularly in tasks such as classifying online comments. From this perspective, this study has been conducted to deploy transformer-based models to identify harmful statements written in the Moroccan dialect (Darija), which poses unique challenges due to its informal and diverse linguistic nature. The Moroccan Arabic Toxic Comments Dataset (MATCD), gathered using Selenium, serves as the foundation for this study, enabling the evaluation of transformer-based models in detecting harmful comments in the Moroccan dialect. As the interest in the BERT (Bidirectional Encoder Representations from Transformers) model increased, several model variations were developed to adjust to different languages, including Arabic, to enhance the ability to handle language-specific tasks effectively. The experimental results exhibit the strengths and weaknesses of each approach, with the model achieving high accuracy, reaching up to 95%, providing valuable insight into the suitability and effectiveness of transformer-based models in this domain.

Record transparency

Publication details

DOI
10.1109/iraset64571.2025.11008323
OpenAlex
W4410738987
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.