conference-paper

Mitigating Online Harassment: Machine Learning Approaches for Hate Speech Detection in Transliterated Bengali Comments

Research footprint

At a glance

Citations
3
References
14
Comments
0
Paper overview

Abstract

In the era of widespread online communication, the detection of hate speech has become increasingly critical for maintaining a healthy digital discourse. This significance is magnified when considering languages with unique characteristics, such as transliterated Bengali, where challenges in distinguishing hate speech abound. This work undertakes the task of exploring machine learning algorithms to tackle this challenge and contribute to the broader effort of fostering a respectful and inclusive online environment. The study introduces a novel dataset for hate speech detection in transliterated Bengali text, employing two distinct data preprocessing approaches—TF-IDF and Bag of Words. Eight diverse machine learning algorithms are then applied to evaluate their performance under each preprocessing technique. The results showcase the efficacy of specific algorithms, with Multinomial Naive Bayes excelling in the binary dataset and Logistic Regression emerging as a top performer in the multiclass dataset. Despite encountering challenges like imbalanced data and word length distribution, our models demonstrate enhanced precision and recall. This work not only contributes valuable insights to the field but also provides a new dataset, paving the way for future advancements in hate speech detection and model robustness.

Record transparency

Publication details

DOI
10.1109/iccit60459.2023.10441244
OpenAlex
W4392188733
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.