conference-paper

IndoBERT-Based Ensemble Learning for Multi-Level Multi-Label Hate Speech Detection in Indonesian Social Media

Research footprint

At a glance

Citations
2
References
14
Comments
0
Paper overview

Abstract

Hate speech on social media platforms has become a pressing issue, with harmful content often leading to social tensions and emotional harm. In Indonesia, the complex linguistic and cultural context of online discourse presents additional challenges for effective hate speech detection. This study addresses these challenges by presenting an ensemble learning approach for hate speech detection in Indonesian social media. Leveraging IndoBERT for language understanding and combining it with Bi-LSTM and Bi-GRU models for sequence processing, we developed a robust multi-model architecture that effectively captures linguistic patterns and contextual nuances unique to Indonesian. The proposed ensemble framework was tested on a comprehensive dataset with multiple hate speech labels, including categories such as Religion, Race, Gender, and Severity. Experimental results demonstrate that the ensemble model achieved an accuracy of 86% and an Fl-score of 63%, significantly outperforming individual models across most categories. This approach highlights the potential of ensemble learning for automated content moderation in Indonesian social media, providing a promising solution for managing diverse forms of online hate speech.

Record transparency

Publication details

DOI
10.1109/bts-i2c63534.2024.10942204
OpenAlex
W4409047033
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.