conference-paper

Enhancing Hate Speech Detection in Mixed-Language Texts: A Comparative Study of BLOOM and XLM-RoBERTa Models

Research footprint

At a glance

Citations
3
References
6
Comments
0
Paper overview

Abstract

Hate speech detection is essential in combating online toxicity, particularly in mixed-language or code-switched texts prevalent on social media. Traditional natural language processing (NLP) models often struggle with these complex linguistic structures due to blending multiple languages. This paper investigates the effectiveness of BLOOM (BigScience Large Open-science Open-access Multilingual language model) and XLM-RoBERTa, two powerful multilingual models, in addressing these challenges. BLOOM's extensive pre-training across diverse languages and XLM-RoBERTa's robust capabilities allow a nuanced understanding of context in mixed-language environments. We fine-tune both models on an English and Indonesian text dataset containing instances of mixed-language hate speech and evaluate their performance against state-of-the-art benchmarks. Our findings highlight the effectiveness of these models in recognizing hate speech in mixed-language scenarios, with the fine-tuned BLOOM (bloom-560m) performing better than XLM-RoBERTa (xlm-roberta-base).

Record transparency

Publication details

DOI
10.1109/iccae64891.2025.10980554
OpenAlex
W4410228654
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.