conference-paper

An ensemble multilingual model for toxic comment classification

Research footprint

At a glance

Citations
4
References
11
Comments
0
Paper overview

Abstract

The online toxic comments cause enormous harm to the society, where toxicity is defined as anything rude, disrespectful or otherwise likely to make someone leave a discussion. To have a safer, more collaborative internet, grateful contributions are made by a main area of focus on machine learning models to identify toxicity in English, whereas part of misinformation disseminates in other languages. Over the past year, pretraining multilingual language models give rise to impressive gains for cross lingual toxicity classification. This paper presents an approach to build toxicity models applying the Jigsaw Multilingual Toxic Comment Classification dataset provided by Kaggle. We set our ensemble model in three parts based on Besides, we implement subsample, Pseudo-labeling with open-subtitles, translating non-English languages to English language, and Post Processing to improve the classification accuracy indispensably. Our final model achieved an AUC of 0.9469 for the training set and 0.9485 for the validation set, demonstrating the effectiveness of performance under cross-lingual toxicity detectors.

Record transparency

Publication details

DOI
10.1117/12.2636419
OpenAlex
W4229063414
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.