conference-paper Open access

TOXIC COMMENTS DETECTION IN RUSSIAN

  • Computational Linguistics and Intellectual Technologies
Research footprint

At a glance

Citations
17
References
59
Comments
0
Paper overview

Abstract

Currently, social network sites tend to be one of the major communication platforms in both offline and online space. Freedom of expression of various points of view, including toxic, aggressive, and abusive comments, might have a long-term negative impact on people’s opinions and social cohesion. As a consequence, the ability to automatically identify and moderate toxic content on the Internet to eliminate the negative consequences is one of the necessary tasks for modern society. This paper aims at the automatic detection of toxic comments in the Russian language. As a source of data, we utilized anonymously published Kaggle dataset and additionally validated its annotation quality. To build a classification model, we performed fine-tuning of two versions of Multilingual Universal Sentence Encoder, Bidirectional Encoder Representations from Transformers, and ruBERT. Finetuned RuBERT achieved F1 = 92.20%, demonstrating the best classification score. We made trained models and code samples publicly available to the research community.

Record transparency

Publication details

DOI
10.28995/2075-7182-2020-19-1149-1159
OpenAlex
W3101323333
Document type
conference-paper
Language
EN
Source
Computational Linguistics and Intellectual Technologies
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.