article Open access

Explaining Offensive Language Detection

  • LDV-Forum/Journal for language technology and computational linguistics
Research footprint

At a glance

Citations
7
References
46
Comments
0
Paper overview

Abstract

Machine learning approaches have proven to be on or even above human-level accuracy for the task of offensive language detection. In contrast to human experts, however, they often lack the capability of giving explanations for their decisions. This article compares four different approaches to make offensive language detection explainable: an interpretable machine learning model (naive Bayes), a model-agnostic explainability method (LIME), a model-based explainability method (LRP), and a self-explanatory model (LSTM with an attention mechanism). Three different classification methods: SVM, naive Bayes, and LSTM are paired with appropriate explanation methods. To this end, we investigate the trade-off between classification performance and explainability of the respective classifiers. We conclude that, with the appropriate explanation methods, the superior classification performance of more complex models is worth the initial lack of explainability.

Record transparency

Publication details

DOI
10.21248/jlcl.34.2020.223
OpenAlex
W4312760094
Document type
article
Language
EN
Source
LDV-Forum/Journal for language technology and computational linguistics
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.