Hinglish HateBERT: Hate Speech Classification in Hinglish language
At a glance
- Citations
- 0
- References
- 13
- Comments
- 0
Abstract
Abstract Hate speech has become a widespread concern in the contemporary world, surpassing geographical confines and impacting communities on a worldwide level. The role of digital platforms in disseminating hate speech, the obstacles in regulating online hate speech, and the additional complexity due to language diversity collectively demand a comprehensive method to detect and identify hate speech from free speech. In this work, we introduce Hinglish HateBERT, a pre-trained BERT model for Hate speech detection in code-mixed Hindi English language. The model was trained on a large-scale dataset having offensive and non-offensive content to avoid any bias. We finetuned Hinglish HateBERT using CNN and LSTM and performed experimentation on three publicly available dataset. Further, we presented a detailed comparative performance of our model with publicly available pretrained models for classifying hate speech in Hinglish. We observed that our proposed model Hinglish HateBERT significantly outperformed for two datasets. Also, we have released our model and code for research purpose at : Author can release it on request of reviewer or after revision.
Publication details
- DOI
- 10.21203/rs.3.rs-3436709/v1
- OpenAlex
- W4387700450
- Document type
- preprint
- Language
- EN
- Source
- Research Square
- Last metadata update
Comments
Log in to join the discussion.