conference-paper وصول مفتوح

Socially Responsible and Explainable Automated Fact-Checking and Hate Speech Detection

Research footprint

At a glance

الاستشهادات
0
المراجع
29
Comments
0
Paper overview

Abstract

Misinformation and hate speech form a socially harmful cycle. Research shows that misinformation can amplify hate speech targeting social identity groups and reinforce harmful stereotypes. To combat this cycle, a wide range of Natural Language Processing (NLP) methods have been proposed. Nevertheless, while NLP has historically relied on inherently explainable “white-box” techniques, such as rule-based algorithms, decision trees, hidden markov models, and logistic regression, the adoption of Large Language Models (LLMs) and language embeddings (often considered “black-box”) has significantly reduced interpretability. This lack of transparency introduces considerable risks, including biases, which have become a major concern in AI. This Ph.D. thesis addresses these critical gaps by proposing new resources that ensure explainability and bias mitigation in NLP models for these tasks. Specifically, it introduces five benchmark datasets (HateBR, HateBRXplain, HausaHate, MOL, and FactNews), three novel methods (SELFAR, SSA, and B+M), and one web system (NoHateBrazil) designed to improve the explainability and fairness of automated fact-checking and hate speech detection. The proposed models outperform existing baselines for Portuguese and Hausa, both underrepresented languages. This research contributes to ongoing discussions on responsible and explainable AI, bridging the gap between model performance and interpretability for realworld applications. Finally, it has had a significant impact both nationally and internationally, receiving citations from prestigious universities and research institutes abroad, and inspiring new M.Sc. and Ph.D. projects in Brazil.

Record transparency

Publication details

DOI
10.5753/ctd.2025.8511
OpenAlex
W4413774299
Document type
conference-paper
Language
EN
Last metadata update
المجتمع

Comments

تسجيل الدخول للانضمام إلى النقاش.

  1. لا توجد تعليقات بعد. ابدأ النقاش.