Improving Hebrew offensive language classification using LLM-assisted human-in-the-loop annotation!!
At a glance
- Citations
- 0
- References
- 17
- Comments
- 0
Abstract
Abstract This study explores the challenges and opportunities of using large language models (LLMs) for automatic annotation of offensive language in Hebrew. The analysis is based on a six-level taxonomy of offensive discourse – covering offensiveness, target, target presence, vulgarity, offense strength, and specific aspects – enabling systematic examination of nuanced Hebrew patterns. Several prompting strategies were tested, including few-shot, role-based prompting, chain-of-thought, and LLM-as-judge, and their outputs were compared to human annotations using standard evaluation metrics and inter-annotator reliability. Findings reveal that salient categories such as explicit threats show strong interpretive stability, while ambiguous ones, such as discrediting attacks, require higher precision. The study also introduces a two-step classification method: first identifying the two most plausible categories, then selecting the more accurate one. This approach reduces the model’s bias toward general categories and improves fine-grained classification. Overall, the study contributes by (1) offering a methodological framework to assess LLMs’ interpretive limits compared to humans, and (2) laying groundwork for building more refined datasets to advance Hebrew offensive language research.
Publication details
- DOI
- 10.1515/lpp-2025-0093
- OpenAlex
- W7132820080
- Document type
- article
- Language
- EN
- Source
- Lodz Papers in Pragmatics
- Last metadata update
Comments
Log in to join the discussion.