Senja Pollak
4 papers in the PaperMetrix corpus
Papers by this author
-
Investigating cross-lingual training for offensive language detection
2021 · PeerJ Computer Science
Platforms that feature user-generated content (social media, online forums, newspaper comment sections etc.) have to detect and filter offensive speech within large, fast-changing datasets. While many automatic methods have been proposed and achieve good accuracies, …
-
The Recent Advances in Automatic Term Extraction: A survey
2023 · arXiv (Cornell University)
Automatic term extraction (ATE) is a Natural Language Processing (NLP) task that eases the effort of manually identifying terms from domain-specific corpora by providing a list of candidate terms. As units of knowledge in a …
-
Mono- and cross-lingual evaluation of representation language models on less-resourced languages
2025 · Computer Speech & Language
The current dominance of large language models in natural language processing is based on their contextual awareness. For text classification, text representation models, such as ELMo, BERT, and BERT derivatives, are typically fine-tuned for a …