Researcher profile

Arkaitz Zubiaga

6 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Making the Most of Tweet-Inherent Features for Social Spam Detection on Twitter

    2015 · arXiv (Cornell University)

    Social spam produces a great amount of noise on social media services such as Twitter, which reduces the signal-to-noise ratio that both end users and data mining applications observe. Existing techniques on social spam detection …

  2. Hidden behind the obvious: misleading keywords and implicitly abusive language on social media

    2022 · arXiv (Cornell University)

    While social media offers freedom of self-expression, abusive language carry significant negative social impact. Driven by the importance of the issue, research in the automated detection of abusive language has witnessed growth and improvement. However, …

  3. AnnoBERT: Effectively Representing Multiple Annotators' Label Choices to Improve Hate Speech Detection

    2022 · arXiv (Cornell University)

    Supervised approaches generally rely on majority-based labels. However, it is hard to achieve high agreement among annotators in subjective tasks such as hate speech detection. Existing neural network models principally regard labels as categorical variables, …

  4. Hate Speech Detection and Reclaimed Language: Mitigating False Positives and Compounded Discrimination

    2024

    While minimising false negatives in hate speech classification remains an important goal in order to reduce discrimination and increase fairness for online communities, there is a growing need to produce models that are sensitive to …

  5. ALHD: A Large-Scale and Multigenre Benchmark Dataset for Arabic LLM-Generated Text Detection

    2025 · arXiv (Cornell University)

    We introduce ALHD, the first large-scale comprehensive Arabic dataset explicitly designed to distinguish between human- and LLM-generated texts. ALHD spans three genres (news, social media, reviews), covering both MSA and dialectal Arabic, and contains over …

  6. MultiClaimNet: A Massively Multilingual Dataset of Fact-Checked Claim Clusters

    2025 · arXiv (Cornell University)

    In the context of fact-checking, claims are often repeated across various platforms and in different languages, which can benefit from a process that reduces this redundancy. While retrieving previously fact-checked claims has been investigated as …