Researcher profile

Preslav Nakov

19 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Feature-Rich Named Entity Recognition for Bulgarian Using Conditional Random Fields

    2021 · arXiv (Cornell University)

    The paper presents a feature-rich approach to the automatic recognition and categorization of named entities (persons, organizations, locations, and miscellaneous) in news text for Bulgarian. We combine well-established features used for other languages with language-specific …

  2. WERD: Using social text spelling variants for evaluating dialectal speech recognition

    2017 · 2017 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU)

    We study the problem of evaluating automatic speech recognition (ASR) systems that target dialectal speech input. A major challenge in this case is that the orthography of dialects is typically not standardized. From an ASR …

  3. Building Chatbots from Forum Data: Model Selection Using Question Answering Metrics

    2017 · arXiv (Cornell University)

    We propose to use question answering (QA) data from Web forums to train chatbots from scratch, i.e., without dialog training data. First, we extract pairs of question and answer sentences from the typically much longer …

  4. Finding People's Professions and Nationalities Using Distant Supervision - The FMI@SU "goosefoot" team at the WSDM Cup 2017 Triple Scoring Task

    2017 · arXiv (Cornell University)

    We describe the system that our FMI@SU student's team built for participating in the Triple Scoring task at the WSDM Cup 2017. Given a triple from a "type-like" relation, profession or nationality, the goal is …

  5. Language Identification and Morphosyntactic Tagging: The Second VarDial Evaluation Campaign

    2018 · Työväentutkimus Vuosikirja

    We present the results and the findings of the Second VarDial Evaluation Campaign on Natural Language Processing (NLP) for Similar Languages, Varieties and Dialects. The campaign was organized as part of the fifth edition of …

  6. Adversarial Domain Adaptation for Duplicate Question Detection

    2018 · arXiv (Cornell University)

    We address the problem of detecting duplicate questions in forums, which is an important step towards automating the process of answering new questions. As finding and annotating such potential duplicates manually is very tedious and …

  7. SemEval-2010 Task 8: Multi-Way Classification of Semantic Relations Between Pairs of Nominals

    2019 · arXiv (Cornell University)

    SemEval-2 Task 8 focuses on Multi-way classification of semantic relations between pairs of nominals. The task was designed to compare different approaches to semantic relation classification and to provide a standard testbed for future research. …

  8. A Large-Scale Semi-Supervised Dataset for Offensive Language Identification

    2020 · arXiv (Cornell University)

    The use of offensive language is a major problem in social media which has led to an abundance of research in detecting content such as hate speech, cyberbulling, and cyber-aggression. There have been several attempts …

  9. SOLID: A Large-Scale Semi-Supervised Dataset for Offensive Language Identification

    2021

    The widespread use of offensive content in social media has led to an abundance of research in detecting language such as hate speech, cyberbullying, and cyber-aggression. Recent work presented the OLID dataset, which follows a …

  10. Detecting and Understanding Harmful Memes: A Survey

    2022

    The automatic identification of harmful content online is of major concern for social media platforms, policymakers, and society. Researchers have studied textual, visual, and audio content, but typically in isolation. Yet, harmful content often combines …

  11. Few-Shot Cross-Lingual Stance Detection with Sentiment-Based Pre-Training

    2021 · arXiv (Cornell University)

    The goal of stance detection is to determine the viewpoint expressed in a piece of text towards a target. These viewpoints or contexts are often expressed in many different languages depending on the user and …

  12. Feature-Rich Part-of-speech Tagging for Morphologically Complex\n Languages: Application to Bulgarian

    2019 · arXiv (Cornell University)

    We present experiments with part-of-speech tagging for Bulgarian, a Slavic\nlanguage with rich inflectional and derivational morphology. Unlike most\nprevious work, which has used a small number of grammatical categories, we work\nwith 680 morpho-syntactic tags. We combine …

  13. Fact-Checking the Output of Large Language Models via Token-Level Uncertainty Quantification

    2024 · arXiv (Cornell University)

    Large language models (LLMs) are notorious for hallucinating, i.e., producing erroneous claims in their output. Such hallucinations can be dangerous, as occasional factual inaccuracies in the generated text might be obscured by the rest of …

  14. Benchmarking Uncertainty Quantification Methods for Large Language Models with LM-Polygraph

    2024 · arXiv (Cornell University)

    The rapid proliferation of large language models (LLMs) has stimulated researchers to seek effective and efficient approaches to deal with LLM hallucinations and low-quality outputs. Uncertainty quantification (UQ) is a key element of machine learning …

  15. Toxicity Red-Teaming: Benchmarking LLM Safety in Singapore's Low-Resource Languages

    2025 · arXiv (Cornell University)

    The advancement of Large Language Models (LLMs) has transformed natural language processing; however, their safety mechanisms remain under-explored in low-resource, multilingual settings. Here, we aim to bridge this gap. In particular, we introduce \textsf{SGToxicGuard}, a …

  16. SemEval-2016 Task 3: Community Question Answering

    2016

    Preslav Nakov, Lluís Màrquez, Alessandro Moschitti, Walid Magdy, Hamdy Mubarak, Abed Alhakim Freihat, Jim Glass, Bilal Randeree. Proceedings of the 10th International Workshop on Semantic Evaluation (SemEval-2016). 2016.

  17. Overview of the DSL Shared Task 2015

    2015

    We present the results of the 2nd edition of the Discriminating between Similar Lan-guages (DSL) shared task, which was or-ganized as part of the LT4VarDial’2015 workshop and focused on the identifica-tion of very similar languages …

  18. Discriminating between Similar Languages and Arabic Dialect Identification: A Report on the Third DSL Shared Task

    2016 · International Conference on Computational Linguistics

    We present the results of the third edition of the Discriminating between Similar Languages (DSL) shared task, which was organized as part of the VarDial’2016 workshop at COLING’2016. The challenge offered two subtasks: subtask 1 …

  19. SemEval-2017 Task 3: Community Question Answering

    2017

    Preslav Nakov, Doris Hoogeveen, Lluís Màrquez, Alessandro Moschitti, Hamdy Mubarak, Timothy Baldwin, Karin Verspoor. Proceedings of the 11th International Workshop on Semantic Evaluation (SemEval-2017). 2017.