ملف الباحث

Marcos Zampieri

12 ورقة في مجموعة PaperMetrix

المنشورات

أوراق هذا المؤلف

  1. German Dialect Identification in Interview Transcriptions

    2017

    This paper presents three systems submitted to the German Dialect Identification (GDI) task at the VarDial Evaluation Campaign 2017. The task consists of training models to identify the dialect of Swiss-German speech transcripts. The dialects …

  2. Classifier Ensembles for Dialect and Language Variety Identification

    2018 · arXiv (Cornell University)

    In this paper we present ensemble-based systems for dialect and language variety identification using the datasets made available by the organizers of the VarDial Evaluation Campaign 2018. We present a system developed to discriminate between …

  3. Language Identification and Morphosyntactic Tagging: The Second VarDial Evaluation Campaign

    2018 · Työväentutkimus Vuosikirja

    We present the results and the findings of the Second VarDial Evaluation Campaign on Natural Language Processing (NLP) for Similar Languages, Varieties and Dialects. The campaign was organized as part of the fifth edition of …

  4. LIdioms: A Multilingual Linked Idioms Data Set

    2018

    In this paper, we describe the LIDIOMS data set, a multilingual RDF representation of idioms currently containing five languages: English, German, Italian, Portuguese, and Russian.The data set is intended to support natural language processing applications …

  5. A Large-Scale Semi-Supervised Dataset for Offensive Language Identification

    2020 · arXiv (Cornell University)

    The use of offensive language is a major problem in social media which has led to an abundance of research in detecting content such as hate speech, cyberbulling, and cyber-aggression. There have been several attempts …

  6. SOLID: A Large-Scale Semi-Supervised Dataset for Offensive Language Identification

    2021

    The widespread use of offensive content in social media has led to an abundance of research in detecting language such as hate speech, cyberbullying, and cyber-aggression. Recent work presented the OLID dataset, which follows a …

  7. Health Text Simplification: An Annotated Corpus for Digestive Cancer Education and Novel Strategies for Reinforcement Learning

    2024 · arXiv (Cornell University)

    Objective: The reading level of health educational materials significantly influences the understandability and accessibility of the information, particularly for minoritized populations. Many patient educational resources surpass the reading level and complexity of widely accepted standards. …

  8. Deep Contrastive Active Learning for Out-of-domain Filtering in Dialog Systems

    2024

    Task-oriented dialog systems have shown to foster effective human-chatbot collaborations for accomplishing goal-specific tasks through intent classification. In a real-world setting, collecting and training over user intents incurs a labeling-cost challenge for human annotators. While …

  9. Findings of the 2016 Conference on Machine Translation

    2016

    Ondřej Bojar, Rajen Chatterjee, Christian Federmann, Yvette Graham, Barry Haddow, Matthias Huck, Antonio Jimeno Yepes, Philipp Koehn, Varvara Logacheva, Christof Monz, Matteo Negri, Aurélie Névéol, Mariana Neves, Martin Popel, Matt Post, Raphael Rubino, Carolina Scarton, …

  10. Overview of the DSL Shared Task 2015

    2015

    We present the results of the 2nd edition of the Discriminating between Similar Lan-guages (DSL) shared task, which was or-ganized as part of the LT4VarDial’2015 workshop and focused on the identifica-tion of very similar languages …

  11. Discriminating between Similar Languages and Arabic Dialect Identification: A Report on the Third DSL Shared Task

    2016 · International Conference on Computational Linguistics

    We present the results of the third edition of the Discriminating between Similar Languages (DSL) shared task, which was organized as part of the VarDial’2016 workshop at COLING’2016. The challenge offered two subtasks: subtask 1 …

  12. Findings of the 2019 Conference on Machine Translation (WMT19)

    2019

    Loïc Barrault, Ondřej Bojar, Marta R. Costa-jussà, Christian Federmann, Mark Fishel, Yvette Graham, Barry Haddow, Matthias Huck, Philipp Koehn, Shervin Malmasi, Christof Monz, Mathias Müller, Santanu Pal, Matt Post, Marcos Zampieri. Proceedings of the Fourth …