Marcos Zampieri
12 ورقة في مجموعة PaperMetrix
أوراق هذا المؤلف
-
German Dialect Identification in Interview Transcriptions
2017
This paper presents three systems submitted to the German Dialect Identification (GDI) task at the VarDial Evaluation Campaign 2017. The task consists of training models to identify the dialect of Swiss-German speech transcripts. The dialects …
-
Classifier Ensembles for Dialect and Language Variety Identification
2018 · arXiv (Cornell University)
In this paper we present ensemble-based systems for dialect and language variety identification using the datasets made available by the organizers of the VarDial Evaluation Campaign 2018. We present a system developed to discriminate between …
-
Language Identification and Morphosyntactic Tagging: The Second VarDial Evaluation Campaign
2018 · Työväentutkimus Vuosikirja
We present the results and the findings of the Second VarDial Evaluation Campaign on Natural Language Processing (NLP) for Similar Languages, Varieties and Dialects. The campaign was organized as part of the fifth edition of …
-
LIdioms: A Multilingual Linked Idioms Data Set
2018
In this paper, we describe the LIDIOMS data set, a multilingual RDF representation of idioms currently containing five languages: English, German, Italian, Portuguese, and Russian.The data set is intended to support natural language processing applications …
-
A Large-Scale Semi-Supervised Dataset for Offensive Language Identification
2020 · arXiv (Cornell University)
The use of offensive language is a major problem in social media which has led to an abundance of research in detecting content such as hate speech, cyberbulling, and cyber-aggression. There have been several attempts …
-
SOLID: A Large-Scale Semi-Supervised Dataset for Offensive Language Identification
2021
The widespread use of offensive content in social media has led to an abundance of research in detecting language such as hate speech, cyberbullying, and cyber-aggression. Recent work presented the OLID dataset, which follows a …
-
Health Text Simplification: An Annotated Corpus for Digestive Cancer Education and Novel Strategies for Reinforcement Learning
2024 · arXiv (Cornell University)
Objective: The reading level of health educational materials significantly influences the understandability and accessibility of the information, particularly for minoritized populations. Many patient educational resources surpass the reading level and complexity of widely accepted standards. …
-
Deep Contrastive Active Learning for Out-of-domain Filtering in Dialog Systems
2024
Task-oriented dialog systems have shown to foster effective human-chatbot collaborations for accomplishing goal-specific tasks through intent classification. In a real-world setting, collecting and training over user intents incurs a labeling-cost challenge for human annotators. While …
-
Findings of the 2016 Conference on Machine Translation
2016
Ondřej Bojar, Rajen Chatterjee, Christian Federmann, Yvette Graham, Barry Haddow, Matthias Huck, Antonio Jimeno Yepes, Philipp Koehn, Varvara Logacheva, Christof Monz, Matteo Negri, Aurélie Névéol, Mariana Neves, Martin Popel, Matt Post, Raphael Rubino, Carolina Scarton, …
-
Overview of the DSL Shared Task 2015
2015
We present the results of the 2nd edition of the Discriminating between Similar Lan-guages (DSL) shared task, which was or-ganized as part of the LT4VarDial’2015 workshop and focused on the identifica-tion of very similar languages …
-
Discriminating between Similar Languages and Arabic Dialect Identification: A Report on the Third DSL Shared Task
2016 · International Conference on Computational Linguistics
We present the results of the third edition of the Discriminating between Similar Languages (DSL) shared task, which was organized as part of the VarDial’2016 workshop at COLING’2016. The challenge offered two subtasks: subtask 1 …
-
Findings of the 2019 Conference on Machine Translation (WMT19)
2019
Loïc Barrault, Ondřej Bojar, Marta R. Costa-jussà, Christian Federmann, Mark Fishel, Yvette Graham, Barry Haddow, Matthias Huck, Philipp Koehn, Shervin Malmasi, Christof Monz, Mathias Müller, Santanu Pal, Matt Post, Marcos Zampieri. Proceedings of the Fourth …