Shervin Malmasi
9 أوراق في مجموعة PaperMetrix
أوراق هذا المؤلف
-
German Dialect Identification in Interview Transcriptions
2017
This paper presents three systems submitted to the German Dialect Identification (GDI) task at the VarDial Evaluation Campaign 2017. The task consists of training models to identify the dialect of Swiss-German speech transcripts. The dialects …
-
Classifier Ensembles for Dialect and Language Variety Identification
2018 · arXiv (Cornell University)
In this paper we present ensemble-based systems for dialect and language variety identification using the datasets made available by the organizers of the VarDial Evaluation Campaign 2018. We present a system developed to discriminate between …
-
Language Identification and Morphosyntactic Tagging: The Second VarDial Evaluation Campaign
2018 · Työväentutkimus Vuosikirja
We present the results and the findings of the Second VarDial Evaluation Campaign on Natural Language Processing (NLP) for Similar Languages, Varieties and Dialects. The campaign was organized as part of the fifth edition of …
-
CycleNER: An Unsupervised Training Approach for Named Entity Recognition
2022 · Proceedings of the ACM Web Conference 2022
Named Entity Recognition (NER) is a crucial natural language understanding task for many down-stream tasks such as question answering and retrieval. Despite significant progress in developing NER models for multiple languages and domains, scaling to …
-
Reinforced Question Rewriting for Conversational Question Answering
2022 · arXiv (Cornell University)
Conversational Question Answering (CQA) aims to answer questions contained within dialogues, which are not easily interpretable without context. Developing a model to rewrite conversational questions into self-contained ones is an emerging solution in industry settings …
-
Generative Product Recommendations for Implicit Superlative Queries
2025 · arXiv (Cornell University)
In Recommender Systems, users often seek the best products through indirect, vague, or under-specified queries, such as "best shoes for trail running". Such queries, also referred to as implicit superlative queries, pose a significant challenge …
-
Discriminating between Similar Languages and Arabic Dialect Identification: A Report on the Third DSL Shared Task
2016 · International Conference on Computational Linguistics
We present the results of the third edition of the Discriminating between Similar Languages (DSL) shared task, which was organized as part of the VarDial’2016 workshop at COLING’2016. The challenge offered two subtasks: subtask 1 …
-
Findings of the 2019 Conference on Machine Translation (WMT19)
2019
Loïc Barrault, Ondřej Bojar, Marta R. Costa-jussà, Christian Federmann, Mark Fishel, Yvette Graham, Barry Haddow, Matthias Huck, Philipp Koehn, Shervin Malmasi, Christof Monz, Mathias Müller, Santanu Pal, Matt Post, Marcos Zampieri. Proceedings of the Fourth …
-
MultiCoNER: A Large-scale Multilingual dataset for Complex Named Entity Recognition
2022 · arXiv (Cornell University)
We present MultiCoNER, a large multilingual dataset for Named Entity Recognition that covers 3 domains (Wiki sentences, questions, and search queries) across 11 languages, as well as multilingual and code-mixing subsets. This dataset is designed …