Researcher profile

Shervin Malmasi

9 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. German Dialect Identification in Interview Transcriptions

    2017

    This paper presents three systems submitted to the German Dialect Identification (GDI) task at the VarDial Evaluation Campaign 2017. The task consists of training models to identify the dialect of Swiss-German speech transcripts. The dialects …

  2. Classifier Ensembles for Dialect and Language Variety Identification

    2018 · arXiv (Cornell University)

    In this paper we present ensemble-based systems for dialect and language variety identification using the datasets made available by the organizers of the VarDial Evaluation Campaign 2018. We present a system developed to discriminate between …

  3. Language Identification and Morphosyntactic Tagging: The Second VarDial Evaluation Campaign

    2018 · Työväentutkimus Vuosikirja

    We present the results and the findings of the Second VarDial Evaluation Campaign on Natural Language Processing (NLP) for Similar Languages, Varieties and Dialects. The campaign was organized as part of the fifth edition of …

  4. CycleNER: An Unsupervised Training Approach for Named Entity Recognition

    2022 · Proceedings of the ACM Web Conference 2022

    Named Entity Recognition (NER) is a crucial natural language understanding task for many down-stream tasks such as question answering and retrieval. Despite significant progress in developing NER models for multiple languages and domains, scaling to …

  5. Reinforced Question Rewriting for Conversational Question Answering

    2022 · arXiv (Cornell University)

    Conversational Question Answering (CQA) aims to answer questions contained within dialogues, which are not easily interpretable without context. Developing a model to rewrite conversational questions into self-contained ones is an emerging solution in industry settings …

  6. Generative Product Recommendations for Implicit Superlative Queries

    2025 · arXiv (Cornell University)

    In Recommender Systems, users often seek the best products through indirect, vague, or under-specified queries, such as "best shoes for trail running". Such queries, also referred to as implicit superlative queries, pose a significant challenge …

  7. Discriminating between Similar Languages and Arabic Dialect Identification: A Report on the Third DSL Shared Task

    2016 · International Conference on Computational Linguistics

    We present the results of the third edition of the Discriminating between Similar Languages (DSL) shared task, which was organized as part of the VarDial’2016 workshop at COLING’2016. The challenge offered two subtasks: subtask 1 …

  8. Findings of the 2019 Conference on Machine Translation (WMT19)

    2019

    Loïc Barrault, Ondřej Bojar, Marta R. Costa-jussà, Christian Federmann, Mark Fishel, Yvette Graham, Barry Haddow, Matthias Huck, Philipp Koehn, Shervin Malmasi, Christof Monz, Mathias Müller, Santanu Pal, Matt Post, Marcos Zampieri. Proceedings of the Fourth …

  9. MultiCoNER: A Large-scale Multilingual dataset for Complex Named Entity Recognition

    2022 · arXiv (Cornell University)

    We present MultiCoNER, a large multilingual dataset for Named Entity Recognition that covers 3 domains (Wiki sentences, questions, and search queries) across 11 languages, as well as multilingual and code-mixing subsets. This dataset is designed …