Researcher profile

Nizar Habash

7 papers in the PaperMetrix corpus

Publications

Papers by this author

  1. Optimizing Tokenization Choice for Machine Translation across Multiple Target Languages

    2017 · ˜The œPrague Bulletin of Mathematical Linguistics

    Abstract Tokenization is very helpful for Statistical Machine Translation (SMT), especially when translating from morphologically rich languages. Typically, a single tokenization scheme is applied to the entire source-language text and regardless of the target language. …

  2. A Cross-lingual Messenger with Keyword Searchable Phrases for the Travel Domain

    2018 · International Conference on Computational Linguistics

    We present Qutr (Query Translator), a smart cross-lingual communication application for the travel domain. Qutr is a real-time messaging app that automatically translates conversations while supporting keyword-to-sentence matching. Qutr relies on querying a database that …

  3. The Impact of Preprocessing on Arabic-English Statistical and Neural Machine Translation

    2019 · arXiv (Cornell University)

    Neural networks have become the state-of-the-art approach for machine translation (MT) in many languages. While linguistically-motivated tokenization techniques were shown to have significant effects on the performance of statistical MT, it remains unclear if those …

  4. CoNLL 2017 Shared Task: Multilingual Parsing from Raw Text to Universal Dependencies

    2017

    Daniel Zeman, Martin Popel, Milan Straka, Jan Hajič, Joakim Nivre, Filip Ginter, Juhani Luotolahti, Sampo Pyysalo, Slav Petrov, Martin Potthast, Francis Tyers, Elena Badmaeva, Memduh Gokirmak, Anna Nedoluzhko, Silvie Cinková, Jan Hajič jr., Jaroslava Hlaváčová, …

  5. Fine-Grained Arabic Dialect Identification

    2018 · International Conference on Computational Linguistics

    Previous work on the problem of Arabic Dialect Identification typically targeted coarse-grained five dialect classes plus Standard Arabic (6-way classification). This paper presents the first results on a fine-grained dialect classification task covering 25 specific …

  6. The MADAR Shared Task on Arabic Fine-Grained Dialect Identification

    2019

    In this paper, we present the results and findings of the MADAR Shared Task on Arabic Fine-Grained Dialect Identification. This shared task was organized as part of The Fourth Arabic Natural Language Processing Workshop, collocated …

  7. The Interplay of Variant, Size, and Task Type in Arabic Pre-trained Language Models

    2021 · arXiv (Cornell University)

    In this paper, we explore the effects of language variants, data sizes, and fine-tuning task types in Arabic pre-trained language models. To do so, we build three pre-trained language models across three variants of Arabic: …