Nizar Habash
7 أوراق في مجموعة PaperMetrix
أوراق هذا المؤلف
-
Optimizing Tokenization Choice for Machine Translation across Multiple Target Languages
2017 · The Prague Bulletin of Mathematical Linguistics
Abstract Tokenization is very helpful for Statistical Machine Translation (SMT), especially when translating from morphologically rich languages. Typically, a single tokenization scheme is applied to the entire source-language text and regardless of the target language. …
-
A Cross-lingual Messenger with Keyword Searchable Phrases for the Travel Domain
2018 · International Conference on Computational Linguistics
We present Qutr (Query Translator), a smart cross-lingual communication application for the travel domain. Qutr is a real-time messaging app that automatically translates conversations while supporting keyword-to-sentence matching. Qutr relies on querying a database that …
-
The Impact of Preprocessing on Arabic-English Statistical and Neural Machine Translation
2019 · arXiv (Cornell University)
Neural networks have become the state-of-the-art approach for machine translation (MT) in many languages. While linguistically-motivated tokenization techniques were shown to have significant effects on the performance of statistical MT, it remains unclear if those …
-
CoNLL 2017 Shared Task: Multilingual Parsing from Raw Text to Universal Dependencies
2017
Daniel Zeman, Martin Popel, Milan Straka, Jan Hajič, Joakim Nivre, Filip Ginter, Juhani Luotolahti, Sampo Pyysalo, Slav Petrov, Martin Potthast, Francis Tyers, Elena Badmaeva, Memduh Gokirmak, Anna Nedoluzhko, Silvie Cinková, Jan Hajič jr., Jaroslava Hlaváčová, …
-
Fine-Grained Arabic Dialect Identification
2018 · International Conference on Computational Linguistics
Previous work on the problem of Arabic Dialect Identification typically targeted coarse-grained five dialect classes plus Standard Arabic (6-way classification). This paper presents the first results on a fine-grained dialect classification task covering 25 specific …
-
The MADAR Shared Task on Arabic Fine-Grained Dialect Identification
2019
In this paper, we present the results and findings of the MADAR Shared Task on Arabic Fine-Grained Dialect Identification. This shared task was organized as part of The Fourth Arabic Natural Language Processing Workshop, collocated …
-
The Interplay of Variant, Size, and Task Type in Arabic Pre-trained Language Models
2021 · arXiv (Cornell University)
In this paper, we explore the effects of language variants, data sizes, and fine-tuning task types in Arabic pre-trained language models. To do so, we build three pre-trained language models across three variants of Arabic: …