article وصول مفتوح

Optimizing Tokenization Choice for Machine Translation across Multiple Target Languages

  • ˜The œPrague Bulletin of Mathematical Linguistics
  • De Gruyter Open
Research footprint

At a glance

الاستشهادات
19
المراجع
24
Comments
0
Paper overview

Abstract

Abstract Tokenization is very helpful for Statistical Machine Translation (SMT), especially when translating from morphologically rich languages. Typically, a single tokenization scheme is applied to the entire source-language text and regardless of the target language. In this paper, we evaluate the hypothesis that SMT performance may benefit from different tokenization schemes for different words within the same text, and also for different target languages. We apply this approach to Arabic as a source language, with five target languages of varying morphological complexity: English, French, Spanish, Russian and Chinese. Our results show that different target languages indeed require different source-language schemes; and a context-variable tokenization scheme can outperform a context-constant scheme with a statistically significant performance enhancement of about 1.4 BLEU points.

Record transparency

Publication details

DOI
10.1515/pralin-2017-0025
OpenAlex
W2624461416
Document type
article
Language
EN
Source
˜The œPrague Bulletin of Mathematical Linguistics
Last metadata update
المجتمع

Comments

تسجيل الدخول للانضمام إلى النقاش.

  1. لا توجد تعليقات بعد. ابدأ النقاش.