preprint Open access

Character-based NMT with Transformer

  • arXiv (Cornell University)
  • Cornell University
Research footprint

At a glance

Citations
8
References
43
Comments
0
Paper overview

Abstract

Character-based translation has several appealing advantages, but its performance is in general worse than a carefully tuned BPE baseline. In this paper we study the impact of character-based input and output with the Transformer architecture. In particular, our experiments on EN-DE show that character-based Transformer models are more robust than their BPE counterpart, both when translating noisy text, and when translating text from a different domain. To obtain comparable BLEU scores in clean, in-domain data and close the gap with BPE-based models we use known techniques to train deeper Transformer models.

Record transparency

Publication details

DOI
10.48550/arxiv.1911.04997
OpenAlex
W2984171820
Document type
preprint
Language
EN
Source
arXiv (Cornell University)
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.