preprint
Open access
Character-based NMT with Transformer
Research footprint
At a glance
- Citations
- 8
- References
- 43
- Comments
- 0
Paper overview
Abstract
Character-based translation has several appealing advantages, but its performance is in general worse than a carefully tuned BPE baseline. In this paper we study the impact of character-based input and output with the Transformer architecture. In particular, our experiments on EN-DE show that character-based Transformer models are more robust than their BPE counterpart, both when translating noisy text, and when translating text from a different domain. To obtain comparable BLEU scores in clean, in-domain data and close the gap with BPE-based models we use known techniques to train deeper Transformer models.
Record transparency
Publication details
- DOI
- 10.48550/arxiv.1911.04997
- OpenAlex
- W2984171820
- Document type
- preprint
- Language
- EN
- Source
- arXiv (Cornell University)
- Last metadata update
Comments
Log in to join the discussion.