preprint
وصول مفتوح
How Do Source-side Monolingual Word Embeddings Impact Neural Machine Translation?
Research footprint
At a glance
- الاستشهادات
- 2
- المراجع
- 11
- Comments
- 0
Paper overview
Abstract
Using pre-trained word embeddings as input layer is a common practice in many natural language processing (NLP) tasks, but it is largely neglected for neural machine translation (NMT). In this paper, we conducted a systematic analysis on the effect of using pre-trained source-side monolingual word embedding in NMT. We compared several strategies, such as fixing or updating the embeddings during NMT training on varying amounts of data, and we also proposed a novel strategy called dual-embedding that blends the fixing and updating strategies. Our results suggest that pre-trained embeddings can be helpful if properly incorporated into NMT, especially when parallel data is limited or additional in-domain monolingual data is readily available.
Record transparency
Publication details
- DOI
- 10.48550/arxiv.1806.01515
- OpenAlex
- W2806931283
- Document type
- preprint
- Language
- EN
- Source
- arXiv (Cornell University)
- Last metadata update
Comments
تسجيل الدخول للانضمام إلى النقاش.