conference-paper
Open access
Copied Monolingual Data Improves Low-Resource Neural Machine Translation
Research footprint
At a glance
- Citations
- 222
- References
- 24
- Comments
- 0
Paper overview
Abstract
We train a neural machine translation (NMT) system to both translate sourcelanguage text and copy target-language text, thereby exploiting monolingual corpora in the target language. Specifically, we create a bitext from the monolingual text in the target language so that each source sentence is identical to the target sentence. This copied data is then mixed with the parallel corpus and the NMT system is trained like normal, with no metadata to distinguish the two input languages.
Record transparency
Publication details
- DOI
- 10.18653/v1/w17-4715
- OpenAlex
- W2756566411
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
Log in to join the discussion.