conference-paper Open access

Copied Monolingual Data Improves Low-Resource Neural Machine Translation

Research footprint

At a glance

Citations
222
References
24
Comments
0
Paper overview

Abstract

We train a neural machine translation (NMT) system to both translate sourcelanguage text and copy target-language text, thereby exploiting monolingual corpora in the target language. Specifically, we create a bitext from the monolingual text in the target language so that each source sentence is identical to the target sentence. This copied data is then mixed with the parallel corpus and the NMT system is trained like normal, with no metadata to distinguish the two input languages.

Record transparency

Publication details

DOI
10.18653/v1/w17-4715
OpenAlex
W2756566411
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.