preprint Open access

Building a Neural Machine Translation System Using Only Synthetic Parallel Data

  • arXiv (Cornell University)
  • Cornell University
Research footprint

At a glance

Citations
20
References
29
Comments
0
Paper overview

Abstract

Recent works have shown that synthetic parallel data automatically generated by translation models can be effective for various neural machine translation (NMT) issues. In this study, we build NMT systems using only synthetic parallel data. As an efficient alternative to real parallel data, we also present a new type of synthetic parallel corpus. The proposed pseudo parallel data are distinct from previous works in that ground truth and synthetic examples are mixed on both sides of sentence pairs. Experiments on Czech-German and French-German translations demonstrate the efficacy of the proposed pseudo parallel corpus, which shows not only enhanced results for bidirectional translation tasks but also substantial improvement with the aid of a ground truth real parallel corpus.

Record transparency

Publication details

DOI
10.48550/arxiv.1704.00253
OpenAlex
W2604275745
Document type
preprint
Language
EN
Source
arXiv (Cornell University)
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.