article

Empirical Use of Information Retrieval to Build Synthetic Data for SMT Domain Adaptation

  • IEEE/ACM Transactions on Audio Speech and Language Processing
  • Institute of Electrical and Electronics Engineers
Research footprint

At a glance

Citations
12
References
61
Comments
0
Paper overview

Öz

In this paper, we present information retrieval as a powerful tool for addressing an imperative problem in the field of statistical machine translation, i.e., improving translation quality when not enough parallel corpora are available. We devise a framework, which uses information retrieval to create a synthetic corpus from the easily available monolingual corpora. We propose an improved unsupervised training approach with a data selection mechanism, which selects only the most appropriate sentences, thus reducing the amount of data, which is less related to the domain in the additional bitext. We also introduce a new method to choose sentences based on their relative similarity/difference from the query sentence. Using the synthetic corpus created by our method, we are able to improve state-of-the-art statistical machine translation systems.

Record transparency

Publication details

DOI
10.1109/taslp.2016.2517318
OpenAlex
W2294808479
Document type
article
Language
EN
Source
IEEE/ACM Transactions on Audio Speech and Language Processing
Last metadata update
Community

Comments

Oturum Açın to join the discussion.

  1. No comments yet. Start the discussion.