conference-paper

A syntactic knowledge constrained paraphrase extraction for Chinese-English statistical machine translation

Research footprint

At a glance

Citations
1
References
22
Comments
0
Paper overview

Öz

On the basis of general pivot method for paraphrase extraction which might introduce much noise in extracted paraphrases, this paper proposes a syntactic knowledge-enhanced method to extract higher-quality paraphrases to further improve the quality of statistical machine translation. Firstly, the syntactic knowledge is acquired and added to paraphrase extraction algorithm as constraints to obtain higher quality paraphrases from a parallel corpora. Then the extracted paraphrases are used to update the phrase table and reordering table of the translation system based on the unknown words and phrases in the source-side sentences of the development set and test set. Specifically, if a paraphrase of an unknown word or phrase in a source-side sentence of the development set and test set exists in the phrase table or reordering table, then a new phrase pair constructed with this unknown word or phrase with the target-side word or phrase of this paraphrase will be added to the phrase table or reordering table. In doing so, it can reduce the number of our-of-vocabulary words, and improve the coverage of the test set by the paraphrase-enhanced tables. Finally, the translation results are obtained using the updated phrase table and reordering table by translating the source language sentences. Experiments are carried out on NIST 2005 and NIST 2008 test sets, and results show that the method can significantly improve translation performance compared with the baseline system.

Record transparency

Publication details

DOI
10.1109/cac.2017.8243697
OpenAlex
W2781575997
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Oturum Açın to join the discussion.

  1. No comments yet. Start the discussion.