A syntactic knowledge constrained paraphrase extraction for Chinese-English statistical machine translation
At a glance
- Citations
- 1
- References
- 22
- Comments
- 0
Öz
On the basis of general pivot method for paraphrase extraction which might introduce much noise in extracted paraphrases, this paper proposes a syntactic knowledge-enhanced method to extract higher-quality paraphrases to further improve the quality of statistical machine translation. Firstly, the syntactic knowledge is acquired and added to paraphrase extraction algorithm as constraints to obtain higher quality paraphrases from a parallel corpora. Then the extracted paraphrases are used to update the phrase table and reordering table of the translation system based on the unknown words and phrases in the source-side sentences of the development set and test set. Specifically, if a paraphrase of an unknown word or phrase in a source-side sentence of the development set and test set exists in the phrase table or reordering table, then a new phrase pair constructed with this unknown word or phrase with the target-side word or phrase of this paraphrase will be added to the phrase table or reordering table. In doing so, it can reduce the number of our-of-vocabulary words, and improve the coverage of the test set by the paraphrase-enhanced tables. Finally, the translation results are obtained using the updated phrase table and reordering table by translating the source language sentences. Experiments are carried out on NIST 2005 and NIST 2008 test sets, and results show that the method can significantly improve translation performance compared with the baseline system.
Publication details
- DOI
- 10.1109/cac.2017.8243697
- OpenAlex
- W2781575997
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
Oturum Açın to join the discussion.