conference-paper

Paraphrase Generation with Chinese Short Text Dataset

Research footprint

At a glance

Citations
1
References
21
Comments
0
Paper overview

Öz

An obstacle of conducting investigation on paraphrase generation is short of high-quality, publicly-available labeled dataset of sentential paraphrases, which is particularly serious for Chinese paraphrase generation research. Therefore, the study in Chinese paraphrase generation is the starting stage. This paper aimed to use a novel way to create Chinese paraphrase dataset, which contains 8K sentences pairs. The data source comes from a bank QA dataset, in which there are several sentences to express each problem. By calculating the similarity between the same semantic sentences, we can obtain paraphrase pairs to create Chinese paraphrase dataset. Then, we achieve paraphrase generation task by leveraging a classical Seq2Sseq model with attention mechanism. Following previous work and evaluate paraphrase generation result on our Chinese dataset. Experimental results not only show that the dataset is suitable for Chinese paraphrase generation task, but also provides a benchmark for further research on this research area.

Record transparency

Publication details

DOI
10.1109/iccia49625.2020.00019
OpenAlex
W3082886490
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Oturum Açın to join the discussion.

  1. No comments yet. Start the discussion.