conference-paper Open access

Generating Paraphrases with Lean Vocabulary

Research footprint

At a glance

Citations
1
References
19
Comments
0
Paper overview

Abstract

In this work, we examine whether it is possible to achieve the state of the art performance in paraphrase generation with reduced vocabulary. Our approach consists of building a convolution to sequence model (Conv2Seq) partially guided by the reinforcement learning, and training it on the subword representation of the input. The experiment on the Quora dataset, which contains over 140,000 pairs of sentences and corresponding paraphrases, found that with less than 1,000 token types, we were able to achieve performance that exceeded that of the current state of the art.

Record transparency

Publication details

DOI
10.18653/v1/w19-8655
OpenAlex
W2996171631
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.