conference-paper Open access

A Large-Scale Benchmark for Vietnamese Sentence Paraphrases

Research footprint

At a glance

Citations
0
References
0
Comments
0
Paper overview

Abstract

This paper presents ViSP, a high-quality Vietnamese dataset for sentence paraphrasing, consisting of 1.2M original-paraphrase pairs collected from various domains.The dataset was constructed using a hybrid approach that combines automatic paraphrase generation with manual evaluation to ensure high quality.We conducted experiments using methods such as back-translation, EDA, and baseline models like BART and T5, as well as large language models (LLMs), including GPT-4o, Gemini-1.5,Aya, Qwen-2.5, and Meta-Llama-3.1 variants.To the best of our knowledge, this is the first large-scale study on Vietnamese paraphrasing.We hope that our dataset and findings will serve as a valuable foundation for future research and applications in Vietnamese paraphrase tasks.

Record transparency

Publication details

DOI
10.18653/v1/2025.findings-naacl.59
OpenAlex
W4411113031
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.