conference-paper
CQADupStack
Research footprint
At a glance
- الاستشهادات
- 77
- المراجع
- 37
- Comments
- 0
Paper overview
Abstract
This paper presents a benchmark dataset, CQADupStack, for use in community question-answering (cQA) research. It contains threads from twelve StackExchange subforums, annotated with duplicate question information. We provide pre-defined training and test splits, both for retrieval and classification experiments, to ensure maximum comparability between different studies using the set. Furthermore, it comes with a script to manipulate the data in various ways. We give an analysis of the data in the set, and report benchmark results on a duplicate question retrieval task using well established retrieval models.
Record transparency
Publication details
- DOI
- 10.1145/2838931.2838934
- OpenAlex
- W2256784804
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
تسجيل الدخول للانضمام إلى النقاش.