conference-paper Open access

LC-QuAD: A Corpus for Complex Question Answering over Knowledge Graphs

  • Lecture notes in computer science
  • Springer Science+Business Media
Research footprint

At a glance

Citations
252
References
14
Comments
0
Paper overview

Abstract

Being able to access knowledge bases in an intuitive way has been an active area of research over the past years. In particular, several question answering (QA) approaches which allow to query RDF datasets in natural language have been developed as they allow end users to access knowledge without needing to learn the schema of a knowledge base and learn a formal query language. To foster this research area, several training datasets have been created, e.g. in the QALD (Question Answering over Linked Data) initiative. However, existing datasets are insufficient in terms of size, variety or complexity to apply and evaluate a range of machine learning based QA approaches for learning complex SPARQL queries. With the provision of the Large-Scale Complex Question Answering Dataset (LC-QuAD), we close this gap by providing a dataset with 5000 questions and their corresponding SPARQL queries over the DBpedia dataset. In this article, we describe the dataset creation process and how we ensure a high variety of questions, which should enable to assess the robustness and accuracy of the next generation of QA systems for knowledge graphs. These keywords were added by machine and not by the authors. This process is experimental and the keywords may be updated as the learning algorithm improves.

Record transparency

Publication details

DOI
10.1007/978-3-319-68204-4_22
OpenAlex
W2763039547
Document type
conference-paper
Language
EN
Source
Lecture notes in computer science
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.