conference-paper Open access

T-REx: A Large Scale Alignment of Natural Language with Knowledge Base Triples

Research footprint

At a glance

Citations
174
References
17
Comments
0
Paper overview

Öz

Alignments between natural language and Knowledge Base (KB) triples are an essential prerequisite for training machine learning approaches employed in a variety of Natural Language Processing problems.These include Relation Extraction, KB Population, Question Answering and Natural Language Generation from KB triples.Available datasets that provide those alignments are plagued by significant shortcomings -they are of limited size, they exhibit a restricted predicate coverage, and/or they are of unreported quality.To alleviate these shortcomings, we present T-REx, a dataset of large scale alignments between Wikipedia abstracts and Wikidata triples.T-REx consists of 11 million triples aligned with 3.09 million Wikipedia abstracts (6.2 million sentences).T-REx is two orders of magnitude larger than the largest available alignments dataset and covers 2.5 times more predicates.Additionally, we stress the quality of this language resource thanks to an extensive crowdsourcing evaluation.

Record transparency

Publication details

DOI
10.63317/2yu5nu76n7bu
OpenAlex
W2785611959
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Oturum Açın to join the discussion.

  1. No comments yet. Start the discussion.