article Open access

Inference-based schema discovery for RDF data

  • Data & Knowledge Engineering
  • Elsevier BV
Research footprint

At a glance

Citations
0
References
34
Comments
0
Paper overview

Abstract

The Semantic Web represents a huge information space where an increasing number of datasets, described in RDF, are made available to users and applications. In this context, the data is not constrained by a predefined schema. In RDF datasets, the schema may be incomplete or even missing. While this offers high flexibility in creating data sources, it also makes their use difficult. Several works have addressed the problem of automatic schema discovery for RDF datasets, but existing approaches rely only on the explicit information provided by the data source, which may limit the quality of the results. Indeed, in an RDF data source, an entity is described by explicitly declared properties, but also by implicit properties that can be derived using reasoning rules. These implicit properties are not considered by existing schema discovery approaches. In this work, we propose a first contribution towards a hybrid schema discovery approach capable of exploiting all the semantics of a data source, which is represented not only by the explicitly declared triples, but also by the ones that can be inferred through reasoning. By considering both explicit and implicit properties, the quality of the generated schema is improved. We provide a scalable design of our approach to enable the processing of large RDF data sources while improving the quality of the results. We present some experiments which demonstrate the efficiency of our proposal and the quality of the discovered schema.

Record transparency

Publication details

DOI
10.1016/j.datak.2025.102491
OpenAlex
W4412871368
Document type
article
Language
EN
Source
Data & Knowledge Engineering
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.