Exploring Controlled RDF Distribution
At a glance
- Citations
- 9
- References
- 30
- Comments
- 0
Öz
RDF datasets have increased rapidly over the last few years. In order to process SPARQL queries on these large datasets, much effort has been spent on developing horizontally scalable techniques, which involve data partitioning and parallel query processing. While distribution may provide storage scalability, it may also incur high communication costs for processing queries. In this paper, we present a parallel and distributed query rocessing approach that explores the existence of data allocation patterns, provided by a controlled data distribution, that determine how RDF triples should be grouped and stored on the same server. Fragments of the RDF datastore follow a given allocation pattern and correspond also to units of communication among servers. Based on this distribution model, we define two communication strategies for query processing: get-frag, which requests remote servers to send fragments that contain data required by a query, and send-result, which forwards intermediate results. These strategies are combined on a method, called 2ways, that chooses the adequate communication strategy whenever queries traverse fragment boundaries. We provide a cost function used to determine this choice and present experimental results. They show that our proposed technique effectively reduces the communication cost and improves the response time for processing SPARQL queries on a distributed RDF datastore.
Publication details
- DOI
- 10.1109/cloudcom.2016.0038
- OpenAlex
- W2582046608
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
Oturum Açın to join the discussion.