conference-paper

Dynamic Checkpoint Architecture for Reliability Improvement on Distributed Frameworks

Research footprint

At a glance

Citations
3
References
6
Comments
0
Paper overview

Öz

Fault tolerant mechanisms are essential to provide reliable feature for distributed systems. Checkpoint and Recovery is a widely used technique that consists on saving data states for a fast recovery in case of failure. On Apache Hadoop and Apache Spark - distributed high performance frameworks -, checkpoint aims to help on recovery steps after failures. However, wrong configuration of checkpoint attributes can degrade system performance and reliability, thus losing checkpoint purpose. This work proposes a dynamic architecture for checkpoint based on system monitoring and alerts. In order to avoid checkpoint problems on Hadoop and Spark, one implementation of dynamic mechanism is defined for each framework.

Record transparency

Publication details

DOI
10.1109/srds.2018.00038
OpenAlex
W2911210215
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Oturum Açın to join the discussion.

  1. No comments yet. Start the discussion.