conference-paper
Dynamic Checkpoint Architecture for Reliability Improvement on Distributed Frameworks
Research footprint
At a glance
- Citations
- 3
- References
- 6
- Comments
- 0
Paper overview
Öz
Fault tolerant mechanisms are essential to provide reliable feature for distributed systems. Checkpoint and Recovery is a widely used technique that consists on saving data states for a fast recovery in case of failure. On Apache Hadoop and Apache Spark - distributed high performance frameworks -, checkpoint aims to help on recovery steps after failures. However, wrong configuration of checkpoint attributes can degrade system performance and reliability, thus losing checkpoint purpose. This work proposes a dynamic architecture for checkpoint based on system monitoring and alerts. In order to avoid checkpoint problems on Hadoop and Spark, one implementation of dynamic mechanism is defined for each framework.
Record transparency
Publication details
- DOI
- 10.1109/srds.2018.00038
- OpenAlex
- W2911210215
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
Oturum Açın to join the discussion.