conference-paper

A Review on Job Scheduling for Hadoop Mapreduce

Research footprint

At a glance

Citations
4
References
16
Comments
0
Paper overview

Abstract

Hadoop is a distributed computing environment based on java which not only stores but also process the vast volume of data. It's HDFS (Hadoop Distributed File System) is for storing the data and analytics is done by MapReduce. MapReduce is an emerging paradigm for handling huge data sets using shared-nothing clusters. Lot of organizations have already adopted MapReduce for their analytics work. To boost the performance and utilization of the shared cluster, many scheduling mechanism are proposed by different authors. Many problems are faced during MapReduce jobs scheduling such as-locality, synchronization overhead, and fairness. Now, by introducing various scheduling issues concerned with locality, synchronization and fairness this paper surveys the various approaches to handle these problems. In addition, here evaluation of the various scheduling algorithms and for solving overhead during synchronization methods like asynchronous processing and speculative execution are also discussed.

Record transparency

Publication details

DOI
10.1109/icngcis.2017.40
OpenAlex
W2899720425
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.