conference-paper

Optimization of spark storage solutions

Research footprint

At a glance

الاستشهادات
1
المراجع
29
Comments
0
Paper overview

Abstract

With the increasing demands of big data processing, distributed data processing frameworks like Hadoop and Spark are enjoying growing popularity. To make the best use of these frameworks, performance enhancement becomes a key point to focus on. In this paper, we propose a cost-based optimization of Spark' s storage solutions. In Spark, data is presented as RDDs and they have storage levels to indicate their storage mechanism. Our optimization process is an offline optimization method, which consists of data sampling and training processes. From our evaluations, it shows that our cost-based optimization is effective. It can improve the performance of a Spark application by up to 16%.

Record transparency

Publication details

DOI
10.1109/pic.2016.7949547
OpenAlex
W2672903086
Document type
conference-paper
Language
EN
Last metadata update
المجتمع

Comments

تسجيل الدخول للانضمام إلى النقاش.

  1. لا توجد تعليقات بعد. ابدأ النقاش.