conference-paper

Dynamic Tuple Scheduling with Prediction for Data Stream Processing Systems

Research footprint

At a glance

Citations
6
References
31
Comments
0
Paper overview

Abstract

For data stream processing systems such as Apache Heron, workload imbalance across processing instances often causes significant system performance degradation. To mitigate such issues, Apache Heron leverages a naive throttling-based back-pressure scheme, which may lead to unexpected system disruption. This calls for a finer-grained control to distribute data stream units (tuples) between successive instances, a.k.a. tuple scheduling, which well adapts to data stream variations and workload discrepancy. Besides, the benefits of predictive scheduling to data stream processing systems still remain unexplored. In this paper, we formulate tuple scheduling problem as a stochastic network optimization problem, with careful choices in the granularity of system modeling and decision making. With non-trivial transformation, we decouple the problem into a series of online subproblems. By exploiting unique subproblem structure, we propose POTUS, an efficient, online, and distributed scheduling scheme that employs the power of predictive scheduling but requires only limited system dynamics to achieve a tunable trade-off between communication cost reduction and system queue stability. Theoretical analysis and simulations show that POTUS effectively shortens response time with mild-value of future information, even in the face of misprediction. Our solution is also applicable to other data stream processing systems.

Record transparency

Publication details

DOI
10.1109/globecom38437.2019.9013570
OpenAlex
W2995648034
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.