Dynamic Tuple Scheduling with Prediction for Data Stream Processing Systems
At a glance
- الاستشهادات
- 6
- المراجع
- 31
- Comments
- 0
Abstract
For data stream processing systems such as Apache Heron, workload imbalance across processing instances often causes significant system performance degradation. To mitigate such issues, Apache Heron leverages a naive throttling-based back-pressure scheme, which may lead to unexpected system disruption. This calls for a finer-grained control to distribute data stream units (tuples) between successive instances, a.k.a. tuple scheduling, which well adapts to data stream variations and workload discrepancy. Besides, the benefits of predictive scheduling to data stream processing systems still remain unexplored. In this paper, we formulate tuple scheduling problem as a stochastic network optimization problem, with careful choices in the granularity of system modeling and decision making. With non-trivial transformation, we decouple the problem into a series of online subproblems. By exploiting unique subproblem structure, we propose POTUS, an efficient, online, and distributed scheduling scheme that employs the power of predictive scheduling but requires only limited system dynamics to achieve a tunable trade-off between communication cost reduction and system queue stability. Theoretical analysis and simulations show that POTUS effectively shortens response time with mild-value of future information, even in the face of misprediction. Our solution is also applicable to other data stream processing systems.
Publication details
- DOI
- 10.1109/globecom38437.2019.9013570
- OpenAlex
- W2995648034
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
تسجيل الدخول للانضمام إلى النقاش.