conference-paper

Accelerating Containerized Machine Learning Workloads

Research footprint

At a glance

Citations
1
References
54
Comments
0
Paper overview

Abstract

To facilitate various Machine Learning (ML) training and inference tasks, enterprises tend to build large and expensive clusters and share them among different teams for diverse ML workloads. Virtualized platforms (containers/VMs) and schedulers are typically deployed to allow such access, manage heterogeneous resources and schedule ML jobs in these clusters. However, allocating resource budgets for different ML jobs to achieve best performance and cluster resource efficiency remains a significant challenge. This work proposes Nearchus to accelerate distributed ML training while ensuring high resource efficiency by using adaptive resource allocation. Nearchus automatically identifies potential performance bottlenecks for running jobs and re-allocates resources to provide optimized run-time performance with high resource efficiency. Nearchus’s resource configuration significantly improves the training speed of individual jobs up to 71.4%–129.1% against state-of-the-art resource schedulers, and reduces job completion and queuing time by 35.6% and 67.8%, respectively.

Record transparency

Publication details

DOI
10.1109/noms59830.2024.10575188
OpenAlex
W4400237879
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.