Accelerating Containerized Machine Learning Workloads
At a glance
- الاستشهادات
- 1
- المراجع
- 54
- Comments
- 0
Abstract
To facilitate various Machine Learning (ML) training and inference tasks, enterprises tend to build large and expensive clusters and share them among different teams for diverse ML workloads. Virtualized platforms (containers/VMs) and schedulers are typically deployed to allow such access, manage heterogeneous resources and schedule ML jobs in these clusters. However, allocating resource budgets for different ML jobs to achieve best performance and cluster resource efficiency remains a significant challenge. This work proposes Nearchus to accelerate distributed ML training while ensuring high resource efficiency by using adaptive resource allocation. Nearchus automatically identifies potential performance bottlenecks for running jobs and re-allocates resources to provide optimized run-time performance with high resource efficiency. Nearchus’s resource configuration significantly improves the training speed of individual jobs up to 71.4%–129.1% against state-of-the-art resource schedulers, and reduces job completion and queuing time by 35.6% and 67.8%, respectively.
Publication details
- DOI
- 10.1109/noms59830.2024.10575188
- OpenAlex
- W4400237879
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
تسجيل الدخول للانضمام إلى النقاش.