Scavenger: A Cloud Service for Optimizing Cost and Performance of DL Training
At a glance
- Citations
- 0
- References
- 3
- Comments
- 0
Abstract
Deep learning (DL) models learn non-linear functions and relationships by iteratively training on given data. To accelerate training further, data-parallel training [1] launches multiple instances of training process on separate partitions of data and periodically aggregates model updates. With the availability of VMs in the cloud, choosing the “right“ cluster configuration for data-parallel training presents non-trivial challenges. We tackle this problem by considering both the parallel and statistical efficiency of distributed training w.r.t. the cluster size configuration and batch-size in training. We build performance models to evaluate the pareto-relationship between cost and time of DL training across different cluster and batch-size configurations and develop Scavenger as a cloud service for searching optimum cloud configurations in an online, blackbox manner.
Publication details
- DOI
- 10.1109/ccgridw59191.2023.00081
- OpenAlex
- W4384835138
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
Log in to join the discussion.