Enhancing Machine Learning Optimization Algorithms by Leveraging Memory Caching (Research Poster)
At a glance
- Citations
- 0
- References
- 0
- Comments
- 0
Öz
Searching a solution space using Stochastic Gradient Descent (SGD) depends on the examples picked at each iteration of the algorithm. Therefore, best practices suggest randomizing the order of training points to visit after every epoch. This random selection is typically implemented as a random shuffling of the order of the training vectors rather than a genuine random training point selection. The shuffling is usually performed after every epoch which results in an extremely low temporal locality of access to the training set. Indeed, each training point is used once, and not before all the other training points have been visited. This means that a cache layer in the memory hierarchy of a modern HPC computer system will have little benefit for the algorithm unless all the training points fit inside that cache.
Publication details
- DOI
- 10.1109/hpcs.2018.00171
- OpenAlex
- W2899385487
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
Oturum Açın to join the discussion.