ملف الباحث
Amirkeivan Mohtashami
ورقة واحدة في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
Critical Parameters for Scalable Distributed Learning with Large Batches and Asynchronous Updates
2021 · arXiv (Cornell University)
It has been experimentally observed that the efficiency of distributed training with stochastic gradient (SGD) depends decisively on the batch size and -- in asynchronous implementations -- on the gradient staleness. Especially, it has been …