ملف الباحث
Wotao Yin
ورقتان في مجموعة PaperMetrix
المنشورات
أوراق هذا المؤلف
-
CADA: Communication-Adaptive Distributed Adam
2020 · arXiv (Cornell University)
Stochastic gradient descent (SGD) has taken the stage as the primary workhorse for large-scale machine learning. It is often used with its adaptive variants such as AdaGrad, Adam, and AMSGrad. This paper proposes an adaptive …
-
Exponential Graph is Provably Efficient for Decentralized Deep Training
2021 · arXiv (Cornell University)
Decentralized SGD is an emerging training method for deep learning known for its much less (thus faster) communication per iteration, which relaxes the averaging step in parallel SGD to inexact averaging. The less exact the …