Sebastian U. Stich
3 papers in the PaperMetrix corpus
Papers by this author
-
Error Feedback Fixes SignSGD and other Gradient Compression Schemes
2019 · arXiv (Cornell University)
Sign-based algorithms (e.g. signSGD) have been proposed as a biased gradient compression technique to alleviate the communication bottleneck in training large neural networks across multiple workers. We show simple convex counter-examples where signSGD does not …
-
Critical Parameters for Scalable Distributed Learning with Large Batches and Asynchronous Updates
2021 · arXiv (Cornell University)
It has been experimentally observed that the efficiency of distributed training with stochastic gradient (SGD) depends decisively on the batch size and -- in asynchronous implementations -- on the gradient staleness. Especially, it has been …
-
RelaySum for Decentralized Deep Learning on Heterogeneous Data
2021 · arXiv (Cornell University)
In decentralized machine learning, workers compute model updates on their local data. Because the workers only communicate with few neighbors without central coordination, these updates propagate progressively over the network. This paradigm enables distributed training …