Papers
Topics
Authors
Recent
Search
2000 character limit reached

Communication trade-offs for synchronized distributed SGD with large step size

Published 25 Apr 2019 in cs.LG, math.OC, and stat.ML | (1904.11325v1)

Abstract: Synchronous mini-batch SGD is state-of-the-art for large-scale distributed machine learning. However, in practice, its convergence is bottlenecked by slow communication rounds between worker nodes. A natural solution to reduce communication is to use the \emph{`local-SGD'} model in which the workers train their model independently and synchronize every once in a while. This algorithm improves the computation-communication trade-off but its convergence is not understood very well. We propose a non-asymptotic error analysis, which enables comparison to \emph{one-shot averaging} i.e., a single communication round among independent workers, and \emph{mini-batch averaging} i.e., communicating at every step. We also provide adaptive lower bounds on the communication frequency for large step-sizes (t<sup>−α</sup> t<sup>{-\alpha}</sup> , α∈(1/2,1) \alpha\in (1/2 , 1 ) ) and show that \emph{Local-SGD} reduces communication by a factor of O(TP<sup>3/2)O\Big(\frac{\sqrt{T}}{P<sup>{3/2}}\Big), with TT the total number of gradients and PP machines.

Citations (26)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.