Papers
Topics
Authors
Recent
Search
2000 character limit reached

Global Convergence of Gradient Descent for Deep Linear Residual Networks

Published 2 Nov 2019 in cs.LG and stat.ML | (1911.00645v1)

Abstract: We analyze the global convergence of gradient descent for deep linear residual networks by proposing a new initialization: zero-asymmetric (ZAS) initialization. It is motivated by avoiding stable manifolds of saddle points. We prove that under the ZAS initialization, for an arbitrary target matrix, gradient descent converges to an ε\varepsilon-optimal point in O(L<sup>3</sup>log(1/ε))O(L<sup>3</sup> \log(1/\varepsilon)) iterations, which scales polynomially with the network depth LL. Our result and the exp(Ω(L))\exp(\Omega(L)) convergence time for the standard initialization (Xavier or near-identity) [Shamir, 2018] together demonstrate the importance of the residual structure and the initialization in the optimization for deep linear neural networks, especially when LL is large.

Authors (3)
Citations (22)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.