Papers
Topics
Authors
Recent
Search
2000 character limit reached

Learning Over-Parametrized Two-Layer ReLU Neural Networks beyond NTK

Published 9 Jul 2020 in cs.LG, math.OC, and stat.ML | (2007.04596v1)

Abstract: We consider the dynamic of gradient descent for learning a two-layer neural network. We assume the input x∈R<sup>dx\in\mathbb{R}<sup>d is drawn from a Gaussian distribution and the label of xx satisfies f<sup>⋆(x)</sup>=a<sup>⊤∣W<sup>⋆x∣f<sup>{\star}(x)</sup> = a<sup>{\top}|W<sup>{\star}x|, where a∈R<sup>da\in\mathbb{R}<sup>d is a nonnegative vector and W<sup>⋆</sup>∈R<sup>d×</sup>dW<sup>{\star}</sup> \in\mathbb{R}<sup>{d\times</sup> d} is an orthonormal matrix. We show that an over-parametrized two-layer neural network with ReLU activation, trained by gradient descent from random initialization, can provably learn the ground truth network with population loss at most o(1/d)o(1/d) in polynomial time with polynomial samples. On the other hand, we prove that any kernel method, including Neural Tangent Kernel, with a polynomial number of samples in dd, has population loss at least Ω(1/d)\Omega(1 / d).

Citations (26)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.