Papers
Topics
Authors
Recent
Search
2000 character limit reached

Neural network learns low-dimensional polynomials with SGD near the information-theoretic limit

Published 3 Jun 2024 in cs.LG and stat.ML | (2406.01581v2)

Abstract: We study the problem of gradient descent learning of a single-index target function f<em>(x)=σ</em>(⟨x,θ⟩)f_<em>(\boldsymbol{x}) = \textstyle\sigma_</em>\left(\langle\boldsymbol{x},\boldsymbol{\theta}\rangle\right) under isotropic Gaussian data in R<sup>d\mathbb{R}<sup>d, where the unknown link function σ<em>:R→R\sigma_<em>:\mathbb{R}\to\mathbb{R} has information exponent pp (defined as the lowest degree in the Hermite expansion). Prior works showed that gradient-based training of neural networks can learn this target with n≳d<sup>Θ(p)n\gtrsim d<sup>{\Theta(p)} samples, and such complexity is predicted to be necessary by the correlational statistical query lower bound. Surprisingly, we prove that a two-layer neural network optimized by an SGD-based algorithm (on the squared loss) learns f</em>f_</em> with a complexity that is not governed by the information exponent. Specifically, for arbitrary polynomial single-index models, we establish a sample and runtime complexity of n≃T=Θ(d!⋅!polylogd)n \simeq T = \Theta(d!\cdot! \mathrm{polylog} d), where Θ(⋅)\Theta(\cdot) hides a constant only depending on the degree of σ<em>\sigma_<em>; this dimension dependence matches the information theoretic limit up to polylogarithmic factors. More generally, we show that n≳d<sup>(p</sup></em>−1)∨1n\gtrsim d<sup>{(p_</sup></em>-1)\vee 1} samples are sufficient to achieve low generalization error, where p∗≤pp_* \le p is the \textit{generative exponent} of the link function. Core to our analysis is the reuse of minibatch in the gradient computation, which gives rise to higher-order information beyond correlational queries.

Citations (12)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.