Papers
Topics
Authors
Recent
Search
2000 character limit reached

Generalization Bounds for Neural Networks via Approximate Description Length

Published 13 Oct 2019 in cs.LG and stat.ML | (1910.05697v1)

Abstract: We investigate the sample complexity of networks with bounds on the magnitude of its weights. In particular, we consider the class [ H=\left{W_t\circ\rho\circ \ldots\circ\rho\circ W_{1} :W_1,\ldots,W_{t-1}\in M_{d, d}, W_t\in M_{1,d}\right} ] where the spectral norm of each WiW_i is bounded by O(1)O(1), the Frobenius norm is bounded by RR, and ρ\rho is the sigmoid function e<sup>x1+e<sup>x\frac{e<sup>x}{1+e<sup>x} or the smoothened ReLU function ln(1+e<sup>x) \ln (1+e<sup>x). We show that for any depth tt, if the inputs are in [1,1]<sup>d[-1,1]<sup>d, the sample complexity of HH is O~(dR<sup>2ϵ<sup>2)\tilde O\left(\frac{dR<sup>2}{\epsilon<sup>2}\right). This bound is optimal up to log-factors, and substantially improves over the previous state of the art of O~(d<sup>2R<sup>2ϵ<sup>2)\tilde O\left(\frac{d<sup>2R<sup>2}{\epsilon<sup>2}\right). We furthermore show that this bound remains valid if instead of considering the magnitude of the WiW_i's, we consider the magnitude of WiWi<sup>0W_i - W_i<sup>0, where Wi<sup>0W_i<sup>0 are some reference matrices, with spectral norm of O(1)O(1). By taking the Wi<sup>0W_i<sup>0 to be the matrices at the onset of the training process, we get sample complexity bounds that are sub-linear in the number of parameters, in many typical regimes of parameters. To establish our results we develop a new technique to analyze the sample complexity of families HH of predictors. We start by defining a new notion of a randomized approximate description of functions f:XR<sup>df:X\to\mathbb{R}<sup>d. We then show that if there is a way to approximately describe functions in a class HH using dd bits, then d/ϵ<sup>2d/\epsilon<sup>2 examples suffices to guarantee uniform convergence. Namely, that the empirical loss of all the functions in the class is ϵ\epsilon-close to the true loss. Finally, we develop a set of tools for calculating the approximate description length of classes of functions that can be presented as a composition of linear function classes and non-linear functions.

Authors (2)
Citations (19)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.