Papers
Topics
Authors
Recent
Search
2000 character limit reached

On the Optimal Weighted â„“2\ell_2 Regularization in Overparameterized Linear Regression

Published 10 Jun 2020 in stat.ML, cs.LG, math.ST, and stat.TH | (2006.05800v4)

Abstract: We consider the linear model y=Xβ<em>⋆+ϵ\mathbf{y} = \mathbf{X} \mathbf{\beta}<em>\star + \mathbf{\epsilon} with X∈R<sup>n×</sup>p\mathbf{X}\in \mathbb{R}<sup>{n\times</sup> p} in the overparameterized regime $p&gt;n$. We estimate β</em>⋆\mathbf{\beta}</em>\star via generalized (weighted) ridge regression: β^<em>λ=(X<sup>TX</sup>+λΣw)<sup>†</sup>X<sup>Ty\hat{\mathbf{\beta}}<em>\lambda = \left(\mathbf{X}<sup>T\mathbf{X}</sup> + \lambda \mathbf{\Sigma}_w\right)<sup>\dagger</sup> \mathbf{X}<sup>T\mathbf{y}, where Σw\mathbf{\Sigma}_w is the weighting matrix. Under a random design setting with general data covariance Σx\mathbf{\Sigma}_x and anisotropic prior on the true coefficients Eβ</em>⋆β<em>⋆<sup>T</sup>=Σ</em>β\mathbb{E}\mathbf{\beta}</em>\star\mathbf{\beta}<em>\star<sup>T</sup> = \mathbf{\Sigma}</em>\beta, we provide an exact characterization of the prediction risk E(y−x<sup>Tβ^λ)<sup>2\mathbb{E}(y-\mathbf{x}<sup>T\hat{\mathbf{\beta}}_\lambda)<sup>2 in the proportional asymptotic limit p/n→γ∈(1,∞)p/n\rightarrow \gamma \in (1,\infty). Our general setup leads to a number of interesting findings. We outline precise conditions that decide the sign of the optimal setting λopt\lambda_{\rm opt} for the ridge parameter λ\lambda and confirm the implicit ℓ2\ell_2 regularization effect of overparameterization, which theoretically justifies the surprising empirical observation that λopt\lambda_{\rm opt} can be negative in the overparameterized regime. We also characterize the double descent phenomenon for principal component regression (PCR) when both X\mathbf{X} and β<em>⋆\mathbf{\beta}<em>\star are anisotropic. Finally, we determine the optimal weighting matrix Σw\mathbf{\Sigma}_w for both the ridgeless (λ→0\lambda\to 0) and optimally regularized (λ=λ</em>opt\lambda = \lambda</em>{\rm opt}) case, and demonstrate the advantage of the weighted objective over standard ridge regression and PCR.

Authors (2)
Citations (115)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.

Tweets

Sign up for free to view the 2 tweets with 7 likes about this paper.