Papers
Topics
Authors
Recent
Search
2000 character limit reached

How Over-Parameterization Slows Down Gradient Descent in Matrix Sensing: The Curses of Symmetry and Initialization

Published 3 Oct 2023 in cs.LG, math.OC, and stat.ML | (2310.01769v3)

Abstract: This paper rigorously shows how over-parameterization changes the convergence behaviors of gradient descent (GD) for the matrix sensing problem, where the goal is to recover an unknown low-rank ground-truth matrix from near-isotropic linear measurements. First, we consider the symmetric setting with the symmetric parameterization where M<sup>∗</sup>∈R<sup>n</sup>×nM<sup>*</sup> \in \mathbb{R}<sup>{n</sup> \times n} is a positive semi-definite unknown matrix of rank r≪nr \ll n, and one uses a symmetric parameterization XX<sup>⊤XX<sup>\top to learn M<sup>∗M<sup>*. Here X∈R<sup>n</sup>×kX \in \mathbb{R}<sup>{n</sup> \times k} with $k &gt; r$ is the factor matrix. We give a novel Ω(1/T<sup>2)\Omega (1/T<sup>2) lower bound of randomly initialized GD for the over-parameterized case ($k &gt;r$) where TT is the number of iterations. This is in stark contrast to the exact-parameterization scenario (k=rk=r) where the convergence rate is exp⁡(−Ω(T))\exp (-\Omega (T)). Next, we study asymmetric setting where M<sup>∗</sup>∈R<sup>n1</sup>×n2M<sup>*</sup> \in \mathbb{R}<sup>{n_1</sup> \times n_2} is the unknown matrix of rank r≪min⁡n1,n2r \ll \min{n_1,n_2}, and one uses an asymmetric parameterization FG<sup>⊤FG<sup>\top to learn M<sup>∗M<sup>* where F∈R<sup>n1</sup>×kF \in \mathbb{R}<sup>{n_1</sup> \times k} and G∈R<sup>n2</sup>×kG \in \mathbb{R}<sup>{n_2</sup> \times k}. Building on prior work, we give a global exact convergence result of randomly initialized GD for the exact-parameterization case (k=rk=r) with an exp⁡(−Ω(T))\exp (-\Omega(T)) rate. Furthermore, we give the first global exact convergence result for the over-parameterization case ($k&gt;r$) with an exp⁡(−Ω(α<sup>2</sup>T))\exp(-\Omega(\alpha<sup>2</sup> T)) rate where α\alpha is the initialization scale. This linear convergence result in the over-parameterization case is especially significant because one can apply the asymmetric parameterization to the symmetric setting to speed up from Ω(1/T<sup>2)\Omega (1/T<sup>2) to linear convergence. On the other hand, we propose a novel method that only modifies one step of GD and obtains a convergence rate independent of α\alpha, recovering the rate in the exact-parameterization case.

Authors (3)
Citations (6)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.