Papers
Topics
Authors
Recent
Search
2000 character limit reached

Memory-Sample Tradeoffs for Linear Regression with Small Error

Published 18 Apr 2019 in cs.LG and stat.ML | (1904.08544v2)

Abstract: We consider the problem of performing linear regression over a stream of dd-dimensional examples, and show that any algorithm that uses a subquadratic amount of memory exhibits a slower rate of convergence than can be achieved without memory constraints. Specifically, consider a sequence of labeled examples (a1,b1),(a2,b2),(a_1,b_1), (a_2,b_2)\ldots, with aia_i drawn independently from a dd-dimensional isotropic Gaussian, and where bi=ai,x+ηi,b_i = \langle a_i, x\rangle + \eta_i, for a fixed xR<sup>dx \in \mathbb{R}<sup>d with x2=1|x|_2 = 1 and with independent noise ηi\eta_i drawn uniformly from the interval [2<sup>d/5,2<sup>d/5].[-2<sup>{-d/5},2<sup>{-d/5}]. We show that any algorithm with at most d<sup>2/4d<sup>2/4 bits of memory requires at least Ω(dloglog1ϵ)\Omega(d \log \log \frac{1}{\epsilon}) samples to approximate xx to 2\ell_2 error ϵ\epsilon with probability of success at least $2/3$, for ϵ\epsilon sufficiently small as a function of dd. In contrast, for such ϵ\epsilon, xx can be recovered to error ϵ\epsilon with probability $1-o(1)$ with memory O(d<sup>2</sup>log(1/ϵ))O\left(d<sup>2</sup> \log(1/\epsilon)\right) using dd examples. This represents the first nontrivial lower bounds for regression with super-linear memory, and may open the door for strong memory/sample tradeoffs for continuous optimization.

Citations (35)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.