Papers
Topics
Authors
Recent
Search
2000 character limit reached

Optimal bounds for â„“p\ell_p sensitivity sampling via â„“2\ell_2 augmentation

Published 1 Jun 2024 in cs.DS, cs.LG, and stat.ML | (2406.00328v1)

Abstract: Data subsampling is one of the most natural methods to approximate a massively large data set by a small representative proxy. In particular, sensitivity sampling received a lot of attention, which samples points proportional to an individual importance measure called sensitivity. This framework reduces in very general settings the size of data to roughly the VC dimension dd times the total sensitivity S\mathfrak S while providing strong (1±ε)(1\pm\varepsilon) guarantees on the quality of approximation. The recent work of Woodruff & Yasuda (2023c) improved substantially over the general O~(ε<sup>−2</sup>Sd)\tilde O(\varepsilon<sup>{-2}\mathfrak</sup> Sd) bound for the important problem of ℓp\ell_p subspace embeddings to O~(ε<sup>−2</sup>S<sup>2/p)\tilde O(\varepsilon<sup>{-2}\mathfrak</sup> S<sup>{2/p}) for p∈[1,2]p\in[1,2]. Their result was subsumed by an earlier O~(ε<sup>−2</sup>Sd<sup>1−p/2)\tilde O(\varepsilon<sup>{-2}\mathfrak</sup> Sd<sup>{1-p/2}) bound which was implicitly given in the work of Chen & Derezinski (2021). We show that their result is tight when sampling according to plain ℓp\ell_p sensitivities. We observe that by augmenting the ℓp\ell_p sensitivities by ℓ2\ell_2 sensitivities, we obtain better bounds improving over the aforementioned results to optimal linear O~(ε<sup>−2(</sup>S+d))=O~(ε<sup>−2d)\tilde O(\varepsilon<sup>{-2}(\mathfrak</sup> S+d)) = \tilde O(\varepsilon<sup>{-2}d) sampling complexity for all p∈[1,2]p \in [1,2]. In particular, this resolves an open question of Woodruff & Yasuda (2023c) in the affirmative for p∈[1,2]p \in [1,2] and brings sensitivity subsampling into the regime that was previously only known to be possible using Lewis weights (Cohen & Peng, 2015). As an application of our main result, we also obtain an O~(ε<sup>−2μ</sup>d)\tilde O(\varepsilon<sup>{-2}\mu</sup> d) sensitivity sampling bound for logistic regression, where μ\mu is a natural complexity measure for this problem. This improves over the previous O~(ε<sup>−2μ<sup>2</sup></sup>d)\tilde O(\varepsilon<sup>{-2}\mu<sup>2</sup></sup> d) bound of Mai et al. (2021) which was based on Lewis weights subsampling.

Citations (3)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.

Tweets

Sign up for free to view the 1 tweet with 0 likes about this paper.