Papers
Topics
Authors
Recent
Search
2000 character limit reached

Generic Coreset for Scalable Learning of Monotonic Kernels: Logistic Regression, Sigmoid and more

Published 21 Feb 2018 in cs.LG and cs.DS | (1802.07382v3)

Abstract: Coreset (or core-set) is a small weighted \emph{subset} QQ of an input set PP with respect to a given \emph{monotonic} function f:R→Rf:\mathbb{R}\to\mathbb{R} that \emph{provably} approximates its fitting loss ∑p∈Pf(p⋅x)\sum_{p\in P}f(p\cdot x) to \emph{any} given x∈R<sup>dx\in\mathbb{R}<sup>d. Using QQ we can obtain approximation of x<sup>∗x<sup>* that minimizes this loss, by running \emph{existing} optimization algorithms on QQ. In this work we provide: (i) A lower bound which proves that there are sets with no coresets smaller than n=∣P∣n=|P| for general monotonic loss functions. (ii) A proof that, under a natural assumption that holds e.g. for logistic regression and the sigmoid activation functions, a small coreset exists for \emph{any} input PP. (iii) A generic coreset construction algorithm that computes such a small coreset QQ in O(nd+nlog⁡n)O(nd+n\log n) time, and (iv) Experimental results which demonstrate that our coresets are effective and are much smaller in practice than predicted in theory.

Citations (13)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.