Papers
Topics
Authors
Recent
Search
2000 character limit reached

Thresholded Lasso Bandit

Published 22 Oct 2020 in stat.ML and cs.LG | (2010.11994v4)

Abstract: In this paper, we revisit the regret minimization problem in sparse stochastic contextual linear bandits, where feature vectors may be of large dimension dd, but where the reward function depends on a few, say s0ds_0\ll d, of these features only. We present Thresholded Lasso bandit, an algorithm that (i) estimates the vector defining the reward function as well as its sparse support, i.e., significant feature elements, using the Lasso framework with thresholding, and (ii) selects an arm greedily according to this estimate projected on its support. The algorithm does not require prior knowledge of the sparsity index s0s_0 and can be parameter-free under some symmetric assumptions. For this simple algorithm, we establish non-asymptotic regret upper bounds scaling as O(logd+T)\mathcal{O}( \log d + \sqrt{T} ) in general, and as O(logd+logT)\mathcal{O}( \log d + \log T) under the so-called margin condition (a probabilistic condition on the separation of the arm rewards). The regret of previous algorithms scales as O(logd+Tlog(dT))\mathcal{O}( \log d + \sqrt{T \log (d T)}) and O(logTlogd)\mathcal{O}( \log T \log d) in the two settings, respectively. Through numerical experiments, we confirm that our algorithm outperforms existing methods.

Citations (15)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.