Papers
Topics
Authors
Recent
Search
2000 character limit reached

Model selection for contextual bandits

Published 3 Jun 2019 in cs.LG, math.ST, stat.ML, and stat.TH | (1906.00531v3)

Abstract: We introduce the problem of model selection for contextual bandits, where a learner must adapt to the complexity of the optimal policy while balancing exploration and exploitation. Our main result is a new model selection guarantee for linear contextual bandits. We work in the stochastic realizable setting with a sequence of nested linear policy classes of dimension $d_1 &lt; d_2 &lt; \ldots$, where the m<sup>⋆m<sup>\star-th class contains the optimal policy, and we design an algorithm that achieves O~(T<sup>2/3d<sup>1/3m<sup>⋆)\tilde{O}(T<sup>{2/3}d<sup>{1/3}_{m<sup>\star}) regret with no prior knowledge of the optimal dimension dm<sup>⋆d_{m<sup>\star}. The algorithm also achieves regret O~(T<sup>3/4</sup>+Tdm<sup>⋆)\tilde{O}(T<sup>{3/4}</sup> + \sqrt{Td_{m<sup>\star}}), which is optimal for dm<sup>⋆≥Td_{m<sup>{\star}}\geq{}\sqrt{T}. This is the first model selection result for contextual bandits with non-vacuous regret for all values of dm<sup>⋆d_{m<sup>\star}, and to the best of our knowledge is the first positive result of this type for any online learning setting with partial information. The core of the algorithm is a new estimator for the gap in the best loss achievable by two linear policy classes, which we show admits a convergence rate faster than the rate required to learn the parameters for either class.

Citations (88)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.