Papers
Topics
Authors
Recent
Search
2000 character limit reached

Efficient Contextual Bandits in Non-stationary Worlds

Published 5 Aug 2017 in cs.LG and stat.ML | (1708.01799v4)

Abstract: Most contextual bandit algorithms minimize regret against the best fixed policy, a questionable benchmark for non-stationary environments that are ubiquitous in applications. In this work, we develop several efficient contextual bandit algorithms for non-stationary environments by equipping existing methods for i.i.d. problems with sophisticated statistical tests so as to dynamically adapt to a change in distribution. We analyze various standard notions of regret suited to non-stationary environments for these algorithms, including interval regret, switching regret, and dynamic regret. When competing with the best policy at each time, one of our algorithms achieves regret O(ST)\mathcal{O}(\sqrt{ST}) if there are TT rounds with SS stationary periods, or more generally O(Δ<sup>1/3T<sup>2/3)\mathcal{O}(\Delta<sup>{1/3}T<sup>{2/3}) where Δ\Delta is some non-stationarity measure. These results almost match the optimal guarantees achieved by an inefficient baseline that is a variant of the classic Exp4 algorithm. The dynamic regret result is also the first one for efficient and fully adversarial contextual bandit. Furthermore, while the results above require tuning a parameter based on the unknown quantity SS or Δ\Delta, we also develop a parameter free algorithm achieving regret minS<sup>1/4T<sup>3/4,</sup></sup>Δ<sup>1/5T<sup>4/5\min{S<sup>{1/4}T<sup>{3/4},</sup></sup> \Delta<sup>{1/5}T<sup>{4/5}}. This improves and generalizes the best existing result Δ<sup>0.18T<sup>0.82\Delta<sup>{0.18}T<sup>{0.82} by Karnin and Anava (2016) which only holds for the two-armed bandit problem.

Citations (125)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.