Papers
Topics
Authors
Recent
Search
2000 character limit reached

Smooth Non-Stationary Bandits

Published 29 Jan 2023 in cs.LG, cs.AI, math.OC, math.ST, and stat.TH | (2301.12366v3)

Abstract: In many applications of online decision making, the environment is non-stationary and it is therefore crucial to use bandit algorithms that handle changes. Most existing approaches are designed to protect against non-smooth changes, constrained only by total variation or Lipschitzness over time. However, in practice, environments often change {\em smoothly}, so such algorithms may incur higher-than-necessary regret. We study a non-stationary bandits problem where each arm's mean reward sequence can be embedded into a β\beta-H\"older function, i.e., a function that is (β−1)(\beta-1)-times Lipschitz-continuously differentiable. The non-stationarity becomes more smooth as β\beta increases. When β=1\beta=1, this corresponds to the non-smooth regime, where \cite{besbes2014stochastic} established a minimax regret of Θ~(T<sup>2/3)\tilde \Theta(T<sup>{2/3}). We show the first separation between the smooth (i.e., β≥2\beta\ge 2) and non-smooth (i.e., β=1\beta=1) regimes by presenting a policy with O~(k<sup>4/5</sup>T<sup>3/5)\tilde O(k<sup>{4/5}</sup> T<sup>{3/5}) regret on any kk-armed, $2$-H\"older instance. We complement this result by showing that the minimax regret on the β\beta-H\"older family of instances is Ω(T<sup>(β+1)/(2β+1))\Omega(T<sup>{(\beta+1)/(2\beta+1)}) for any integer β≥1\beta\ge 1. This matches our upper bound for β=2\beta=2 up to logarithmic factors. Furthermore, we validated the effectiveness of our policy through a comprehensive numerical study using real-world click-through rate data.

Citations (6)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.