Papers
Topics
Authors
Recent
Search
2000 character limit reached

Improved Regret Analysis for Variance-Adaptive Linear Bandits and Horizon-Free Linear Mixture MDPs

Published 5 Nov 2021 in stat.ML, cs.LG, math.ST, and stat.TH | (2111.03289v4)

Abstract: In online learning problems, exploiting low variance plays an important role in obtaining tight performance guarantees yet is challenging because variances are often not known a priori. Recently, considerable progress has been made by Zhang et al. (2021) where they obtain a variance-adaptive regret bound for linear bandits without knowledge of the variances and a horizon-free regret bound for linear mixture Markov decision processes (MDPs). In this paper, we present novel analyses that improve their regret bounds significantly. For linear bandits, we achieve O~(mindK,d<sup>1.5k=1<sup>K</sup></sup>σk<sup>2</sup>+d<sup>2)\tilde O(\min{d\sqrt{K}, d<sup>{1.5}\sqrt{\sum_{k=1}<sup>K</sup></sup> \sigma_k<sup>2}}</sup> + d<sup>2) where dd is the dimension of the features, KK is the time horizon, and σk<sup>2\sigma_k<sup>2 is the noise variance at time step kk, and O~\tilde O ignores polylogarithmic dependence, which is a factor of d<sup>3d<sup>3 improvement. For linear mixture MDPs with the assumption of maximum cumulative reward in an episode being in [0,1][0,1], we achieve a horizon-free regret bound of O~(dK+d<sup>2)\tilde O(d \sqrt{K} + d<sup>2) where dd is the number of base models and KK is the number of episodes. This is a factor of d<sup>3.5d<sup>{3.5} improvement in the leading term and d<sup>7d<sup>7 in the lower order term. Our analysis critically relies on a novel peeling-based regret analysis that leverages the elliptical potential `count' lemma.

Citations (30)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.