Papers
Topics
Authors
Recent
Search
2000 character limit reached

Variance-aware robust reinforcement learning with linear function approximation under heavy-tailed rewards

Published 9 Mar 2023 in cs.LG, cs.AI, math.ST, stat.ML, and stat.TH | (2303.05606v2)

Abstract: This paper presents two algorithms, AdaOFUL and VARA, for online sequential decision-making in the presence of heavy-tailed rewards with only finite variances. For linear stochastic bandits, we address the issue of heavy-tailed rewards by modifying the adaptive Huber regression and proposing AdaOFUL. AdaOFUL achieves a state-of-the-art regret bound of O~(d(∑t=1<sup>T</sup>νt<sup>2)<sup>1/2+d)\widetilde{O}\big(d\big(\sum_{t=1}<sup>T</sup> \nu_{t}<sup>2\big)<sup>{1/2}+d\big) as if the rewards were uniformly bounded, where νt<sup>2\nu_{t}<sup>2 is the observed conditional variance of the reward at round tt, dd is the feature dimension, and O~(⋅)\widetilde{O}(\cdot) hides logarithmic dependence. Building upon AdaOFUL, we propose VARA for linear MDPs, which achieves a tighter variance-aware regret bound of O~(dHG<sup>∗K)\widetilde{O}(d\sqrt{HG<sup>*K}). Here, HH is the length of episodes, KK is the number of episodes, and G<sup>∗G<sup>* is a smaller instance-dependent quantity that can be bounded by other instance-dependent quantities when additional structural conditions on the MDP are satisfied. Our regret bound is superior to the current state-of-the-art bounds in three ways: (1) it depends on a tighter instance-dependent quantity and has optimal dependence on dd and HH, (2) we can obtain further instance-dependent bounds of G<sup>∗G<sup>* under additional structural conditions on the MDP, and (3) our regret bound is valid even when rewards have only finite variances, achieving a level of generality unmatched by previous works. Overall, our modified adaptive Huber regression algorithm may serve as a useful building block in the design of algorithms for online problems with heavy-tailed rewards.

Authors (2)
Citations (7)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.