Papers
Topics
Authors
Recent
Search
2000 character limit reached

Breaking the Sample Complexity Barrier to Regret-Optimal Model-Free Reinforcement Learning

Published 9 Oct 2021 in cs.LG, math.ST, stat.ML, and stat.TH | (2110.04645v2)

Abstract: Achieving sample efficiency in online episodic reinforcement learning (RL) requires optimally balancing exploration and exploitation. When it comes to a finite-horizon episodic Markov decision process with SS states, AA actions and horizon length HH, substantial progress has been achieved towards characterizing the minimax-optimal regret, which scales on the order of H<sup>2SAT\sqrt{H<sup>2SAT} (modulo log factors) with TT the total number of samples. While several competing solution paradigms have been proposed to minimize regret, they are either memory-inefficient, or fall short of optimality unless the sample size exceeds an enormous threshold (e.g., S<sup>6A<sup>4</sup></sup> poly(H)S<sup>6A<sup>4</sup></sup> \,\mathrm{poly}(H) for existing model-free methods). To overcome such a large sample size barrier to efficient RL, we design a novel model-free algorithm, with space complexity O(SAH)O(SAH), that achieves near-optimal regret as soon as the sample size exceeds the order of SA poly(H)SA\,\mathrm{poly}(H). In terms of this sample size requirement (also referred to the initial burn-in cost), our method improves -- by at least a factor of S<sup>5A<sup>3S<sup>5A<sup>3 -- upon any prior memory-efficient algorithm that is asymptotically regret-optimal. Leveraging the recently introduced variance reduction strategy (also called {\em reference-advantage decomposition}), the proposed algorithm employs an {\em early-settled} reference update rule, with the aid of two Q-learning sequences with upper and lower confidence bounds. The design principle of our early-settled variance reduction method might be of independent interest to other RL settings that involve intricate exploration-exploitation trade-offs.

Authors (4)
Citations (47)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.