Papers
Topics
Authors
Recent
Search
2000 character limit reached

Posterior sampling for reinforcement learning: worst-case regret bounds

Published 19 May 2017 in cs.LG | (1705.07041v3)

Abstract: We present an algorithm based on posterior sampling (aka Thompson sampling) that achieves near-optimal worst-case regret bounds when the underlying Markov Decision Process (MDP) is communicating with a finite, though unknown, diameter. Our main result is a high probability regret upper bound of O~(DSAT)\tilde{O}(DS\sqrt{AT}) for any communicating MDP with SS states, AA actions and diameter DD. Here, regret compares the total reward achieved by the algorithm to the total expected reward of an optimal infinite-horizon undiscounted average reward policy, in time horizon TT. This result closely matches the known lower bound of Ω(DSAT)\Omega(\sqrt{DSAT}). Our techniques involve proving some novel results about the anti-concentration of Dirichlet distribution, which may be of independent interest.

Authors (2)
Citations (35)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.