Papers
Topics
Authors
Recent
Search
2000 character limit reached

Kernel-Based Reinforcement Learning: A Finite-Time Analysis

Published 12 Apr 2020 in cs.LG and stat.ML | (2004.05599v3)

Abstract: We consider the exploration-exploitation dilemma in finite-horizon reinforcement learning problems whose state-action space is endowed with a metric. We introduce Kernel-UCBVI, a model-based optimistic algorithm that leverages the smoothness of the MDP and a non-parametric kernel estimator of the rewards and transitions to efficiently balance exploration and exploitation. For problems with KK episodes and horizon HH, we provide a regret bound of O~(H<sup>3</sup>K<sup>2d2d+1)\widetilde{O}\left( H<sup>3</sup> K<sup>{\frac{2d}{2d+1}}\right), where dd is the covering dimension of the joint state-action space. This is the first regret bound for kernel-based RL using smoothing kernels, which requires very weak assumptions on the MDP and has been previously applied to a wide range of tasks. We empirically validate our approach in continuous MDPs with sparse rewards.

Citations (18)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.