Papers
Topics
Authors
Recent
Search
2000 character limit reached

Asymptotically Efficient Off-Policy Evaluation for Tabular Reinforcement Learning

Published 29 Jan 2020 in cs.LG, cs.AI, and stat.ML | (2001.10742v1)

Abstract: We consider the problem of off-policy evaluation for reinforcement learning, where the goal is to estimate the expected reward of a target policy π\pi using offline data collected by running a logging policy μ\mu. Standard importance-sampling based approaches for this problem suffer from a variance that scales exponentially with time horizon HH, which motivates a splurge of recent interest in alternatives that break the "Curse of Horizon" (Liu et al. 2018, Xie et al. 2019). In particular, it was shown that a marginalized importance sampling (MIS) approach can be used to achieve an estimation error of order O(H<sup>3/</sup>n)O(H<sup>3/</sup> n) in mean square error (MSE) under an episodic Markov Decision Process model with finite states and potentially infinite actions. The MSE bound however is still a factor of HH away from a Cramer-Rao lower bound of order Ω(H<sup>2/n)\Omega(H<sup>2/n). In this paper, we prove that with a simple modification to the MIS estimator, we can asymptotically attain the Cramer-Rao lower bound, provided that the action space is finite. We also provide a general method for constructing MIS estimators with high-probability error bounds.

Authors (2)
Citations (77)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.