Papers
Topics
Authors
Recent
Search
2000 character limit reached

Towards Minimax Optimality of Model-based Robust Reinforcement Learning

Published 10 Feb 2023 in cs.LG and stat.ML | (2302.05372v3)

Abstract: We study the sample complexity of obtaining an ϵ\epsilon-optimal policy in \emph{Robust} discounted Markov Decision Processes (RMDPs), given only access to a generative model of the nominal kernel. This problem is widely studied in the non-robust case, and it is known that any planning approach applied to an empirical MDP estimated with O~(H<sup>3</sup>∣S∣∣A∣ϵ<sup>2)\tilde{\mathcal{O}}(\frac{H<sup>3</sup> \mid S \mid\mid A \mid}{\epsilon<sup>2}) samples provides an ϵ\epsilon-optimal policy, which is minimax optimal. Results in the robust case are much more scarce. For sasa- (resp ss-)rectangular uncertainty sets, the best known sample complexity is O~(H<sup>4</sup>∣S∣<sup>2∣</sup>A∣ϵ<sup>2)\tilde{\mathcal{O}}(\frac{H<sup>4</sup> \mid S \mid<sup>2\mid</sup> A \mid}{\epsilon<sup>2}) (resp. O~(H<sup>4</sup>∣S∣<sup>2∣</sup>A∣<sup>2ϵ<sup>2)\tilde{\mathcal{O}}(\frac{H<sup>4</sup> \mid S \mid<sup>2\mid</sup> A \mid<sup>2}{\epsilon<sup>2})), for specific algorithms and when the uncertainty set is based on the total variation (TV), the KL or the Chi-square divergences. In this paper, we consider uncertainty sets defined with an LpL_p-ball (recovering the TV case), and study the sample complexity of \emph{any} planning algorithm (with high accuracy guarantee on the solution) applied to an empirical RMDP estimated using the generative model. In the general case, we prove a sample complexity of O~(H<sup>4</sup>∣S∣∣A∣ϵ<sup>2)\tilde{\mathcal{O}}(\frac{H<sup>4</sup> \mid S \mid\mid A \mid}{\epsilon<sup>2}) for both the sasa- and ss-rectangular cases (improvements of ∣S∣\mid S \mid and ∣S∣∣A∣\mid S \mid\mid A \mid respectively). When the size of the uncertainty is small enough, we improve the sample complexity to O~(H<sup>3</sup>∣S∣∣A∣ϵ<sup>2)\tilde{\mathcal{O}}(\frac{H<sup>3</sup> \mid S \mid\mid A \mid }{\epsilon<sup>2}), recovering the lower-bound for the non-robust case for the first time and a robust lower-bound when the size of the uncertainty is small enough.

Citations (11)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.

Tweets

Sign up for free to view the 1 tweet with 1 like about this paper.