Papers
Topics
Authors
Recent
Search
2000 character limit reached

Distributionally Robust Offline Reinforcement Learning with Linear Function Approximation

Published 14 Sep 2022 in cs.LG, cs.AI, and stat.ML | (2209.06620v3)

Abstract: Among the reasons hindering reinforcement learning (RL) applications to real-world problems, two factors are critical: limited data and the mismatch between the testing environment (real environment in which the policy is deployed) and the training environment (e.g., a simulator). This paper attempts to address these issues simultaneously with distributionally robust offline RL, where we learn a distributionally robust policy using historical data obtained from the source environment by optimizing against a worst-case perturbation thereof. In particular, we move beyond tabular settings and consider linear function approximation. More specifically, we consider two settings, one where the dataset is well-explored and the other where the dataset has sufficient coverage of the optimal policy. We propose two algorithms~-- one for each of the two settings~-- that achieve error bounds O~(d<sup>1/2/N<sup>1/2)\tilde{O}(d<sup>{1/2}/N<sup>{1/2}) and O~(d<sup>3/2/N<sup>1/2)\tilde{O}(d<sup>{3/2}/N<sup>{1/2}) respectively, where dd is the dimension in the linear function approximation and NN is the number of trajectories in the dataset. To the best of our knowledge, they provide the first non-asymptotic results of the sample complexity in this setting. Diverse experiments are conducted to demonstrate our theoretical findings, showing the superiority of our algorithm against the non-robust one.

Citations (16)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.