Papers
Topics
Authors
Recent
Search
2000 character limit reached

Truncated Variance Reduced Value Iteration

Published 21 May 2024 in cs.LG, cs.DS, and math.OC | (2405.12952v1)

Abstract: We provide faster randomized algorithms for computing an ϵ\epsilon-optimal policy in a discounted Markov decision process with AtotA_{\text{tot}}-state-action pairs, bounded rewards, and discount factor γ\gamma. We provide an O~(Atot[(1−γ)<sup>−3ϵ<sup>−2</sup></sup>+(1−γ)<sup>−2])\tilde{O}(A_{\text{tot}}[(1 - \gamma)<sup>{-3}\epsilon<sup>{-2}</sup></sup> + (1 - \gamma)<sup>{-2}])-time algorithm in the sampling setting, where the probability transition matrix is unknown but accessible through a generative model which can be queried in O~(1)\tilde{O}(1)-time, and an O~(s+(1−γ)<sup>−2)\tilde{O}(s + (1-\gamma)<sup>{-2})-time algorithm in the offline setting where the probability transition matrix is known and ss-sparse. These results improve upon the prior state-of-the-art which either ran in O~(Atot[(1−γ)<sup>−3ϵ<sup>−2</sup></sup>+(1−γ)<sup>−3])\tilde{O}(A_{\text{tot}}[(1 - \gamma)<sup>{-3}\epsilon<sup>{-2}</sup></sup> + (1 - \gamma)<sup>{-3}]) time [Sidford, Wang, Wu, Ye 2018] in the sampling setting, O~(s+Atot(1−γ)<sup>−3)\tilde{O}(s + A_{\text{tot}} (1-\gamma)<sup>{-3}) time [Sidford, Wang, Wu, Yang, Ye 2018] in the offline setting, or time at least quadratic in the number of states using interior point methods for linear programming. We achieve our results by building upon prior stochastic variance-reduced value iteration methods [Sidford, Wang, Wu, Yang, Ye 2018]. We provide a variant that carefully truncates the progress of its iterates to improve the variance of new variance-reduced sampling procedures that we introduce to implement the steps. Our method is essentially model-free and can be implemented in O~(Atot)\tilde{O}(A_{\text{tot}})-space when given generative model access. Consequently, our results take a step in closing the sample-complexity gap between model-free and model-based methods.

Citations (1)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.

Tweets

Sign up for free to view the 2 tweets with 1 like about this paper.