Papers
Topics
Authors
Recent
Search
2000 character limit reached

Improved No-Regret Algorithms for Stochastic Shortest Path with Linear MDP

Published 18 Dec 2021 in cs.LG | (2112.09859v1)

Abstract: We introduce two new no-regret algorithms for the stochastic shortest path (SSP) problem with a linear MDP that significantly improve over the only existing results of (Vial et al., 2021). Our first algorithm is computationally efficient and achieves a regret bound O~(d<sup>3B<sup>2T</sup></sup>K)\widetilde{O}\left(\sqrt{d<sup>3B_{\star}<sup>2T_{\star}</sup></sup> K}\right), where dd is the dimension of the feature space, BB_{\star} and TT_{\star} are upper bounds of the expected costs and hitting time of the optimal policy respectively, and KK is the number of episodes. The same algorithm with a slight modification also achieves logarithmic regret of order O(d<sup>3B<sup>4cmin<sup>2gap<em>minln<sup>5dB</sup></em></sup></sup></sup>Kcmin)O\left(\frac{d<sup>3B_{\star}<sup>4}{c_{\min}<sup>2\text{gap}<em>{\min}}\ln<sup>5\frac{dB</sup></em>{\star}</sup></sup></sup> K}{c_{\min}} \right), where gap<em>min\text{gap}<em>{\min} is the minimum sub-optimality gap and c</em>minc</em>{\min} is the minimum cost over all state-action pairs. Our result is obtained by developing a simpler and improved analysis for the finite-horizon approximation of (Cohen et al., 2021) with a smaller approximation error, which might be of independent interest. On the other hand, using variance-aware confidence sets in a global optimization problem, our second algorithm is computationally inefficient but achieves the first "horizon-free" regret bound O~(d<sup>3.5BK)\widetilde{O}(d<sup>{3.5}B_{\star}\sqrt{K}) with no polynomial dependency on TT_{\star} or 1/cmin1/c_{\min}, almost matching the Ω(dBK)\Omega(dB_{\star}\sqrt{K}) lower bound from (Min et al., 2021).

Authors (3)
Citations (14)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.