Papers
Topics
Authors
Recent
Search
2000 character limit reached

Rate-Optimal Policy Optimization for Linear Markov Decision Processes

Published 28 Aug 2023 in cs.LG | (2308.14642v3)

Abstract: We study regret minimization in online episodic linear Markov Decision Processes, and obtain rate-optimal O~(K)\widetilde O (\sqrt K) regret where KK denotes the number of episodes. Our work is the first to establish the optimal (w.r.t.~KK) rate of convergence in the stochastic setting with bandit feedback using a policy optimization based approach, and the first to establish the optimal (w.r.t.~KK) rate in the adversarial setup with full information feedback, for which no algorithm with an optimal rate guarantee is currently known.

Citations (7)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.

Tweets

Sign up for free to view the 2 tweets with 1 like about this paper.