Recurrent Natural Policy Gradient for POMDPs (2405.18221v1)

Published 28 May 2024 in math.OC, cs.LG, and stat.ML

Abstract: In this paper, we study a natural policy gradient method based on recurrent neural networks (RNNs) for partially-observable Markov decision processes, whereby RNNs are used for policy parameterization and policy evaluation to address curse of dimensionality in non-Markovian reinforcement learning. We present finite-time and finite-width analyses for both the critic (recurrent temporal difference learning), and correspondingly-operated recurrent natural policy gradient method in the near-initialization regime. Our analysis demonstrates the efficiency of RNNs for problems with short-term memory with explicit bounds on the required network widths and sample complexity, and points out the challenges in the case of long-term dependencies.

References (1)

M. Telgarsky, Deep learning theory lecture notes. https://mjt.cs.illinois.edu/dlt/, 2021. Version: 2021-10-27 v0.0-e7150f2d (alpha).

Summary

We haven't generated a summary for this paper yet.

Summarize Now

Recurrent Natural Policy Gradient for POMDPs (2405.18221v1)

Summary

Related Papers

Tweets