Proximal Policy Distillation (2407.15134v2)

Published 21 Jul 2024 in cs.LG and cs.AI

Abstract: We introduce Proximal Policy Distillation (PPD), a novel policy distillation method that integrates student-driven distillation and Proximal Policy Optimization (PPO) to increase sample efficiency and to leverage the additional rewards that the student policy collects during distillation. To assess the efficacy of our method, we compare PPD with two common alternatives, student-distill and teacher-distill, over a wide range of reinforcement learning environments that include discrete actions and continuous control (ATARI, Mujoco, and Procgen). For each environment and method, we perform distillation to a set of target student neural networks that are smaller, identical (self-distillation), or larger than the teacher network. Our findings indicate that PPD improves sample efficiency and produces better student policies compared to typical policy distillation approaches. Moreover, PPD demonstrates greater robustness than alternative methods when distilling policies from imperfect demonstrations. The code for the paper is released as part of a new Python library built on top of stable-baselines3 to facilitate policy distillation: `sb3-distill'.

Summary

We haven't generated a summary for this paper yet.

Summarize Now

Related Papers

Dual Policy Distillation (2020)
Generative Adversarial Simulator (2020)
Real-time Policy Distillation in Deep Reinforcement Learning (2019)
Distillation Strategies for Proximal Policy Optimization (2019)
Online Policy Distillation with Decision-Attention (2024)

Proximal Policy Distillation (2407.15134v2)

Summary

Related Papers

Tweets