PPD integrates PPO into policy distillation so the student collects and uses its own rewards, yielding better sample efficiency and robustness than standard student-distill or teacher-distill on ATARI, Mujoco, and Procgen tasks.
Kazuma Tsuji, Ken’ichiro Tanaka, and Sebastian Pokutta
6 Pith papers cite this work. Polarity classification is still indexing.
representative citing papers
Episodic kernel quadrature compresses batches of episodes via GP-modeled returns to enable efficient policy gradient updates without evaluating rewards on every sample.
Neural regression collapse occurs when last-layer feature intrinsic dimension falls below target intrinsic dimension, creating over-compressed and under-compressed regimes that govern generalization based on data quantity and noise.
Gimitest is an open-source tool that decorates RL environment APIs to enable search-based, metamorphic, and adversarial testing of single- and multi-agent policies.
A three-stage curriculum RL policy for end-to-end quadrotor stabilization outperforms single-stage training in sample efficiency and robustness in simulation.
Themis is an XAI-enabled framework for RL from human feedback that supports 200+ environments and includes a scalable cloud platform for collecting human preferences.
citing papers explorer
-
Proximal Policy Distillation
PPD integrates PPO into policy distillation so the student collects and uses its own rewards, yielding better sample efficiency and robustness than standard student-distill or teacher-distill on ATARI, Mujoco, and Procgen tasks.
-
Policy Gradient with Kernel Quadrature
Episodic kernel quadrature compresses batches of episodes via GP-modeled returns to enable efficient policy gradient updates without evaluating rewards on every sample.
-
Geometric Analysis of Neural Regression Collapse via Intrinsic Dimension
Neural regression collapse occurs when last-layer feature intrinsic dimension falls below target intrinsic dimension, creating over-compressed and under-compressed regimes that govern generalization based on data quantity and noise.
-
Gimitest: A Comprehensive Tool for Testing Reinforcement Learning Policies
Gimitest is an open-source tool that decorates RL environment APIs to enable search-based, metamorphic, and adversarial testing of single- and multi-agent policies.
-
Curriculum-based Sample Efficient Reinforcement Learning for Robust Stabilization of a Quadrotor
A three-stage curriculum RL policy for end-to-end quadrotor stabilization outperforms single-stage training in sample efficiency and robustness in simulation.
-
Themis: An explainable AI-enabled framework for Reinforcement Learning with Human Feedback
Themis is an XAI-enabled framework for RL from human feedback that supports 200+ environments and includes a scalable cloud platform for collecting human preferences.