REVIEW 3 cited by
Transductive Off-policy Proximal Policy Optimization
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Proximal Policy Optimization (PPO) is a popular model-free reinforcement learning algorithm, esteemed for its simplicity and efficacy. However, due to its inherent on-policy nature, its proficiency in harnessing data from disparate policies is constrained. This paper introduces a novel off-policy extension to the original PPO method, christened Transductive Off-policy PPO (ToPPO). Herein, we provide theoretical justification for incorporating off-policy data in PPO training and prudent guidelines for its safe application. Our contribution includes a novel formulation of the policy improvement lower bound for prospective policies derived from off-policy data, accompanied by a computationally efficient mechanism to optimize this bound, underpinned by assurances of monotonic improvement. Comprehensive experimental results across six representative tasks underscore ToPPO's promising performance.
Forward citations
Cited by 3 Pith papers
-
Revisiting Group Relative Policy Optimization: Insights into On-Policy and Off-Policy Training
Off-policy GRPO, which estimates advantages from a slightly stale policy, matches or improves on-policy GRPO on math reasoning tasks and is supported by a new policy-improvement lower bound.
-
Explore or Converge? Stage-Guided Per-Step Optimization for Diffusion Models
SGPO is a stage-aware RL fine-tuning method for diffusion models that assigns a different optimization objective to each denoising stage, reducing reward hacking and improving quality, diversity, and convergence speed.
-
Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges
A broad survey of LLM alignment that catalogs objectives, benchmarks, SFT/RLHF/DPO methods, and safety challenges, without contributing new experimental or theoretical results.
Discussion (0). Continue with ORCID to comment.