Pith. sign in

REVIEW 3 cited by

Transductive Off-policy Proximal Policy Optimization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.03894 v1 pith:4YLGKXFH submitted 2024-06-06 cs.LG

classification cs.LG
keywords off-policydatapolicyboundimprovementnoveloptimizationpolicies
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Proximal Policy Optimization (PPO) is a popular model-free reinforcement learning algorithm, esteemed for its simplicity and efficacy. However, due to its inherent on-policy nature, its proficiency in harnessing data from disparate policies is constrained. This paper introduces a novel off-policy extension to the original PPO method, christened Transductive Off-policy PPO (ToPPO). Herein, we provide theoretical justification for incorporating off-policy data in PPO training and prudent guidelines for its safe application. Our contribution includes a novel formulation of the policy improvement lower bound for prospective policies derived from off-policy data, accompanied by a computationally efficient mechanism to optimize this bound, underpinned by assurances of monotonic improvement. Comprehensive experimental results across six representative tasks underscore ToPPO's promising performance.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Revisiting Group Relative Policy Optimization: Insights into On-Policy and Off-Policy Training

    cs.LG 2025-05 conditional novelty 5.0 of 10

    Off-policy GRPO, which estimates advantages from a slightly stale policy, matches or improves on-policy GRPO on math reasoning tasks and is supported by a new policy-improvement lower bound.

  2. Explore or Converge? Stage-Guided Per-Step Optimization for Diffusion Models

    cs.CV 2026-08 conditional novelty 4.0 of 10

    SGPO is a stage-aware RL fine-tuning method for diffusion models that assigns a different optimization objective to each denoising stage, reducing reward hacking and improving quality, diversity, and convergence speed.

  3. Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges

    cs.AI 2025-07 reject novelty 1.0 of 10

    A broad survey of LLM alignment that catalogs objectives, benchmarks, SFT/RLHF/DPO methods, and safety challenges, without contributing new experimental or theoretical results.

Pith tools