Pith. sign in

REVIEW 5 cited by

Penalized Proximal Policy Optimization for Safe Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2205.11814 v2 pith:G3IYBP4O submitted 2022-05-24 cs.LG cs.AImath.OC

classification cs.LGcs.AImath.OC
keywords policyconstraintalgorithmsconstrainedconstraintslearningoptimizationpenalized
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Safe reinforcement learning aims to learn the optimal policy while satisfying safety constraints, which is essential in real-world applications. However, current algorithms still struggle for efficient policy updates with hard constraint satisfaction. In this paper, we propose Penalized Proximal Policy Optimization (P3O), which solves the cumbersome constrained policy iteration via a single minimization of an equivalent unconstrained problem. Specifically, P3O utilizes a simple-yet-effective penalty function to eliminate cost constraints and removes the trust-region constraint by the clipped surrogate objective. We theoretically prove the exactness of the proposed method with a finite penalty factor and provide a worst-case analysis for approximate error when evaluated on sample trajectories. Moreover, we extend P3O to more challenging multi-constraint and multi-agent scenarios which are less studied in previous work. Extensive experiments show that P3O outperforms state-of-the-art algorithms with respect to both reward improvement and constraint satisfaction on a set of constrained locomotive tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. End-to-End Humanoid Robot Safe and Comfortable Locomotion Policy

    cs.RO 2025-08 conditional novelty 5.0 of 10

    An end-to-end humanoid locomotion policy maps raw LiDAR point clouds to motor commands using P3O with CBF-inspired safety costs and comfort rewards, with sim-to-real tests on a Unitree G1.

  2. Fairness Aware Reinforcement Learning via Proximal Policy Optimization

    cs.MA 2025-02 conditional novelty 5.0 of 10

    Adding retrospective and prospective reward-disparity penalties to PPO lowers demographic parity and conditional statistical parity disparities in two multi-agent simulations, at a measurable efficiency cost.

  3. Action Mapping for Reinforcement Learning in Continuous Environments with Constraints

    cs.LG 2024-12 conditional novelty 5.0 of 10

    Decoupling feasibility from objective optimization by training the RL policy over latent actions that map to feasible actions improves sample efficiency and constraint satisfaction in continuous constrained RL.

  4. FAWAC: Feasibility Informed Advantage Weighted Regression for Persistent Safety in Offline Reinforcement Learning

    cs.LG 2024-12 reject novelty 4.0 of 10

    FAWAC adds a cost-advantage penalty to advantage weighted regression to keep offline-trained policies within a safety budget, with variants for standard and high-reward-but-unsafe datasets.

  5. RSL-RL: A Learning Library for Robotics Research

    cs.RO 2025-09 conditional novelty 3.0 of 10

    RSL-RL is a compact, GPU-accelerated open-source RL library for robotics, providing PPO, DAgger-style behavior cloning, and auxiliary techniques in an easily modifiable codebase.

Pith tools