Pith. sign in

REVIEW 2 cited by

Embedding Safety into RL: A New Take on Trust Region Methods

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.02957 v4 pith:MLYUAHER submitted 2024-11-05 cs.LG cs.SYeess.SY

classification cs.LGcs.SYeess.SY
keywords policyconstrainedtrustc-trpoconstraintmethodsoptimizationregion
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Reinforcement Learning (RL) agents can solve diverse tasks but often exhibit unsafe behavior. Constrained Markov Decision Processes (CMDPs) address this by enforcing safety constraints, yet existing methods either sacrifice reward maximization or allow unsafe training. We introduce Constrained Trust Region Policy Optimization (C-TRPO), which reshapes the policy space geometry to ensure trust regions contain only safe policies, guaranteeing constraint satisfaction throughout training. We analyze its theoretical properties and connections to TRPO, Natural Policy Gradient (NPG), and Constrained Policy Optimization (CPO). Experiments show that C-TRPO reduces constraint violations while maintaining competitive returns.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The Geometry of Nonlinear Reinforcement Learning

    cs.LG 2025-09 conditional novelty 5.0 of 10

    Actor-critic reinforcement learning methods are reformulated as mirror descent on the occupancy manifold, and a Hessian-based update is proposed for nonlinear and constrained objectives.

  2. Central Path Proximal Policy Optimization

    cs.LG 2025-05 conditional novelty 4.0 of 10

    C3PO augments the PPO loss with a receding ReLU penalty on the cost advantage, approximating C-TRPO's central path and improving reward-constraint trade-offs in Safety Gymnasium tasks.

Pith tools