Pith. sign in

REVIEW 1 cited by

Responsive Safety in Reinforcement Learning by PID Lagrangian Methods

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2007.03964 v1 pith:7LLK7SN3 submitted 2020-07-08 math.OC cs.AIcs.LG

classification math.OCcs.AIcs.LG
keywords lagrangianlearningemphmethodsalgorithmscontroldynamicsintroduce
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Lagrangian methods are widely used algorithms for constrained optimization problems, but their learning dynamics exhibit oscillations and overshoot which, when applied to safe reinforcement learning, leads to constraint-violating behavior during agent training. We address this shortcoming by proposing a novel Lagrange multiplier update method that utilizes derivatives of the constraint function. We take a controls perspective, wherein the traditional Lagrange multiplier update behaves as \emph{integral} control; our terms introduce \emph{proportional} and \emph{derivative} control, achieving favorable learning dynamics through damping and predictive measures. We apply our PID Lagrangian methods in deep RL, setting a new state of the art in Safety Gym, a safe RL benchmark. Lastly, we introduce a new method to ease controller tuning by providing invariance to the relative numerical scales of reward and cost. Our extensive experiments demonstrate improved performance and hyperparameter robustness, while our algorithms remain nearly as simple to derive and implement as the traditional Lagrangian approach.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Central Path Proximal Policy Optimization

    cs.LG 2025-05 conditional novelty 4.0 of 10

    C3PO augments the PPO loss with a receding ReLU penalty on the cost advantage, approximating C-TRPO's central path and improving reward-constraint trade-offs in Safety Gymnasium tasks.

Pith tools