C3PO augments the PPO loss with a receding ReLU penalty on the cost advantage, approximating C-TRPO's central path and improving reward-constraint trade-offs in Safety Gymnasium tasks.
Dimitri P Bertsekas
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Central Path Proximal Policy Optimization
C3PO augments the PPO loss with a receding ReLU penalty on the cost advantage, approximating C-TRPO's central path and improving reward-constraint trade-offs in Safety Gymnasium tasks.