C3PO augments the PPO loss with a receding ReLU penalty on the cost advantage, approximating C-TRPO's central path and improving reward-constraint trade-offs in Safety Gymnasium tasks.
Yiming Zhang, Quan Vuong, and Keith Ross
1 Pith paper cite this work, alongside 12 external citations. Polarity classification is still indexing.
1
Pith paper citing it
12
external citations · OpenAlex
citation-role summary
background 1
citation-polarity summary
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Central Path Proximal Policy Optimization
C3PO augments the PPO loss with a receding ReLU penalty on the cost advantage, approximating C-TRPO's central path and improving reward-constraint trade-offs in Safety Gymnasium tasks.