TraC trains an offline safe RL policy by classifying trajectories as desirable (safe, high-reward) versus undesirable (unsafe or low-reward) using a logistic loss on a policy-ratio score.
Additionally, SafetyGymnasium includes five velocity- constrained tasks for the agents, Ant, HalfCheetah, Hopper, Walker2d, and Swimmer
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Offline Safe Reinforcement Learning Using Trajectory Classification
TraC trains an offline safe RL policy by classifying trajectories as desirable (safe, high-reward) versus undesirable (unsafe or low-reward) using a logistic loss on a policy-ratio score.