SSkP uses PU learning to build a skill risk predictor from demonstrations and a cross-entropy risk planner to select safe skills during online RL, outperforming prior methods on most tested MuJoCo tasks.
Constrained Markov decision processes: stochastic modeling
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Skill-based Safe Reinforcement Learning with Risk Planning
SSkP uses PU learning to build a skill risk predictor from demonstrations and a cross-entropy risk planner to select safe skills during online RL, outperforming prior methods on most tested MuJoCo tasks.