A dynamic safety shield uses an RL supervisor to adaptively weight obstacle-avoidance and action-matching terms in an MPC cost, improving the goals-to-collisions ratio in navigation RL.
Safe reinforcement learning via shielding under partial observability
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.RO 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
A Dynamic Safety Shield for Safe and Efficient Reinforcement Learning of Navigation Tasks
A dynamic safety shield uses an RL supervisor to adaptively weight obstacle-avoidance and action-matching terms in an MPC cost, improving the goals-to-collisions ratio in navigation RL.