REVIEW 2 cited by
Safe Reinforcement Learning Using Robust Control Barrier Functions
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Reinforcement Learning (RL) has been shown to be effective in many scenarios. However, it typically requires the exploration of a sufficiently large number of state-action pairs, some of which may be unsafe. Consequently, its application to safety-critical systems remains a challenge. An increasingly common approach to address safety involves the addition of a safety layer that projects the RL actions onto a safe set of actions. In turn, a difficulty for such frameworks is how to effectively couple RL with the safety layer to improve the learning performance. In this paper, we frame safety as a differentiable robust-control-barrier-function layer in a model-based RL framework. Moreover, we also propose an approach to modularly learn the underlying reward-driven task, independent of safety constraints. We demonstrate that this approach both ensures safety and effectively guides exploration during training in a range of experiments, including zero-shot transfer when the reward is learned in a modular way.
Forward citations
Cited by 2 Pith papers
-
Q-learning-based Model-free Safety Filter
A Q-learning safety filter with a time-dependent reward blocks unsafe actions from arbitrary task policies, but its theoretical guarantee is not valid as written.
-
Enforcing Cooperative Safety for Reinforcement Learning-based Mixed-Autonomy Platoon Control
A safe MARL framework for mixed-autonomy platoons that filters RL actions through a cooperative control barrier function with conformal prediction bounds, improving simulated system-level safety with little efficiency loss.
Discussion (0). Continue with ORCID to comment.