REVIEW 3 cited by
Conservative Safety Critics for Exploration
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Safe exploration presents a major challenge in reinforcement learning (RL): when active data collection requires deploying partially trained policies, we must ensure that these policies avoid catastrophically unsafe regions, while still enabling trial and error learning. In this paper, we target the problem of safe exploration in RL by learning a conservative safety estimate of environment states through a critic, and provably upper bound the likelihood of catastrophic failures at every training iteration. We theoretically characterize the tradeoff between safety and policy improvement, show that the safety constraints are likely to be satisfied with high probability during training, derive provable convergence guarantees for our approach, which is no worse asymptotically than standard RL, and demonstrate the efficacy of the proposed approach on a suite of challenging navigation, manipulation, and locomotion tasks. Empirically, we show that the proposed approach can achieve competitive task performance while incurring significantly lower catastrophic failure rates during training than prior methods. Videos are at this url https://sites.google.com/view/conservative-safety-critics/home
Forward citations
Cited by 3 Pith papers
-
Optimistic Exploration for Risk-Averse Constrained Reinforcement Learning
ORAC combines upper-confidence-bound reward exploration with lower-confidence-bound risk-averse cost constraints and adaptive cost weighting to improve exploration in risk-averse constrained RL.
-
Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving
C-HAC combines human demonstrations and reward-based RL for driving, using distributional return estimates to decide when the agent should follow the human-guided policy versus its self-learned policy.
-
Safe and Performant Deployment of Autonomous Systems via Model Predictive Control and Hamilton-Jacobi Reachability Analysis
Adding a Hamilton-Jacobi reachability safety value as a terminal constraint in model predictive control makes the controller recursively feasible and reduces safety violations in car and robot arm simulations.
Discussion (0). Sign in to comment.