REVIEW 8 cited by
Learning to be Safe: Deep RL with a Safety Critic
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Safety is an essential component for deploying reinforcement learning (RL) algorithms in real-world scenarios, and is critical during the learning process itself. A natural first approach toward safe RL is to manually specify constraints on the policy's behavior. However, just as learning has enabled progress in large-scale development of AI systems, learning safety specifications may also be necessary to ensure safety in messy open-world environments where manual safety specifications cannot scale. Akin to how humans learn incrementally starting in child-safe environments, we propose to learn how to be safe in one set of tasks and environments, and then use that learned intuition to constrain future behaviors when learning new, modified tasks. We empirically study this form of safety-constrained transfer learning in three challenging domains: simulated navigation, quadruped locomotion, and dexterous in-hand manipulation. In comparison to standard deep RL techniques and prior approaches to safe RL, we find that our method enables the learning of new tasks and in new environments with both substantially fewer safety incidents, such as falling or dropping an object, and faster, more stable learning. This suggests a path forward not only for safer RL systems, but also for more effective RL systems.
Forward citations
Cited by 8 Pith papers
-
SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions
An unbiased PPO gradient that ignores recovery-policy density, plus analytic recovery values and success-gated imitation, cuts training falls by 26–233× on locomotion tasks without sacrificing reward.
-
ARMOR: Robust Reinforcement Learning-based Control for UAVs under Physical Attacks
A teacher-student latent representation method lets an RL drone controller stay stable under GPS, gyroscope, and other sensor attacks, including attacks never seen in training.
-
SafeMimic: Towards Safe and Autonomous Human-to-Robot Imitation for Mobile Manipulation
SafeMimic enables a mobile robot to safely and autonomously adapt a single third-person human video into a successful multi-step manipulation strategy.
-
Approximating Safety Feedback Without a Safety Oracle via Model Predictive Control
RL-SA VMPC shields an RL policy by planning, via MPPI in a black-box simulator, a path from the next state back to the previous state, aborting when no such path exists.
-
Verifiable Safety Q-Filters via Hamilton-Jacobi Reachability and Multiplicative Q-Networks
Learned Q-function safety filters are certified by verifying two sufficient conditions with a mixed-integer optimizer, using a multiplicative Q-network to prevent safe-set collapse during fine-tuning.
-
Learning Fast, Tool aware Collision Avoidance for Collaborative Robots
A real-time, tool-aware collision avoidance system for cobots that blends a learned perception-safety critic with classical IK, achieving low collision rates in dynamic partially-observed environments.
-
Safe and Performant Controller Synthesis using Gradient-based Model Predictive Control and Control Barrier Functions
A two-stage controller that uses L-BFGS gradient-based MPC for performance and a CBF-QP filter for hard safety constraints is demonstrated on simulated unicycle and planar quadrotor navigation.
-
Safe and Performant Deployment of Autonomous Systems via Model Predictive Control and Hamilton-Jacobi Reachability Analysis
Adding a Hamilton-Jacobi reachability safety value as a terminal constraint in model predictive control makes the controller recursively feasible and reduces safety violations in car and robot arm simulations.
Discussion (0). Continue with ORCID to comment.