Pith. sign in

REVIEW 1 cited by

Verified Safe Reinforcement Learning for Neural Network Dynamic Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.15994 v2 pith:MPAZKAFH submitted 2024-05-25 cs.LG cs.AI

classification cs.LGcs.AI
keywords safeverifiedlearningcontrolsafetyachieveapproachcontroller
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Learning reliably safe autonomous control is one of the core problems in trustworthy autonomy. However, training a controller that can be formally verified to be safe remains a major challenge. We introduce a novel approach for learning verified safe control policies in nonlinear neural dynamical systems while maximizing overall performance. Our approach aims to achieve safety in the sense of finite-horizon reachability proofs, and is comprised of three key parts. The first is a novel curriculum learning scheme that iteratively increases the verified safe horizon. The second leverages the iterative nature of gradient-based learning to leverage incremental verification, reusing information from prior verification runs. Finally, we learn multiple verified initial-state-dependent controllers, an idea that is especially valuable for more complex domains where learning a single universal verified safe controller is extremely challenging. Our experiments on five safe control problems demonstrate that our trained controllers can achieve verified safety over horizons that are as much as an order of magnitude longer than state-of-the-art baselines, while maintaining high reward, as well as a perfect safety record over entire episodes. Our code is available at https://github.com/jlwu002/VSRL.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. BaB-ND: Long-Horizon Motion Planning with Branch-and-Bound and Neural Dynamics

    cs.RO 2024-12 conditional novelty 6.0 of 10

    A GPU-accelerated branch-and-bound planner over neural dynamics models uses adapted CROWN bounds to prune subdomains and beat sampling-based and MIP baselines on long-horizon manipulation tasks.

Pith tools