Pith. sign in

REVIEW 1 cited by

Learning to Provably Satisfy High Relative Degree Constraints for Black-Box Systems

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.20456 v1 pith:KFB3KQM4 submitted 2024-07-29 eess.SY cs.SY

classification eess.SYcs.SY
keywords affineconstraintdegreerelativeblack-boxconstraintshighcontrol
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper, we develop a method for learning a control policy guaranteed to satisfy an affine state constraint of high relative degree in closed loop with a black-box system. Previous reinforcement learning (RL) approaches to satisfy safety constraints either require access to the system model, or assume control affine dynamics, or only discourage violations with reward shaping. Only recently have these issues been addressed with POLICEd RL, which guarantees constraint satisfaction for black-box systems. However, this previous work can only enforce constraints of relative degree 1. To address this gap, we build a novel RL algorithm explicitly designed to enforce an affine state constraint of high relative degree in closed loop with a black-box control system. Our key insight is to make the learned policy be affine around the unsafe set and to use this affine region to dissipate the inertia of the high relative degree constraint. We prove that such policies guarantee constraint satisfaction for deterministic systems while being agnostic to the choice of the RL training algorithm. Our results demonstrate the capacity of our approach to enforce hard constraints in the Gym inverted pendulum and on a space shuttle landing simulation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. mPOLICE: Provable Enforcement of Multi-Region Affine Constraints in Deep Neural Networks

    cs.LG 2025-02 conditional novelty 6.0 of 10

    mPOLICE generalizes POLICE to enforce exact affine output constraints inside multiple disjoint convex regions of a ReLU network's input by giving each region its own activation pattern.

Pith tools