Pith. sign in

REVIEW 1 cited by

POLICEd RL: Learning Closed-Loop Robot Control Policies with Provable Satisfaction of Hard Constraints

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.13297 v3 pith:2R3UBJNA submitted 2024-03-20 cs.RO

classification cs.RO
keywords constraintsconstraintaffinehardpolicedsatisfactionalgorithmclosed-loop
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper, we seek to learn a robot policy guaranteed to satisfy state constraints. To encourage constraint satisfaction, existing RL algorithms typically rely on Constrained Markov Decision Processes and discourage constraint violations through reward shaping. However, such soft constraints cannot offer verifiable safety guarantees. To address this gap, we propose POLICEd RL, a novel RL algorithm explicitly designed to enforce affine hard constraints in closed-loop with a black-box environment. Our key insight is to force the learned policy to be affine around the unsafe set and use this affine region as a repulsive buffer to prevent trajectories from violating the constraint. We prove that such policies exist and guarantee constraint satisfaction. Our proposed framework is applicable to both systems with continuous and discrete state and action spaces and is agnostic to the choice of the RL training algorithm. Our results demonstrate the capacity of POLICEd RL to enforce hard constraints in robotic tasks while significantly outperforming existing methods.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. mPOLICE: Provable Enforcement of Multi-Region Affine Constraints in Deep Neural Networks

    cs.LG 2025-02 conditional novelty 6.0 of 10

    mPOLICE generalizes POLICE to enforce exact affine output constraints inside multiple disjoint convex regions of a ReLU network's input by giving each region its own activation pattern.

Pith tools