Pith. sign in

REVIEW 3 cited by

Value Functions are Control Barrier Functions: Verification of Safe Policies using Control Theory

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.04026 v4 pith:P2CVPWCD submitted 2023-06-06 cs.LG cs.AIcs.RO

classification cs.LGcs.AIcs.RO
keywords controlfunctionsvaluelearningpoliciessafetheoryverification
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Guaranteeing safe behaviour of reinforcement learning (RL) policies poses significant challenges for safety-critical applications, despite RL's generality and scalability. To address this, we propose a new approach to apply verification methods from control theory to learned value functions. By analyzing task structures for safety preservation, we formalize original theorems that establish links between value functions and control barrier functions. Further, we propose novel metrics for verifying value functions in safe control tasks and practical implementation details to improve learning. Our work presents a novel method for certificate learning, which unlocks a diversity of verification techniques from control theory for RL policies, and marks a significant step towards a formal framework for the general, scalable, and verifiable design of RL-based control systems. Code and videos are available at this https url: https://rl-cbf.github.io/

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CBF-RL: Safety Filtering Reinforcement Learning in Training with Control Barrier Functions

    cs.RO 2025-10 conditional novelty 6.0 of 10

    Training RL policies with a closed-form CBF safety filter plus CBF reward lets a Unitree G1 humanoid avoid obstacles and climb stairs without a runtime safety filter.

  2. Towards Safe Robot Foundation Models Using Inductive Biases

    cs.RO 2025-05 conditional novelty 5.0 of 10

    A modular ATACOM safety layer is added to robot foundation models pi0 and OCTO, yielding provably safe actions with minimal performance loss.

  3. Learning Ensembles of Vision-based Safety Control Filters

    cs.LG 2024-12 conditional novelty 5.0 of 10

    Ensembles of vision-based safety filters with diverse backbones and aggregation methods improve safe/unsafe classification accuracy over individual models on the DeepAccident dataset.

Pith tools