Pith. sign in

REVIEW 2 cited by

Infinite-Horizon Reach-Avoid Zero-Sum Games via Deep Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2203.10142 v3 pith:LOCID7AL submitted 2022-03-18 eess.SY cs.AIcs.LGcs.SYmath.OC

classification eess.SYcs.AIcs.LGcs.SYmath.OC
keywords reach-avoidfunctionvaluegivenstatesapproximationconservativeconstraints
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper, we consider the infinite-horizon reach-avoid zero-sum game problem, where the goal is to find a set in the state space, referred to as the reach-avoid set, such that the system starting at a state therein could be controlled to reach a given target set without violating constraints under the worst-case disturbance. We address this problem by designing a new value function with a contracting Bellman backup, where the super-zero level set, i.e., the set of states where the value function is evaluated to be non-negative, recovers the reach-avoid set. Building upon this, we prove that the proposed method can be adapted to compute the viability kernel, or the set of states which could be controlled to satisfy given constraints, and the backward reachable set, or the set of states that could be driven towards a given target set. Finally, we propose to alleviate the curse of dimensionality issue in high-dimensional problems by extending Conservative Q-Learning, a deep reinforcement learning technique, to learn a value function such that the super-zero level set of the learned value function serves as a (conservative) approximation to the reach-avoid set. Our theoretical and empirical results suggest that the proposed method could learn reliably the reach-avoid set and the optimal control policy even with neural network approximation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Safe Multi-Agent Reinforcement Learning with Convergence to Generalized Nash Equilibrium

    cs.LG 2024-11 conditional novelty 7.0 of 10

    A new safe MARL framework with state-wise constraints that converges to a generalized Nash equilibrium by jointly learning controlled invariant sets and task policies.

  2. Bridging Adaptivity and Safety: Learning Agile Collision-Free Locomotion Across Varied Physics

    cs.RO 2025-01 conditional novelty 6.0 of 10

    A legged-robot controller that estimates payload and friction online and uses those estimates to switch between agile and recovery policies achieves lower collision rates and higher speeds than non-adaptive baselines.

Pith tools