Pith. sign in

REVIEW 1 cited by

Safe Reinforcement Learning for Constrained Markov Decision Processes with Stochastic Stopping Time

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.15928 v1 pith:NAYKQJD7 submitted 2024-03-23 cs.LG math.OC

classification cs.LGmath.OC
keywords learningalgorithmpolicysafesafetyconstrainedconstraintsdecision
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper, we present an online reinforcement learning algorithm for constrained Markov decision processes with a safety constraint. Despite the necessary attention of the scientific community, considering stochastic stopping time, the problem of learning optimal policy without violating safety constraints during the learning phase is yet to be addressed. To this end, we propose an algorithm based on linear programming that does not require a process model. We show that the learned policy is safe with high confidence. We also propose a method to compute a safe baseline policy, which is central in developing algorithms that do not violate the safety constraints. Finally, we provide simulation results to show the efficacy of the proposed algorithm. Further, we demonstrate that efficient exploration can be achieved by defining a subset of the state-space called proxy set.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Distributionally Robust Safety Verification for Markov Decision Processes

    eess.SY 2024-11 conditional novelty 4.0 of 10

    An upper bound on the worst-case probability of hitting an unsafe state under Wasserstein-ambiguous transition kernels is computed via a convex-program robust Q-iteration for MDPs.

Pith tools