Pith. sign in

REVIEW 1 cited by

Safe Exploration in Reinforcement Learning: A Generalized Formulation and Algorithms

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.03225 v1 pith:3O7WDVFP submitted 2023-10-05 cs.LG cs.AIcs.RO

classification cs.LGcs.AIcs.RO
keywords safetyexplorationsafealgorithmgeneralizedmasealgorithmscombines
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Safe exploration is essential for the practical use of reinforcement learning (RL) in many real-world scenarios. In this paper, we present a generalized safe exploration (GSE) problem as a unified formulation of common safe exploration problems. We then propose a solution of the GSE problem in the form of a meta-algorithm for safe exploration, MASE, which combines an unconstrained RL algorithm with an uncertainty quantifier to guarantee safety in the current episode while properly penalizing unsafe explorations before actual safety violation to discourage them in future episodes. The advantage of MASE is that we can optimize a policy while guaranteeing with a high probability that no safety constraint will be violated under proper assumptions. Specifically, we present two variants of MASE with different constructions of the uncertainty quantifier: one based on generalized linear models with theoretical guarantees of safety and near-optimality, and another that combines a Gaussian process to ensure safety with a deep RL algorithm to maximize the reward. Finally, we demonstrate that our proposed algorithm achieves better performance than state-of-the-art algorithms on grid-world and Safety Gym benchmarks without violating any safety constraints, even during training.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Optimistic Exploration for Risk-Averse Constrained Reinforcement Learning

    cs.LG 2025-07 conditional novelty 6.0 of 10

    ORAC combines upper-confidence-bound reward exploration with lower-confidence-bound risk-averse cost constraints and adaptive cost weighting to improve exploration in risk-averse constrained RL.

Pith tools