Pith. sign in

REVIEW 2 cited by

A CMDP-within-online framework for Meta-Safe Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.16601 v1 pith:ITWZZ2KF submitted 2024-05-26 cs.LG

classification cs.LG
keywords learningconstraintframeworkoptimalityviolationsapproachbeenbounds
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Meta-reinforcement learning has widely been used as a learning-to-learn framework to solve unseen tasks with limited experience. However, the aspect of constraint violations has not been adequately addressed in the existing works, making their application restricted in real-world settings. In this paper, we study the problem of meta-safe reinforcement learning (Meta-SRL) through the CMDP-within-online framework to establish the first provable guarantees in this important setting. We obtain task-averaged regret bounds for the reward maximization (optimality gap) and constraint violations using gradient-based meta-learning and show that the task-averaged optimality gap and constraint satisfaction improve with task-similarity in a static environment or task-relatedness in a dynamic environment. Several technical challenges arise when making this framework practical. To this end, we propose a meta-algorithm that performs inexact online learning on the upper bounds of within-task optimality gap and constraint violations estimated by off-policy stationary distribution corrections. Furthermore, we enable the learning rates to be adapted for every task and extend our approach to settings with a competing dynamically changing oracle. Finally, experiments are conducted to demonstrate the effectiveness of our approach.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Adjustment Speed as a Safety Constraint for Nonstationary Reinforcement Learning

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Forecasting context shifts and treating adaptation demand against calibrated recovery capacity as a safety gate reduces transient violations in a nonstationary highway-driving simulator.

  2. Estimation of Regions of Attraction for Nonlinear Systems via Coordinate-Transformed TS Models and Piecewise Quadratic Lyapunov Functions

    math.DS 2025-07 reject novelty 3.0 of 10

    The paper claims that combining coordinate transformations with piecewise quadratic Lyapunov functions enlarges Takagi-Sugeno region-of-attraction estimates, but the demonstration is a single example with inconsistent...

Pith tools