Pith. sign in

REVIEW 3 cited by

OmniSafe: An Infrastructure for Accelerating Safe Reinforcement Learning Research

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.09304 v1 pith:FQOLMSCU submitted 2023-05-16 cs.LG cs.AI

classification cs.LGcs.AI
keywords researchsaferllearningagentsframeworkreinforcementsafesafety
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

AI systems empowered by reinforcement learning (RL) algorithms harbor the immense potential to catalyze societal advancement, yet their deployment is often impeded by significant safety concerns. Particularly in safety-critical applications, researchers have raised concerns about unintended harms or unsafe behaviors of unaligned RL agents. The philosophy of safe reinforcement learning (SafeRL) is to align RL agents with harmless intentions and safe behavioral patterns. In SafeRL, agents learn to develop optimal policies by receiving feedback from the environment, while also fulfilling the requirement of minimizing the risk of unintended harm or unsafe behavior. However, due to the intricate nature of SafeRL algorithm implementation, combining methodologies across various domains presents a formidable challenge. This had led to an absence of a cohesive and efficacious learning framework within the contemporary SafeRL research milieu. In this work, we introduce a foundational framework designed to expedite SafeRL research endeavors. Our comprehensive framework encompasses an array of algorithms spanning different RL domains and places heavy emphasis on safety elements. Our efforts are to make the SafeRL-related research process more streamlined and efficient, therefore facilitating further research in AI safety. Our project is released at: https://github.com/PKU-Alignment/omnisafe.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation

    cs.LG 2024-12 conditional novelty 6.0 of 10

    The paper introduces Gradient-based Estimation (GBE), a Taylor-expansion method using analytic trajectory gradients, and the trust-region algorithm CGPO for safe RL with finite-horizon constraints.

  2. A Dynamic Safety Shield for Safe and Efficient Reinforcement Learning of Navigation Tasks

    cs.RO 2024-12 conditional novelty 6.0 of 10

    A dynamic safety shield uses an RL supervisor to adaptively weight obstacle-avoidance and action-matching terms in an MPC cost, improving the goals-to-collisions ratio in navigation RL.

  3. Offline Safe Reinforcement Learning Using Trajectory Classification

    cs.LG 2024-12 conditional novelty 5.0 of 10

    TraC trains an offline safe RL policy by classifying trajectories as desirable (safe, high-reward) versus undesirable (unsafe or low-reward) using a logistic loss on a policy-ratio score.

Pith tools