Pith. sign in

REVIEW 1 cited by

SAFER: Data-Efficient and Safe Reinforcement Learning via Skill Acquisition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2202.04849 v2 pith:7FI66XKC submitted 2022-02-10 cs.LG

classification cs.LG
keywords learningsafetaskssaferexperiencesmethodspolicyskills
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Methods that extract policy primitives from offline demonstrations using deep generative models have shown promise at accelerating reinforcement learning(RL) for new tasks. Intuitively, these methods should also help to trainsafeRLagents because they enforce useful skills. However, we identify these techniques are not well equipped for safe policy learning because they ignore negative experiences(e.g., unsafe or unsuccessful), focusing only on positive experiences, which harms their ability to generalize to new tasks safely. Rather, we model the latentsafetycontextusing principled contrastive training on an offline dataset of demonstrations from many tasks, including both negative and positive experiences. Using this late variable, our RL framework, SAFEty skill pRiors (SAFER) extracts task-specific safe primitive skills to safely and successfully generalize to new tasks. In the inference stage, policies trained with SAFER learn to compose safe skills into successful policies. We theoretically characterize why SAFER can enforce safe policy learning and demonstrate its effectiveness on several complex safety-critical robotic grasping tasks inspired by the game Operation, in which SAFERoutperforms state-of-the-art primitive learning methods in success and safety.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Enhancing Offline Reinforcement Learning with Curriculum Learning-Based Trajectory Valuation

    cs.LG 2025-02 reject novelty 4.0 of 10

    CLTV selects source trajectories that resemble a small target dataset (using learned transition scores and a KL-plus-return trajectory value) and trains the offline RL agent on those trajectories together with the tar...

Pith tools