Pith. sign in

REVIEW 2 cited by

Reset-Free Lifelong Learning with Skill-Space Planning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2012.03548 v3 pith:6Q7NBQOR submitted 2020-12-07 cs.LG cs.AIcs.RO

classification cs.LGcs.AIcs.RO
keywords lifelongplanningenvironmentsnon-episodicskillsagentsevenframework
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The objective of lifelong reinforcement learning (RL) is to optimize agents which can continuously adapt and interact in changing environments. However, current RL approaches fail drastically when environments are non-stationary and interactions are non-episodic. We propose Lifelong Skill Planning (LiSP), an algorithmic framework for non-episodic lifelong RL based on planning in an abstract space of higher-order skills. We learn the skills in an unsupervised manner using intrinsic rewards and plan over the learned skills using a learned dynamics model. Moreover, our framework permits skill discovery even from offline data, thereby reducing the need for excessive real-world interactions. We demonstrate empirically that LiSP successfully enables long-horizon planning and learns agents that can avoid catastrophic failures even in challenging non-stationary and non-episodic environments derived from gridworld and MuJoCo benchmarks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Reinforcement Learning: From Algorithms To Foundation Models

    cs.AI 2026-07 conditional novelty 3.0 of 10

    A dissertation uniting the author's published results: non-exploitable Nash-DQN policies and the FightLadder benchmark for games, plus diffusion/consistency-model world models for RL — a compilation rather than new results.

  2. Advancements and Challenges in Continual Reinforcement Learning: A Comprehensive Review

    cs.LG 2025-06 conditional novelty 2.0 of 10

    A survey that categorizes continual reinforcement learning methods, environments, and evaluation metrics for deep RL, with a focus on robotics.

Pith tools