Pith. sign in

REVIEW 2 cited by

Reverse Forward Curriculum Learning for Extreme Sample and Demonstration Efficiency in Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.03379 v1 pith:LMQY3EFK submitted 2024-05-06 cs.LG cs.AIcs.RO

classification cs.LGcs.AIcs.RO
keywords curriculumdemonstrationreverseforwardinitialdataefficiencylearning
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Reinforcement learning (RL) presents a promising framework to learn policies through environment interaction, but often requires an infeasible amount of interaction data to solve complex tasks from sparse rewards. One direction includes augmenting RL with offline data demonstrating desired tasks, but past work often require a lot of high-quality demonstration data that is difficult to obtain, especially for domains such as robotics. Our approach consists of a reverse curriculum followed by a forward curriculum. Unique to our approach compared to past work is the ability to efficiently leverage more than one demonstration via a per-demonstration reverse curriculum generated via state resets. The result of our reverse curriculum is an initial policy that performs well on a narrow initial state distribution and helps overcome difficult exploration problems. A forward curriculum is then used to accelerate the training of the initial policy to perform well on the full initial state distribution of the task and improve demonstration and sample efficiency. We show how the combination of a reverse curriculum and forward curriculum in our method, RFCL, enables significant improvements in demonstration and sample efficiency compared against various state-of-the-art learning-from-demonstration baselines, even solving previously unsolvable tasks that require high precision and control.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. RL-AWB: Deep Reinforcement Learning for Auto White Balance Correction in Low-Light Night-time Scenes

    cs.CV 2026-01 unverdicted novelty 7.0 of 10

    RL-AWB uses reinforcement learning to optimize parameters of a statistical white-balance estimator for nighttime scenes and reports better generalization on a new multi-sensor dataset.

  2. Umbrella Reinforcement Learning -- computationally efficient tool for hard non-linear problems

    cs.LG 2024-11 reject novelty 6.0 of 10

    Umbrella RL adds an ensemble-entropy bonus to policy gradient to solve sparse-reward, trap-heavy RL tasks, and reports large gains over PPO, RND, iLQR, and value iteration on two toy benchmarks.

Pith tools