REVIEW 4 cited by
Learning by Playing - Solving Sparse Reward Tasks from Scratch
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We propose Scheduled Auxiliary Control (SAC-X), a new learning paradigm in the context of Reinforcement Learning (RL). SAC-X enables learning of complex behaviors - from scratch - in the presence of multiple sparse reward signals. To this end, the agent is equipped with a set of general auxiliary tasks, that it attempts to learn simultaneously via off-policy RL. The key idea behind our method is that active (learned) scheduling and execution of auxiliary policies allows the agent to efficiently explore its environment - enabling it to excel at sparse reward RL. Our experiments in several challenging robotic manipulation settings demonstrate the power of our approach.
Forward citations
Cited by 4 Pith papers
-
Reinforcement Learning via Implicit Imitation Guidance
A reinforcement learning method that learns a state-dependent covariance from expert-policy action differences and uses it as exploration noise, improving sample efficiency on sparse-reward continuous control tasks.
-
Learning to Explore in Motion and Interaction Tasks
A learned generative model of past task motions, used as exploration noise in DDPG, speeds up learning of new robot manipulation and contact tasks by more than two times in simulation.
-
Learning to combine primitive skills: A step towards versatile robotic manipulation
RLBC, a hierarchical reinforcement learning method, combines behavioral-cloned primitive skills into composite manipulation tasks using only sparse rewards and no full-task demonstrations, with sim-to-real transfer de...
-
Solving Rubik's Cube Without Tricky Sampling
The paper presents a PPO policy trained with rewards from a learned cost model, claiming 99.4% success on the 2x2x2 Rubik's Cube without search or solved-state sampling, but with weak evidential support.
Discussion (0). Continue with ORCID to comment.