Pith. sign in

REVIEW 1 cited by

Solving Compositional Reinforcement Learning Problems via Task Reduction

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2103.07607 v2 pith:QF6UDVO5 submitted 2021-03-13 cs.LG

classification cs.LG
keywords taskreductionlearningagentcompositionalproblemsreinforcementself-imitation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We propose a novel learning paradigm, Self-Imitation via Reduction (SIR), for solving compositional reinforcement learning problems. SIR is based on two core ideas: task reduction and self-imitation. Task reduction tackles a hard-to-solve task by actively reducing it to an easier task whose solution is known by the RL agent. Once the original hard task is successfully solved by task reduction, the agent naturally obtains a self-generated solution trajectory to imitate. By continuously collecting and imitating such demonstrations, the agent is able to progressively expand the solved subspace in the entire task space. Experiment results show that SIR can significantly accelerate and improve learning on a variety of challenging sparse-reward continuous-control problems with compositional structures. Code and videos are available at https://sites.google.com/view/sir-compositional.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Umbrella Reinforcement Learning -- computationally efficient tool for hard non-linear problems

    cs.LG 2024-11 reject novelty 6.0 of 10

    Umbrella RL adds an ensemble-entropy bonus to policy gradient to solve sparse-reward, trap-heavy RL tasks, and reports large gains over PPO, RND, iLQR, and value iteration on two toy benchmarks.

Pith tools