Pith. sign in

REVIEW 1 cited by

Continual Model-Based Reinforcement Learning with Hypernetworks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2009.11997 v3 pith:RSGSXFHL submitted 2020-09-25 cs.LG cs.AIcs.RO

classification cs.LGcs.AIcs.RO
keywords learningdynamicsexperiencecontinualhypercrlhypernetworksmodelmodel-based
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Effective planning in model-based reinforcement learning (MBRL) and model-predictive control (MPC) relies on the accuracy of the learned dynamics model. In many instances of MBRL and MPC, this model is assumed to be stationary and is periodically re-trained from scratch on state transition experience collected from the beginning of environment interactions. This implies that the time required to train the dynamics model - and the pause required between plan executions - grows linearly with the size of the collected experience. We argue that this is too slow for lifelong robot learning and propose HyperCRL, a method that continually learns the encountered dynamics in a sequence of tasks using task-conditional hypernetworks. Our method has three main attributes: first, it includes dynamics learning sessions that do not revisit training data from previous tasks, so it only needs to store the most recent fixed-size portion of the state transition experience; second, it uses fixed-capacity hypernetworks to represent non-stationary and task-aware dynamics; third, it outperforms existing continual learning alternatives that rely on fixed-capacity networks, and does competitively with baselines that remember an ever increasing coreset of past experience. We show that HyperCRL is effective in continual model-based reinforcement learning in robot locomotion and manipulation scenarios, such as tasks involving pushing and door opening. Our project website with videos is at this link https://rvl.cs.toronto.edu/blog/hypercrl

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. When Does Neuroevolution Outcompete Reinforcement Learning in Transfer Learning Tasks?

    cs.LG 2025-05 conditional novelty 6.0 of 10

    On two new curriculum benchmarks, direct-encoding neuroevolution (NEAT) transfers skills across levels better than PPO reinforcement learning, while indirect encodings like HyperNEAT transfer poorly.

Pith tools