Pith. sign in

REVIEW 2 cited by

Fast and Slow Learning of Recurrent Independent Mechanisms

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2105.08710 v2 pith:CHTHUQZN submitted 2021-05-18 cs.LG cs.AI

classification cs.LGcs.AI
keywords knowledgemodulespiecesattentionlearningparametersadaptationagent
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Decomposing knowledge into interchangeable pieces promises a generalization advantage when there are changes in distribution. A learning agent interacting with its environment is likely to be faced with situations requiring novel combinations of existing pieces of knowledge. We hypothesize that such a decomposition of knowledge is particularly relevant for being able to generalize in a systematic manner to out-of-distribution changes. To study these ideas, we propose a particular training framework in which we assume that the pieces of knowledge an agent needs and its reward function are stationary and can be re-used across tasks. An attention mechanism dynamically selects which modules can be adapted to the current task, and the parameters of the selected modules are allowed to change quickly as the learner is confronted with variations in what it experiences, while the parameters of the attention mechanisms act as stable, slowly changing, meta-parameters. We focus on pieces of knowledge captured by an ensemble of modules sparsely communicating with each other via a bottleneck of attention. We find that meta-learning the modular aspects of the proposed system greatly helps in achieving faster adaptation in a reinforcement learning setup involving navigation in a partially observed grid world with image-level input. We also find that reversing the role of parameters and meta-parameters does not work nearly as well, suggesting a particular role for fast adaptation of the dynamically selected modules.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Dreamweaver: Learning Compositional World Models from Pixels

    cs.CV 2025-01 conditional novelty 6.0 of 10

    An unsupervised recurrent block-slot model that discovers static and dynamic concept blocks from raw video and recombines them to imagine novel future videos.

  2. CAREL: Instruction-guided reinforcement learning with cross-modal auxiliary objectives

    cs.LG 2024-11 conditional novelty 6.0 of 10

    CAREL improves instruction-following RL sample efficiency by aligning observation sequences with instruction tokens via an X-CLIP style contrastive loss and masking completed subtasks.

Pith tools