Pith. sign in

REVIEW 1 cited by

Learning with AMIGo: Adversarially Motivated Intrinsic Goals

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2006.12122 v2 pith:22B4N2D2 submitted 2020-06-22 cs.LG cs.AIstat.ML

classification cs.LGcs.AIstat.ML
keywords goalsintrinsiclearningadversariallyagentamigochallengingenvironment
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

A key challenge for reinforcement learning (RL) consists of learning in environments with sparse extrinsic rewards. In contrast to current RL methods, humans are able to learn new skills with little or no reward by using various forms of intrinsic motivation. We propose AMIGo, a novel agent incorporating -- as form of meta-learning -- a goal-generating teacher that proposes Adversarially Motivated Intrinsic Goals to train a goal-conditioned "student" policy in the absence of (or alongside) environment reward. Specifically, through a simple but effective "constructively adversarial" objective, the teacher learns to propose increasingly challenging -- yet achievable -- goals that allow the student to learn general skills for acting in a new environment, independent of the task to be solved. We show that our method generates a natural curriculum of self-proposed goals which ultimately allows the agent to solve challenging procedurally-generated tasks where other forms of intrinsic motivation and state-of-the-art RL methods fail.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Generalized Back-Stepping Experience Replay in Sparse-Reward Environments

    cs.LG 2024-12 conditional novelty 4.0 of 10

    GBER combines back-stepping experience replay with HER-style relabeling and rfaab sampling, improving learning speed and stability in sparse-reward mazes.

Pith tools