Pith. sign in

REVIEW 1 cited by

Generalized Hindsight for Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2002.11708 v1 pith:ZNCI6NSI submitted 2020-02-26 cs.LG cs.AIcs.NEcs.ROstat.ML

classification cs.LGcs.AIcs.NEcs.ROstat.ML
keywords taskgeneralizedhindsightbehaviordatalearningreinforcementtasks
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

One of the key reasons for the high sample complexity in reinforcement learning (RL) is the inability to transfer knowledge from one task to another. In standard multi-task RL settings, low-reward data collected while trying to solve one task provides little to no signal for solving that particular task and is hence effectively wasted. However, we argue that this data, which is uninformative for one task, is likely a rich source of information for other tasks. To leverage this insight and efficiently reuse data, we present Generalized Hindsight: an approximate inverse reinforcement learning technique for relabeling behaviors with the right tasks. Intuitively, given a behavior generated under one task, Generalized Hindsight returns a different task that the behavior is better suited for. Then, the behavior is relabeled with this new task before being used by an off-policy RL optimizer. Compared to standard relabeling techniques, Generalized Hindsight provides a substantially more efficient reuse of samples, which we empirically demonstrate on a suite of multi-task navigation and manipulation tasks. Videos and code can be accessed here: https://sites.google.com/view/generalized-hindsight.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Hindsight Planner: A Closed-Loop Few-Shot Planner for Embodied Instruction Following

    cs.AI 2024-12 conditional novelty 6.0 of 10

    A few-shot LLM planner for ALFRED that relabels suboptimal trajectories with hindsight prompts reaches 25.51 SR on Test Seen, approaching or beating the full-shot HLSM baseline.

Pith tools