Pith. sign in

REVIEW 3 cited by

Contrastive Learning as Goal-Conditioned Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2206.07568 v2 pith:WSLXEFE5 submitted 2022-06-15 cs.LG cs.AI

classification cs.LGcs.AI
keywords learningcontrastivepriorrepresentationmethodsrepresentationsalgorithmsgoal-conditioned
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In reinforcement learning (RL), it is easier to solve a task if given a good representation. While deep RL should automatically acquire such good representations, prior work often finds that learning representations in an end-to-end fashion is unstable and instead equip RL algorithms with additional representation learning parts (e.g., auxiliary losses, data augmentation). How can we design RL algorithms that directly acquire good representations? In this paper, instead of adding representation learning parts to an existing RL algorithm, we show (contrastive) representation learning methods can be cast as RL algorithms in their own right. To do this, we build upon prior work and apply contrastive representation learning to action-labeled trajectories, in such a way that the (inner product of) learned representations exactly corresponds to a goal-conditioned value function. We use this idea to reinterpret a prior RL method as performing contrastive learning, and then use the idea to propose a much simpler method that achieves similar performance. Across a range of goal-conditioned RL tasks, we demonstrate that contrastive RL methods achieve higher success rates than prior non-contrastive methods, including in the offline RL setting. We also show that contrastive RL outperforms prior methods on image-based tasks, without using data augmentation or auxiliary objectives.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Good Rankers, Bad Objectives: Bilinear Contrastive Critics under Expressive Policy Search

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Bilinear contrastive critics remain good compatibility rankers but are unsafe to maximize for action selection; cosine bounding does not fix value decalibration, while Bellman TD-Q does.

  2. Visual Pre-Training on Unlabeled Images using Reinforcement Learning

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Casting image-crop consistency as temporal-difference value learning improves visual representations on unlabeled web, scene, and video data.

  3. Offline RL with Hierarchical Action Chunking

    cs.LG 2026-07 conditional novelty 5.0 of 10

    HiQC, a hierarchy of latent subgoal planning and chunked action execution, achieves the best OGBench aggregate score (53%) and an O(sqrt(T/k)) error bound under a bootstrap-chain model.

Pith tools