Pith. sign in

REVIEW 5 cited by

C-Learning: Learning to Achieve Goals via Recursive Classification

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2011.08909 v2 pith:QX5SFY6L submitted 2020-11-17 cs.LG cs.AI

classification cs.LGcs.AI
keywords futuredensitygoal-conditionedstatedistributionfunctionlearningallows
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We study the problem of predicting and controlling the future state distribution of an autonomous agent. This problem, which can be viewed as a reframing of goal-conditioned reinforcement learning (RL), is centered around learning a conditional probability density function over future states. Instead of directly estimating this density function, we indirectly estimate this density function by training a classifier to predict whether an observation comes from the future. Via Bayes' rule, predictions from our classifier can be transformed into predictions over future states. Importantly, an off-policy variant of our algorithm allows us to predict the future state distribution of a new policy, without collecting new experience. This variant allows us to optimize functionals of a policy's future state distribution, such as the density of reaching a particular goal state. While conceptually similar to Q-learning, our work lays a principled foundation for goal-conditioned RL as density estimation, providing justification for goal-conditioned methods used in prior work. This foundation makes hypotheses about Q-learning, including the optimal goal-sampling ratio, which we confirm experimentally. Moreover, our proposed method is competitive with prior goal-conditioned RL methods.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Equivariant Goal Conditioned Contrastive Reinforcement Learning

    cs.RO 2025-07 conditional novelty 6.0 of 10

    Equivariant Contrastive RL imposes C8 rotation symmetry on the critic and actor, improving sample efficiency and goal generalization in simulated manipulation.

  2. Efficient Skill Discovery via Regret-Aware Optimization

    cs.LG 2025-06 conditional novelty 6.0 of 10

    A regret-aware skill discovery algorithm, RSD, improves sample efficiency and zero-shot goal-reaching in high-dimensional continuous control by focusing exploration on unmastered skills.

  3. Reachability Weighted Offline Goal-conditioned Resampling

    cs.LG 2025-06 conditional novelty 6.0 of 10

    RWS trains a positive-unlabeled reachability classifier on goal-conditioned Q-values and uses it to re-weight goal sampling, improving offline goal-conditioned RL performance on robotic manipulation benchmarks.

  4. GCHR : Goal-Conditioned Hindsight Regularization for Sample-Efficient Reinforcement Learning

    cs.LG 2025-08 conditional novelty 5.0 of 10

    GCHR combines hindsight self-imitation with a hindsight-goal KL prior in off-policy RL, reporting large sample-efficiency gains on Fetch and Shadow Hand tasks.

  5. Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization

    cs.LG 2025-06 conditional novelty 5.0 of 10

    GCReinSL adds Q-conditioned maximization to supervised offline RL, using normalizing flows to estimate goal-reaching probabilities and expectile regression to condition actions on the best in-distribution value, impro...

Pith tools