Pith. sign in

REVIEW 18 cited by

Learning Invariant Representations for Reinforcement Learning without Reconstruction

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2006.10742 v2 pith:GNISDVH6 submitted 2020-06-18 cs.LG cs.AIstat.ML

classification cs.LGcs.AIstat.ML
keywords learningbisimulationmethodrepresentationsdistancesinformationinvariancelatent
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We study how representation learning can accelerate reinforcement learning from rich observations, such as images, without relying either on domain knowledge or pixel-reconstruction. Our goal is to learn representations that both provide for effective downstream control and invariance to task-irrelevant details. Bisimulation metrics quantify behavioral similarity between states in continuous MDPs, which we propose using to learn robust latent representations which encode only the task-relevant information from observations. Our method trains encoders such that distances in latent space equal bisimulation distances in state space. We demonstrate the effectiveness of our method at disregarding task-irrelevant information using modified visual MuJoCo tasks, where the background is replaced with moving distractors and natural videos, while achieving SOTA performance. We also test a first-person highway driving task where our method learns invariance to clouds, weather, and time of day. Finally, we provide generalization results drawn from properties of bisimulation metrics, and links to causal inference.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 18 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 77 citations worldwide. Full citation record

  1. Same physical state, different collective dynamics: state encodings select synchronization outcomes in language-model agents

    physics.soc-ph 2026-08 accept novelty 7.0 of 10

    State encodings, not just the physical state, determine the collective synchronization outcomes of language-model agent populations, and the effect is model-family dependent.

  2. When Does Reward Teach State? A Hidden-Automaton Instrument and a Group-Language Warning Signal

    cs.LG 2026-07 conditional novelty 7.0 of 10

    High reward in sparse RL does not imply latent-state recovery; a hidden-DFA instrument separates perception from planning gaps and flags group-language structure as a pre-training warning.

  3. Geometry of Neural Reinforcement Learning in Continuous State and Action Spaces

    cs.LG 2025-07 conditional novelty 7.0 of 10

    For wide two-layer linearized neural policies in deterministic continuous RL, the locally attainable states concentrate on a manifold of dimension at most 2da+1, independent of the state dimension.

  4. Self-Predictive Dynamics for Generalization of Vision-based Reinforcement Learning

    cs.CV 2025-06 conditional novelty 7.0 of 10

    A self-supervised auxiliary loss combining weak and strong augmentations, an adversarial discriminator, and inverse-then-forward latent dynamics improves both data efficiency and zero-shot generalization in vision-based RL.

  5. Understanding Behavioral Metric Learning: A Large-Scale Study on Distracting Reinforcement Learning Environments

    cs.LG 2025-05 conditional novelty 7.0 of 10

    Across noisy DeepMind Control tasks, explicit bisimulation-metric losses add little denoising benefit beyond plain self-prediction and feature normalization, which dominate performance.

  6. TaskSense: Focusing on What Matters in World Models

    cs.AI 2026-08 conditional novelty 6.0 of 10

    TaskSense filters visual observations with latent-conditioned stochastic spatial attention before encoding, improving world-model control under visual distractions relative to DreamerV3.

  7. Next-Latent Prediction Transformers Learn Compact World Models

    cs.LG 2025-11 unverdicted novelty 6.0 of 10

    NextLat augments next-token prediction with latent next-state prediction, theoretically converging latents to belief states and showing empirical gains in world modeling, reasoning, planning, and faster inference via ...

  8. Reimagination with Test-time Observation Interventions: Distractor-Robust World Model Predictions for Visual Model Predictive Control

    cs.RO 2025-06 conditional novelty 6.0 of 10

    Removing unfamiliar visual distractors from the input observation before world-model rollout, then compositing them back into predicted frames, makes action verification robust and raises real-robot pick-and-place suc...

  9. Towards Empowerment Gain through Causal Structure Learning in Model-Based RL

    cs.AI 2025-02 conditional novelty 6.0 of 10

    A model-based RL framework that alternates causal structure learning with empowerment-driven exploration, plus a curiosity reward, improves sample efficiency and asymptotic performance in six environments.

  10. Latent Action Learning Requires Supervision in the Presence of Distractors

    cs.CV 2025-02 conditional novelty 6.0 of 10

    Latent action models need at least a small amount of action supervision to learn useful actions when observations contain distractors, as shown on the Distracting Control Suite.

  11. Hierarchical Successor Representation for Robust Transfer

    cs.LG 2026-02 conditional novelty 5.0 of 10

    Successor representations built from temporally extended options are less sensitive to policy changes and, after non-negative matrix factorization, yield sparse, topologically interpretable features that speed transfe...

  12. Physics-informed Temporal Difference Metric Learning for Robot Motion Planning

    cs.RO 2025-05 conditional novelty 5.0 of 10

    A self-supervised neural planner combines Eikonal, temporal difference, obstacle alignment, and causality losses with a learned L1 and L-infinity metric, improving success rates and generalization in robot motion planning.

  13. LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing

    cs.LG 2025-02 conditional novelty 5.0 of 10

    A preference-conditioned PPO routing policy with IRT-based model identity vectors selects cost-effective LLMs per query and generalizes to unseen models from a handful of evaluation prompts.

  14. Contrastive Representation for Interactive Recommendation

    cs.IR 2024-12 conditional novelty 5.0 of 10

    CRIR adds a preference-ranking contrastive loss to a DDPG recommender, improving sample efficiency in simulated interactive recommendation.

  15. Mask-based Predictive Representations for Reinforcement Learning

    cs.LG 2026-07 conditional novelty 4.0 of 10

    Mask-based predictive representations (MPR) as an auxiliary self-supervised task improve sample efficiency of vision-based RL over prior SOTA on continuous and discrete control benchmarks.

  16. Efficient and Generalizable Environmental Understanding for Visual Navigation

    cs.AI 2025-06 conditional novelty 4.0 of 10

    Adding an auxiliary next-state prediction loss to EmbCLIP substantially improves object and point navigation in RoboTHOR and Habitat and boosts supervised vision-and-language navigation baselines.

  17. Decoupled Hierarchical Reinforcement Learning with State Abstraction for Discrete Grids

    cs.LG 2025-06 conditional novelty 4.0 of 10

    A decoupled hierarchical RL framework with a rule-based low-level policy and DeepMDP state abstraction outperforms PPO on two custom discrete grid environments, but with a single baseline and sparse experimental detail.

  18. Approximated Behavioral Metric-based State Projection for Federated Reinforcement Learning

    cs.LG 2025-05 conditional novelty 4.0 of 10

    Federated averaging of behavior-metric-based state projection networks improves cross-environment generalization in federated reinforcement learning, while the claimed privacy protection is not demonstrated.

Pith tools