Pith. sign in

REVIEW 8 cited by

Representation Learning via Invariant Causal Mechanisms

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2010.07922 v1 pith:IAMMPMFQ submitted 2020-10-15 cs.LG cs.CVstat.ML

classification cs.LGcs.CVstat.ML
keywords learningmethodsself-supervisedaugmentationscausaldatainvariantproxy
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Self-supervised learning has emerged as a strategy to reduce the reliance on costly supervised signal by pretraining representations only using unlabeled data. These methods combine heuristic proxy classification tasks with data augmentations and have achieved significant success, but our theoretical understanding of this success remains limited. In this paper we analyze self-supervised representation learning using a causal framework. We show how data augmentations can be more effectively utilized through explicit invariance constraints on the proxy classifiers employed during pretraining. Based on this, we propose a novel self-supervised objective, Representation Learning via Invariant Causal Mechanisms (ReLIC), that enforces invariant prediction of proxy targets across augmentations through an invariance regularizer which yields improved generalization guarantees. Further, using causality we generalize contrastive learning, a particular kind of self-supervised method, and provide an alternative theoretical explanation for the success of these methods. Empirically, ReLIC significantly outperforms competing methods in terms of robustness and out-of-distribution generalization on ImageNet, while also significantly outperforming these methods on Atari achieving above human-level performance on $51$ out of $57$ games.

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Entity-Centric World Models: Interaction-Aware Masking for Causal Video Prediction

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    IA-JEPA applies motion-centric masking in JEPA to focus on entity interactions, reporting 14.26% causal reasoning accuracy on CLEVRER versus 3.22% for standard baselines plus higher latent entropy and R²=0.43 energy l...

  2. Entity-Centric World Models: Interaction-Aware Masking for Causal Video Prediction

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    IA-JEPA applies interaction-aware masking to JEPA, raising causal reasoning accuracy on CLEVRER from 3.22% to 14.26% while producing a higher-entropy latent space that better aligns with physical energy.

  3. Trustworthy AI Suffers from Invariance Conflicts and Causality is The Solution

    cs.AI 2026-05 unverdicted novelty 6.0 of 10

    Causality provides a unifying framework for resolving trade-offs in trustworthy AI by managing invariance conflicts under changes to the data-generating process.

  4. A Recipe for Causal Graph Regression: Confounding Effects Revisited

    cs.LG 2025-07 conditional novelty 6.0 of 10

    The paper proposes a contrastive-learning-based causal graph regression framework that explicitly models the predictive power of confounding subgraphs and achieves state-of-the-art OOD generalization on graph regressi...

  5. Revisiting Feature Prediction for Learning Visual Representations from Video

    cs.CV 2024-02 conditional novelty 6.0 of 10

    V-JEPA models trained only on feature prediction from 2 million public videos achieve 81.9% on Kinetics-400, 72.2% on Something-Something-v2, and 77.9% on ImageNet-1K using frozen ViT-H/16 backbones.

  6. Geometric Deep Learning: Grids, Groups, Graphs, Geodesics, and Gauges

    cs.LG 2021-04 accept novelty 6.0 of 10

    Geometric deep learning provides a unified mathematical framework based on grids, groups, graphs, geodesics, and gauges to explain and extend neural network architectures by incorporating physical regularities.

  7. Dirichlet-Guided Group Forecasting for Alleviating Over-smoothing in Time Series Forecasting

    cs.LG 2026-06 unverdicted novelty 5.0 of 10

    DGF explicitly models multiple mode-conditioned predictive distributions via Dirichlet-guided sampling and reward optimization to preserve dynamical features in time series forecasts.

  8. Trustworthy AI Suffers from Invariance Conflicts and Causality is The Solution

    cs.AI 2026-05 unverdicted novelty 4.0 of 10

    Causality resolves trade-offs in trustworthy AI by treating them as invariance conflicts under different data-generating process changes.

Pith tools