Pith. sign in

REVIEW 10 cited by

Identifiability Results for Multimodal Contrastive Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.09166 v1 pith:HNKNNFZQ submitted 2023-03-16 cs.LG stat.ML

Identifiability Results for Multimodal Contrastive Learning

classification cs.LG stat.ML
keywords learningmultimodalcontrastivefactorsidentifiabilityresultsworklatent
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Contrastive learning is a cornerstone underlying recent progress in multi-view and multimodal learning, e.g., in representation learning with image/caption pairs. While its effectiveness is not yet fully understood, a line of recent work reveals that contrastive learning can invert the data generating process and recover ground truth latent factors shared between views. In this work, we present new identifiability results for multimodal contrastive learning, showing that it is possible to recover shared factors in a more general setup than the multi-view setting studied previously. Specifically, we distinguish between the multi-view setting with one generative mechanism (e.g., multiple cameras of the same type) and the multimodal setting that is characterized by distinct mechanisms (e.g., cameras and microphones). Our work generalizes previous identifiability results by redefining the generative process in terms of distinct mechanisms with modality-specific latent variables. We prove that contrastive learning can block-identify latent factors shared between modalities, even when there are nontrivial dependencies between factors. We empirically verify our identifiability results with numerical simulations and corroborate our findings on a complex multimodal dataset of image/text pairs. Zooming out, our work provides a theoretical basis for multimodal representation learning and explains in which settings multimodal contrastive learning can be effective in practice.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. When Do Diffusion Models learn to Generate Multiple Objects?

    cs.CV 2026-04 unverdicted novelty 8.0

    Diffusion models' multi-object generation is limited primarily by scene complexity and held-out combinations rather than imbalance, with counting difficult in low data and compositional generalization collapsing as mo...

  2. Mechanistic Independence: A Principle for Identifiable Disentangled Representations

    cs.LG 2025-09 unverdicted novelty 7.0

    Mechanistic independence criteria yield identifiability of latent subspaces under nonlinear mixing by focusing on action-based independence rather than latent distributions, with a hierarchy and graph-theoretic view o...

  3. Unsupervised Causal Abstractions Discovery

    cs.LG 2026-06 unverdicted novelty 6.0

    Low-rank graphs induce latents that form causal abstractions, with identifiability results and a practical objective enabling unsupervised learning of high-level SCMs from low-level measurements.

  4. Understanding Self-Supervised Learning via Latent Distribution Matching

    cs.LG 2026-05 unverdicted novelty 6.0

    Self-supervised learning is recast as latent distribution matching that unifies multiple SSL families and yields a sampling-free Kalman-based predictor plus an identifiability proof for predictive variants under mild ...

  5. Understanding Self-Supervised Learning via Latent Distribution Matching

    cs.LG 2026-05 unverdicted novelty 6.0

    Self-supervised learning is cast as latent distribution matching that aligns representations to a model while enforcing uniformity, unifying multiple SSL families and proving identifiability for predictive variants ev...

  6. Understanding Self-Supervised Learning via Latent Distribution Matching

    cs.LG 2026-05 conditional novelty 6.0

    Self-supervised learning can be understood as latent distribution matching, and under a Gaussian predictive model this yields identifiable representations up to affine transformations.

  7. When Do Diffusion Models learn to Generate Multiple Objects?

    cs.CV 2026-04 unverdicted novelty 6.0

    Using the mosaic controlled dataset framework, experiments show scene complexity dominates over concept imbalance in diffusion model failures for multi-object generation, with counting especially hard in low-data regi...

  8. MoVA: Learning Asymmetric Dual Projections for Modular Long Video-Text Alignment

    cs.CV 2026-07 unverdicted novelty 5.0

    MoVA introduces modular asymmetric dual projections to handle temporal misalignment and semantic asymmetry in long video-text alignment.

  9. Understanding Self-Supervised Learning via Latent Distribution Matching

    cs.LG 2026-05 unverdicted novelty 5.0

    Self-supervised learning is recast as latent distribution matching that unifies ICA with contrastive and predictive methods and derives an identifiable nonlinear Bayesian filtering model for timeseries.

  10. Foundation Models for Astrophysics

    astro-ph.IM 2026-08 conditional novelty 3.0

    Astronomical 'foundation models' largely reuse transformers and self-supervised pretraining, but evidence of transfer to new instruments, populations, or tasks remains rare; the paper argues such evidence, not archite...