Pith. sign in

REVIEW 4 cited by

InfoNCE: Identifying the Gap Between Theory and Practice

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.00143 v2 pith:5GNANMHD submitted 2024-06-28 cs.LG cs.CVstat.ML

classification cs.LGcs.CVstat.ML
keywords factorsinfoncelatentpracticeacrossaninfonceassumptionscertain
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Prior theory work on Contrastive Learning via the InfoNCE loss showed that, under certain assumptions, the learned representations recover the ground-truth latent factors. We argue that these theories overlook crucial aspects of how CL is deployed in practice. Specifically, they either assume equal variance across all latents or that certain latents are kept invariant. However, in practice, positive pairs are often generated using augmentations such as strong cropping to just a few pixels. Hence, a more realistic assumption is that all latent factors change with a continuum of variability across all factors. We introduce AnInfoNCE, a generalization of InfoNCE that can provably uncover the latent factors in this anisotropic setting, broadly generalizing previous identifiability results in CL. We validate our identifiability results in controlled experiments and show that AnInfoNCE increases the recovery of previously collapsed information in CIFAR10 and ImageNet, albeit at the cost of downstream accuracy. Finally, we discuss the remaining mismatches between theoretical assumptions and practical implementations.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Estimating the Empowerment of Language Model Agents

    cs.AI 2025-09 conditional novelty 6.0 of 10

    EELMA estimates the mutual information between an LM agent's actions and future text states, and this 'empowerment' is shown to correlate with task performance across toy games and WebArena.

  2. Contrastive Self-Supervised Network Intrusion Detection using Augmented Negative Pairs

    cs.LG 2025-09 conditional novelty 5.0 of 10

    CLAN clusters genuine benign network flows while repelling augmented copies, then classifies new flows by distance to the cluster centroid; on Lycos2017 it reports the highest mean AUROC among compared SSL and anomaly...

  3. Skill Learning via Policy Diversity Yields Identifiable Representations for Reinforcement Learning

    cs.LG 2025-07 conditional novelty 4.0 of 10

    Contrastive Successor Features recover ground-truth RL states up to a linear map whenever the skill-conditioned transition differences follow a von Mises-Fisher distribution and policies are diverse.

  4. Zero Shot Composed Image Retrieval

    cs.CV 2025-06 reject novelty 2.0 of 10

    Fine-tuning BLIP-2 with a Q-Former raises FashionIQ validation Recall@10 to roughly 45%, but the zero-shot framing is inaccurate and the DPO variant omits the reference image.

Pith tools