REVIEW 3 major objections 6 minor 25 references
The modality gap is a low-temperature failure of InfoNCE with independent encoders; mixing same-modality negatives (xNCE) removes it while improving zero-shot transfer.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-14 09:54 UTC pith:NHT2TL3C
load-bearing objection Clean causal experiment + simple fix that actually improves zero-shot; the low-temp mode-failure story is plausible but still rests on a stylized MNIST proxy and a qualitative limit argument. the 3 major comments →
On the modality gap and the contrastive loss in multi-modal representation learning
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The modality gap is a mode-failure of the multi-modal InfoNCE objective that appears only at low temperature: with independent encoders the denominator can be satisfied by pushing the two modalities apart rather than by learning aligned features. Replacing a controllable fraction of the negatives with same-modality samples (xNCE) removes this divergence incentive, collapses the gap, and still yields the low-temperature geometry required for strong zero-shot classification.
What carries the argument
xNCE: the InfoNCE softmax is computed on a mixed batch in which a fraction r_mix of the embeddings of one modality are swapped with those of the other, so that both inter- and intra-modality pairs appear as negatives. The same temperature regime can then be used without inducing a gap.
Load-bearing premise
That the two-encoder MNIST experiment (identical nets, identical start weights, same-family augmentations) and the low-temperature limit analysis correctly capture the optimization dynamics that create the gap inside a real high-dimensional pretrained CLIP model.
What would settle it
Train a CLIP-scale model from scratch with xNCE at the usual low temperature and measure whether the modality-centroid distance stays near zero while retrieval and zero-shot accuracy remain at least as high as the InfoNCE baseline; if the gap reappears or accuracy collapses, the claimed mechanism fails.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper argues that the modality gap in CLIP-style dual-encoder contrastive learning is a low-temperature mode-failure of InfoNCE with independent encoders, not merely an artifact of initialization or data. A controlled uni-modal MNIST experiment with two identically initialized, non-shared encoders shows that InfoNCE actively separates the two embedding clouds at low τ. A low-temperature limit analysis of the InfoNCE loss (Eqs. 3–4) is used to explain margin collapse into separated cones. The authors propose xNCE, which mixes a fraction r_mix of intra-modality negatives into the contrastive denominator, and show on COCO-finetuned CLIP ViT-B/16 that small mixing reduces the gap, nearly matches InfoNCE retrieval, and improves zero-shot accuracy on Caltech-101, CIFAR-10/100, and ImageNet relative to standard InfoNCE, high-temperature InfoNCE, and a regularized InfoNCE baseline.
Significance. If the mechanism is correct, the paper supplies a simple, loss-level fix for a widely observed pathology of dual-encoder contrastive models, with a favorable trade-off between gap reduction and transfer geometry that temperature scaling and explicit alignment/uniformity regularizers do not achieve. Strengths include a clean controlled MNIST construction that isolates the dual-encoder InfoNCE dynamics under matched initialization, an explicit low-τ limit argument, multi-run COCO tables with standard deviations, and head-to-head comparison against the natural alternatives (high-τ InfoNCE and Fahim-style regularized InfoNCE). The zero-shot gains at small r_mix are practically relevant and falsifiable. The contribution is incremental rather than foundational, but it is useful and actionable for multimodal representation learning.
major comments (3)
- [§5.2, Eqs. (3)–(4)] §5.2, Eqs. (3)–(4): The limit argument correctly shows that imperfect ranking plus τ→0+ incentivizes shrinking all similarities toward their mean m_i, which can produce a tight negative cluster. The dual-encoder-specific step—that this optimum is realized as two separated modality cones rather than within-modality collapse or other geometries—is only sketched in §3.3 and is not derived from the same limit. The manuscript should either (i) make the dual-encoder geometry a formal consequence of independent parameters under the low-τ objective, or (ii) clearly label the modality-gap interpretation as a plausible reading of the limit rather than a proven mode-failure unique to independent encoders.
- [§4.1–4.2, Fig. 6, Hypothesis H1] §4.1 vs §4.2 / Fig. 6: The claim that InfoNCE “actively generates” the gap is demonstrated only in the MNIST two-encoder proxy (identical architecture, identical init, same-family augmentations). All COCO results finetune pretrained CLIP weights that already exhibit a large gap (Fig. 6: distance starts high; Org maintains ~0.88). Thus the multimodal evidence shows that xNCE reduces an existing gap and improves transfer under finetuning, not that the same low-τ ranking-error mechanism induces the gap from scratch in high-dimensional image–text training. Either a from-scratch (or randomly re-initialized) multimodal run, or a carefully scoped claim limited to gap maintenance/reduction under finetuning, is needed for the central causal statement in the abstract and H1.
- [Tables 1–3, §4.2] Tables 1–3: xNCE is run with learnable τ initialized at 0.02 (clipped), while Org uses 0.01 and Reg uses 0.01. Given that temperature is itself a primary control on gap and geometry (§5.1, Fig. 2), the zero-shot and gap comparisons should include an Org ablation at the same τ schedule/init as Mix, or a joint sweep, so that gains cannot be attributed to the slight temperature difference rather than to mixed negatives.
minor comments (6)
- [Figure 2] Fig. 2 pseudocode: labels use loss_i / loss_t in the text but loss_j appears in the last line (loss = (loss_i + loss_j) / 2). Fix the variable name.
- [Acknowledgements] Acknowledgements: typo “futher” → “further”.
- [Abstract] Code availability is listed as “...” in the abstract; provide a concrete link or anonymized repo for review.
- [§3.1] §3.1 / Eq. (1): a_ij is introduced as cosine similarity in [−1,1], but the theoretical discussion sometimes treats a_ii → 1 and E[a_ik] = 0 as if they were simultaneous optima under dual encoders; a short remark that these are the uni-modal ideal (Wang & Isola) would avoid confusion.
- [Appendix D, Figure 7] Fig. 7 and §D: the sensitivity plot is useful; stating the exact r_mix grid and number of seeds in the caption would improve reproducibility.
- [§2] Related work cites Liang et al., Shi et al., Fahim et al. appropriately; a one-sentence contrast with Oh et al.’s modality-mix hard negatives (already cited) would clarify that xNCE mixes negatives inside the batch softmax rather than via geodesic mixup at finetuning time.
Circularity Check
No circularity: theory, controlled MNIST experiment, and COCO/zero-shot results are independent of one another and of any fitted target.
full rationale
The paper’s central claim (modality gap as low-temperature mode-failure of dual-encoder InfoNCE) is supported by three independent pillars that do not reduce to one another by construction. (1) The MNIST two-encoder experiment (identical architecture, identical initialization, same-family augmentations) is a controlled construction that actively generates a gap under InfoNCE; the gap is measured, not assumed. (2) The lim τ→0+ analysis of Li (Eq. 4) is a standard asymptotic expansion of the softmax loss under an irreducible ranking error; it derives the collapse of similarities toward their mean as a consequence of the loss, not as a definition of the gap. (3) xNCE is a simple change of the negative set that is then evaluated on external, held-out benchmarks (COCO retrieval R@K, four zero-shot classification datasets) whose metrics are independent of the loss derivation or of any fitted parameter. Temperature and mixing ratio are free hyperparameters swept on a grid; they are not fitted to force the claimed gap reduction. No self-citation is load-bearing, no uniqueness theorem is imported from the authors, and no known empirical pattern is merely renamed. The derivation chain is therefore self-contained against external evidence.
Axiom & Free-Parameter Ledger
free parameters (3)
- temperature τ =
0.01 (CLIP default / learned); 0.02 (xNCE initial); grid 0.01-1.0 (MNIST)
- mixing ratio r_mix =
0.01 / 0.05 / 0.5
- regularization weights λ_u, λ_xu, λ_a =
1
axioms (3)
- domain assumption InfoNCE with independent encoders admits a degenerate optimum in which all cross-modal similarities collapse toward their mean while preserving ranking (low-τ limit).
- domain assumption A dual-encoder MNIST experiment with identical weights and same-family augmentations isolates the contribution of the loss from data heterogeneity and architecture mismatch.
- standard math Cosine similarity on ℓ2-normalized embeddings and the symmetric InfoNCE objective are the correct geometric setting for analyzing the modality gap.
invented entities (1)
-
xNCE (Mixed Noise-Contrastive Estimation)
independent evidence
read the original abstract
We study the modality gap in CLIP-style dual-encoder contrastive learning, where image and text embeddings remain misaligned despite being trained in a shared space. We argue that the gap is induced by a failure of the InfoNCE formulation with independent encoders. We conduct a uni-modal experiment with two independent encoders and identical initialization conditions and find that InfoNCE actively generates a gap at low temperatures. We provide a theoretical analysis of this phenomenon and show that the modality gap is indeed a mode-failure of InfoNCE, but only at low temperatures. We propose a simple modification called xNCE, which uses intermodal as well as intra-modality negative contrastive pairs. xNCE matches retrieval performance on MS-COCO while consistently reducing the gap even at low temperatures. Notably, xNCE improves zero-shot classification over the InfoNCE baseline across all benchmarks, whereas high-temperature InfoNCE and regularized InfoNCE both fail to do so, demonstrating that xNCE reduces the modality gap without sacrificing the discriminative geometry needed for transfer.
Figures
Reference graph
Works this paper leans on
-
[1]
Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning , author =
-
[2]
International conference on machine learning , pages=
Understanding contrastive representation learning through alignment and uniformity on the hypersphere , author=. International conference on machine learning , pages=. 2020 , organization=
2020
-
[3]
Welle and M
Peiyang Shi and Michael C. Welle and M. Towards understanding the modality gap in. ICLR 2023 Workshop on Multimodal Representation Learning: Perks and Pitfalls , year=
2023
-
[4]
Learning Transferable Visual Models From Natural Language Supervision , issn =
Radford, Alec and Kim, Jong Wook and Hallacy, Chris and Ramesh, Aditya and Goh, Gabriel and Agarwal, Sandhini and Sastry, Girish and Askell, Amanda and Mishkin, Pamela and Clark, Jack and Krueger, Gretchen and Sutskever, Ilya , urldate =. Learning Transferable Visual Models From Natural Language Supervision , issn =. Proceedings of the 38th International ...
-
[5]
Mind the Gap: Preserving and Compensating for the Modality Gap in
Huang, Linlan and Cao, Xusheng and Lu, Haori and Meng, Yifan and Yang, Fei and Liu, Xialei , langid =. Mind the Gap: Preserving and Compensating for the Modality Gap in
-
[6]
Role, François and Meyer, Sébastien and Amblard, Victor , urldate =. Fill the Gap: Quantifying and Reducing the Modality Gap in Image-Text Representation Learning , url =. doi:10.48550/arXiv.2505.03703 , shorttitle =. 2505.03703 [cs] , keywords =
-
[7]
It's Not a Modality Gap: Characterizing and Addressing the Contrastive Gap , url =
Fahim, Abrar and Murphy, Alex and Fyshe, Alona , urldate =. It's Not a Modality Gap: Characterizing and Addressing the Contrastive Gap , url =. doi:10.48550/arXiv.2405.18570 , shorttitle =. 2405.18570 [cs] , keywords =
-
[8]
arXiv preprint arXiv:2406.17639 , year=
Mitigate the gap: Investigating approaches for improving cross-modal alignment in clip , author=. arXiv preprint arXiv:2406.17639 , year=
-
[9]
The Double-Ellipsoid Geometry of
Levi, Meir Yossef and Gilboa, Guy , urldate =. The Double-Ellipsoid Geometry of. doi:10.48550/arXiv.2411.14517 , abstract =. 2411.14517 [cs] , keywords =
-
[10]
Understanding the Behaviour of Contrastive Loss , rights =
Wang, Feng and Liu, Huaping , urldate =. Understanding the Behaviour of Contrastive Loss , rights =. 2021. doi:10.1109/CVPR46437.2021.00252 , abstract =
-
[11]
IEEE Transactions on Artificial Intelligence , volume=
Dynamically scaled temperature in self-supervised contrastive learning , author=. IEEE Transactions on Artificial Intelligence , volume=. 2025 , publisher=
2025
-
[12]
Wu, Zhirong and Xiong, Yuanjun and Yu, Stella X. and Lin, Dahua , urldate =. Unsupervised Feature Learning via Non-parametric Instance Discrimination , isbn =. 2018. doi:10.1109/CVPR.2018.00393 , abstract =
-
[13]
Representation Learning with Contrastive Predictive Coding , url =
Oord, Aaron van den and Li, Yazhe and Vinyals, Oriol , urldate =. Representation Learning with Contrastive Predictive Coding , url =. doi:10.48550/arXiv.1807.03748 , abstract =. 1807.03748 [cs] , keywords =
-
[14]
Learning Representations by Maximizing Mutual Information Across Views , volume =
Bachman, Philip and Hjelm, R Devon and Buchwalter, William , urldate =. Learning Representations by Maximizing Mutual Information Across Views , volume =. Advances in Neural Information Processing Systems , publisher =
-
[15]
A Simple Framework for Contrastive Learning of Visual Representations , url =
Chen, Ting and Kornblith, Simon and Norouzi, Mohammad and Hinton, Geoffrey , urldate =. A Simple Framework for Contrastive Learning of Visual Representations , url =. doi:10.48550/arXiv.2002.05709 , abstract =. 2002.05709 [cs] , keywords =
-
[16]
Momentum
He, Kaiming and Fan, Haoqi and Wu, Yuxin and Xie, Saining and Girshick, Ross , abstract =. Momentum
-
[17]
doi:10.48550/arXiv.1702.05373 , shorttitle =
Cohen, Gregory and Afshar, Saeed and Tapson, Jonathan and Schaik, André van , urldate =. doi:10.48550/arXiv.1702.05373 , shorttitle =. 1702.05373 [cs] , note =
-
[18]
Lawrence and Dollár, Piotr , urldate =
Lin, Tsung-Yi and Maire, Michael and Belongie, Serge and Bourdev, Lubomir and Girshick, Ross and Hays, James and Perona, Pietro and Ramanan, Deva and Zitnick, C. Lawrence and Dollár, Piotr , urldate =. Microsoft. doi:10.48550/arXiv.1405.0312 , shorttitle =. 1405.0312 [cs] , keywords =
-
[19]
Geodesic Multi-Modal Mixup for Robust Fine-Tuning
Oh, Changdae and So, Junhyuk and Byun, Hoyoon and Lim, YongTaek and Shin, Minchul and Jeon, Jong-June and Song, Kyungwoo , date =. Geodesic. doi:10.48550/arXiv.2203.03897 , url =. 2203.03897 , eprinttype =
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.2203.03897
-
[20]
Second Workshop on Representational Alignment at ICLR 2025 , year=
Closing the modality gap enables novel multimodal learning applications , author=. Second Workshop on Representational Alignment at ICLR 2025 , year=
2025
-
[21]
2020 , eprint=
UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction , author=. 2020 , eprint=
2020
-
[22]
2018 , eprint=
VSE++: Improving Visual-Semantic Embeddings with Hard Negatives , author=. 2018 , eprint=
2018
-
[23]
2015 , eprint=
Deep Visual-Semantic Alignments for Generating Image Descriptions , author=. 2015 , eprint=
2015
-
[24]
Delphboy/Karpathy-Splits , author =
-
[25]
arXiv preprint arXiv:2010.11929 , year=
An image is worth 16x16 words: Transformers for image recognition at scale , author=. arXiv preprint arXiv:2010.11929 , year=
Pith/arXiv arXiv 2010
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.