Pith. sign in

REVIEW 6 cited by

Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2203.02053 v2 pith:GKZOGVER submitted 2022-03-03 cs.CL cs.AIcs.CVcs.LGcs.MM

classification cs.CLcs.AIcs.CVcs.LGcs.MM
keywords modelmulti-modalrepresentationcontrastivelearningmodalitiesmodalitydata
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present modality gap, an intriguing geometric phenomenon of the representation space of multi-modal models. Specifically, we show that different data modalities (e.g. images and text) are embedded at arm's length in their shared representation in multi-modal models such as CLIP. Our systematic analysis demonstrates that this gap is caused by a combination of model initialization and contrastive learning optimization. In model initialization, we show empirically and theoretically that the representation of a common deep neural network is restricted to a narrow cone. As a consequence, in a multi-modal model with two encoders, the representations of the two modalities are clearly apart when the model is initialized. During optimization, contrastive learning keeps the different modalities separate by a certain distance, which is influenced by the temperature parameter in the loss function. Our experiments further demonstrate that varying the modality gap distance has a significant impact in improving the model's downstream zero-shot classification performance and fairness. Our code and data are available at https://modalitygap.readthedocs.io/

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 98 citations worldwide. Full citation record

  1. DETR-ViP: Detection Transformer with Robust Discriminative Visual Prompts

    cs.CV 2026-04 unverdicted novelty 7.0 of 10

    Global prompt integration, visual-textual relation distillation and selective fusion make visual prompts discriminative enough for DETR-ViP to beat prior visual-prompt detectors by several mAP points.

  2. QuASH: Using Natural-Language Heuristics to Query Visual-Language Robotic Maps

    cs.RO 2025-10 conditional novelty 6.0 of 10

    Querying VLM robot maps with an SVM trained on LLM-generated synonym/antonym embeddings outperforms cosine-threshold and single-antonym baselines on images and OpenSeg maps, but not consistently on LSeg maps.

  3. The Hyperspherical Geometry of CLIP Latent Space: A Semantic Mixture Model

    cs.LG 2026-07 conditional novelty 5.0 of 10

    CLIP embeddings are modeled as a mixture of von Mises-Fisher distributions on the unit sphere, improving out-of-distribution detection and semantic decomposition over single-Gaussian baselines.

  4. Sparse Neuron Ablation Triggers Catastrophic Collapse of the Language Core in Large Vision-Language Models

    cs.AI 2025-11 conditional novelty 5.0 of 10

    Ablating just four neurons in LLaVA-1.5-7b's language-model down-projection layer triggers complete output collapse, with critical neurons concentrated in the language backbone.

  5. OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A shared-backbone transformer with pairwise modality training reports top results across 25 datasets spanning 12 modalities.

  6. Foundation Models for Astrophysics

    astro-ph.IM 2026-08 conditional novelty 3.0 of 10

    Astronomical 'foundation models' largely reuse transformers and self-supervised pretraining, but evidence of transfer to new instruments, populations, or tasks remains rare; the paper argues such evidence, not archite...

Pith tools