Pith. sign in

REVIEW 5 cited by

Mitigate the Gap: Investigating Approaches for Improving Cross-Modal Alignment in CLIP

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.17639 v3 pith:KJ43RWG2 submitted 2024-06-25 cs.CV cs.AIcs.CLcs.LG

classification cs.CVcs.AIcs.CLcs.LG
keywords clipcross-modalmodalityspacealignclipalignmentapproachesdownstream
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Contrastive Language--Image Pre-training (CLIP) has manifested remarkable improvements in zero-shot classification and cross-modal vision-language tasks. Yet, from a geometrical point of view, the CLIP embedding space has been found to have a pronounced modality gap. This gap renders the embedding space overly sparse and disconnected, with different modalities being densely distributed in distinct subregions of the hypersphere. In this work, we aim at answering three main questions: 1. Does sharing the parameter space between the multi-modal encoders reduce the modality gap? 2. Can the gap be mitigated by pushing apart the uni-modal embeddings via intra-modality separation? 3. How do these gap reduction approaches affect the downstream performance? We design AlignCLIP, in order to answer these questions and through extensive experiments, we show that AlignCLIP achieves noticeable enhancements in the cross-modal alignment of the embeddings, and thereby, reduces the modality gap, while improving the performance across several zero-shot and fine-tuning downstream evaluations.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs

    cs.CL 2026-05 conditional novelty 7.0 of 10

    TokenSwap measures and mitigates the MLLM modality gap: swapping textual concepts for matched images lowers accuracy by 4-47% across 42 models, and training with such swaps reduces the gap.

  2. On the modality gap and the contrastive loss in multi-modal representation learning

    cs.LG 2026-07 conditional novelty 6.5 of 10

    InfoNCE with independent encoders actively creates a modality gap at low temperature; mixing intra- and inter-modality negatives (xNCE) removes the gap while improving zero-shot transfer.

  3. ScalablePromptus: Scalable and High-Fidelity Prompt-Based Video Streaming

    eess.IV 2026-07 conditional novelty 6.0 of 10

    Dropout-trained, rank-ordered prompt embeddings let a video receiver reconstruct useful frames from truncated prompts, cutting truncation-induced LPIPS degradation by 82–95% versus Promptus.

  4. Mind the Gap: Preserving and Compensating for the Modality Gap in CLIP-Based Continual Learning

    cs.CV 2025-07 conditional novelty 6.0 of 10

    MG-CLIP preserves CLIP's modality gap by adaptively limiting fine-tuning epochs and compensates for its limits with a visual-space classifier, improving class-incremental learning without replay.

  5. AspectCLIP: Optimizing CLIP Representation Space via Aspect-Guided Consistency Regularization

    cs.CV 2026-07 reject novelty 5.0 of 10

    AspectCLIP partitions captions into four SimCSE-based clusters and restricts cyclic consistency regularization to within clusters, reporting modest gains over CyCLIP on several CLIP benchmarks.

Pith tools