REVIEW 5 cited by
Mitigate the Gap: Investigating Approaches for Improving Cross-Modal Alignment in CLIP
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Contrastive Language--Image Pre-training (CLIP) has manifested remarkable improvements in zero-shot classification and cross-modal vision-language tasks. Yet, from a geometrical point of view, the CLIP embedding space has been found to have a pronounced modality gap. This gap renders the embedding space overly sparse and disconnected, with different modalities being densely distributed in distinct subregions of the hypersphere. In this work, we aim at answering three main questions: 1. Does sharing the parameter space between the multi-modal encoders reduce the modality gap? 2. Can the gap be mitigated by pushing apart the uni-modal embeddings via intra-modality separation? 3. How do these gap reduction approaches affect the downstream performance? We design AlignCLIP, in order to answer these questions and through extensive experiments, we show that AlignCLIP achieves noticeable enhancements in the cross-modal alignment of the embeddings, and thereby, reduces the modality gap, while improving the performance across several zero-shot and fine-tuning downstream evaluations.
Forward citations
Cited by 5 Pith papers
-
TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs
TokenSwap measures and mitigates the MLLM modality gap: swapping textual concepts for matched images lowers accuracy by 4-47% across 42 models, and training with such swaps reduces the gap.
-
On the modality gap and the contrastive loss in multi-modal representation learning
InfoNCE with independent encoders actively creates a modality gap at low temperature; mixing intra- and inter-modality negatives (xNCE) removes the gap while improving zero-shot transfer.
-
ScalablePromptus: Scalable and High-Fidelity Prompt-Based Video Streaming
Dropout-trained, rank-ordered prompt embeddings let a video receiver reconstruct useful frames from truncated prompts, cutting truncation-induced LPIPS degradation by 82–95% versus Promptus.
-
Mind the Gap: Preserving and Compensating for the Modality Gap in CLIP-Based Continual Learning
MG-CLIP preserves CLIP's modality gap by adaptively limiting fine-tuning epochs and compensates for its limits with a visual-space classifier, improving class-incremental learning without replay.
-
AspectCLIP: Optimizing CLIP Representation Space via Aspect-Guided Consistency Regularization
AspectCLIP partitions captions into four SimCSE-based clusters and restricts cyclic consistency regularization to within clusters, reporting modest gains over CyCLIP on several CLIP benchmarks.
Discussion (0). Continue with ORCID to comment.