UniCVR is the first unified zero-shot framework that handles composed image, multi-turn image, and video retrieval by MLLM-VLP alignment plus dual-level reranking.
Modality Gap-Driven Subspace Alignment Training Paradigm For Multimodal Large Language Models
9 Pith papers cite this work. Polarity classification is still indexing.
abstract
Despite the success of multimodal contrastive learning in aligning visual and linguistic representations, a persistent geometric anomaly, the Modality Gap, remains: embeddings of distinct modalities expressing identical semantics occupy systematically offset regions. Prior approaches to bridge this gap are largely limited by oversimplified isotropic assumptions, hindering their application in large-scale scenarios. In this paper, we address these limitations by precisely characterizing the geometric shape of the modality gap and leveraging it for efficient model scaling. First, we propose the Fixed-frame Modality Gap Theory, which decomposes the modality gap within a frozen reference frame into stable biases and anisotropic residuals. Guided by this precise modeling, we introduce ReAlign, a training-free modality alignment strategy. Utilizing statistics from massive unpaired data, ReAlign aligns text representation into the image representation distribution via a three-step process comprising Anchor, Trace, and Centroid Alignment, thereby explicitly rectifying geometric misalignment. Building on ReAlign, we propose ReVision, a scalable training paradigm for Multimodal Large Language Models~(MLLMs). ReVision integrates ReAlign into the pretraining stage, enabling the model to learn the distribution of visual representations from unpaired text before visual instruction tuning, without the need for large-scale, high-quality image-text pairs. Our framework demonstrates that statistically aligned unpaired data can effectively substitute for expensive image-text pairs, offering a robust path for the efficient scaling of MLLMs.
citation-role summary
citation-polarity summary
years
2026 9roles
background 1polarities
background 1representative citing papers
ProjAgent introduces procedural similarity—retrieving code with matching computational logic—via LLM hidden-state projections, improving repository-level code generation to 41.14% Pass@1 on REPOCOD.
A two-stage object-aware gaze estimation method with multi-scale feature fusion and geometric constraints reports AUC scores of 0.961, 0.948, 0.987, and 0.977 on GazeFollow, VideoAttentionTarget, ChildPlay, and GOO-Real with a 7.1M parameter model.
CARE uses exponential moving average competence estimates to progressively shift RL rewards from exploration-oriented long reasoning to efficiency-oriented concise reasoning in video-MLLMs, with batch normalization and posterior amplification, yielding accuracy gains and shorter traces.
Modality representations share dominant semantic geometry but have an anisotropic residual gap; AnisoAlign corrects source representations boundedly using target geometry for unpaired alignment.
Decoder-based VLMs over-align visual embeddings to text manifold causing linguistic bias in top PCs of a universal text subspace; projecting out this subspace reduces hallucinations on POPE/CHAIR/AMBER and improves CLAIR.
The survey formalizes MLLM perception as a unified vision-language capability and traces its evolution via a new five-stage taxonomy while outlining future challenges.
Spectral decomposition of VLM embedding covariances isolates a shared noise subspace whose removal leaves or boosts task performance across data subsets.
A two-level reference alignment framework uses complete-modality samples and prototype voting to reduce decision drift and improve robustness in multimodal sentiment analysis under missing modalities.
citing papers explorer
-
UniCVR: From Alignment to Reranking for Unified Zero-Shot Composed Visual Retrieval
UniCVR is the first unified zero-shot framework that handles composed image, multi-turn image, and video retrieval by MLLM-VLP alignment plus dual-level reranking.
-
ProjAgent: Procedural Similarity Retrieval for Repository-Level Code Generation
ProjAgent introduces procedural similarity—retrieving code with matching computational logic—via LLM hidden-state projections, improving repository-level code generation to 41.14% Pass@1 on REPOCOD.
-
Multi-scale Object-Aware Gaze Estimation via Geometric Reasoning
A two-stage object-aware gaze estimation method with multi-scale feature fusion and geometric constraints reports AUC scores of 0.961, 0.948, 0.987, and 0.977 on GazeFollow, VideoAttentionTarget, ChildPlay, and GOO-Real with a 7.1M parameter model.
-
CARE: Competence-Aware Reward Shaping for Adaptive Reasoning Length in Video-MLLMs
CARE uses exponential moving average competence estimates to progressively shift RL rewards from exploration-oriented long reasoning to efficiency-oriented concise reasoning in video-MLLMs, with batch normalization and posterior amplification, yielding accuracy gains and shorter traces.
-
Anisotropic Modality Align
Modality representations share dominant semantic geometry but have an anisotropic residual gap; AnisoAlign corrects source representations boundedly using target geometry for unpaired alignment.
-
When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models
Decoder-based VLMs over-align visual embeddings to text manifold causing linguistic bias in top PCs of a universal text subspace; projecting out this subspace reduces hallucinations on POPE/CHAIR/AMBER and improves CLAIR.
-
From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models
The survey formalizes MLLM perception as a unified vision-language capability and traces its evolution via a new five-stage taxonomy while outlining future challenges.
-
Your CLIP has 164 dimensions of noise: Exploring the embeddings covariance eigenspectrum of contrastively pretrained vision-language transformers
Spectral decomposition of VLM embedding covariances isolates a shared noise subspace whose removal leaves or boosts task performance across data subsets.
-
Controlling Decision Drift in Multimodal Sentiment Analysis with Missing Modalities
A two-level reference alignment framework uses complete-modality samples and prototype voting to reduce decision drift and improve robustness in multimodal sentiment analysis under missing modalities.