Pith. sign in

Modality Gap-Driven Subspace Alignment Training Paradigm For Multimodal Large Language Models

9 Pith papers cite this work. Polarity classification is still indexing.

9 Pith papers citing it
abstract

Despite the success of multimodal contrastive learning in aligning visual and linguistic representations, a persistent geometric anomaly, the Modality Gap, remains: embeddings of distinct modalities expressing identical semantics occupy systematically offset regions. Prior approaches to bridge this gap are largely limited by oversimplified isotropic assumptions, hindering their application in large-scale scenarios. In this paper, we address these limitations by precisely characterizing the geometric shape of the modality gap and leveraging it for efficient model scaling. First, we propose the Fixed-frame Modality Gap Theory, which decomposes the modality gap within a frozen reference frame into stable biases and anisotropic residuals. Guided by this precise modeling, we introduce ReAlign, a training-free modality alignment strategy. Utilizing statistics from massive unpaired data, ReAlign aligns text representation into the image representation distribution via a three-step process comprising Anchor, Trace, and Centroid Alignment, thereby explicitly rectifying geometric misalignment. Building on ReAlign, we propose ReVision, a scalable training paradigm for Multimodal Large Language Models~(MLLMs). ReVision integrates ReAlign into the pretraining stage, enabling the model to learn the distribution of visual representations from unpaired text before visual instruction tuning, without the need for large-scale, high-quality image-text pairs. Our framework demonstrates that statistically aligned unpaired data can effectively substitute for expensive image-text pairs, offering a robust path for the efficient scaling of MLLMs.

citation-role summary

background 1

citation-polarity summary

years

2026 9

roles

background 1

polarities

background 1

representative citing papers

Multi-scale Object-Aware Gaze Estimation via Geometric Reasoning

cs.CV · 2026-06-28 · unverdicted · novelty 6.0

A two-stage object-aware gaze estimation method with multi-scale feature fusion and geometric constraints reports AUC scores of 0.961, 0.948, 0.987, and 0.977 on GazeFollow, VideoAttentionTarget, ChildPlay, and GOO-Real with a 7.1M parameter model.

CARE: Competence-Aware Reward Shaping for Adaptive Reasoning Length in Video-MLLMs

cs.CV · 2026-06-18 · unverdicted · novelty 6.0

CARE uses exponential moving average competence estimates to progressively shift RL rewards from exploration-oriented long reasoning to efficiency-oriented concise reasoning in video-MLLMs, with batch normalization and posterior amplification, yielding accuracy gains and shorter traces.

Anisotropic Modality Align

cs.MM · 2026-05-08 · unverdicted · novelty 6.0

Modality representations share dominant semantic geometry but have an anisotropic residual gap; AnisoAlign corrects source representations boundedly using target geometry for unpaired alignment.

citing papers explorer

Showing 9 of 9 citing papers.