Pith. sign in

REVIEW 2 cited by

Identifying and Mitigating Position Bias of Multi-image Vision-Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.13792 v1 pith:JAJHWQZ7 submitted 2025-03-18 cs.CV

classification cs.CV
keywords biaslvlmsmodelspositionreasoningimagesattentionbeginning
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The evolution of Large Vision-Language Models (LVLMs) has progressed from single to multi-image reasoning. Despite this advancement, our findings indicate that LVLMs struggle to robustly utilize information across multiple images, with predictions significantly affected by the alteration of image positions. To further explore this issue, we introduce Position-wise Question Answering (PQA), a meticulously designed task to quantify reasoning capabilities at each position. Our analysis reveals a pronounced position bias in LVLMs: open-source models excel in reasoning with images positioned later but underperform with those in the middle or at the beginning, while proprietary models show improved comprehension for images at the beginning and end but struggle with those in the middle. Motivated by this, we propose SoFt Attention (SoFA), a simple, training-free approach that mitigates this bias by employing linear interpolation between inter-image causal attention and bidirectional counterparts. Experimental results demonstrate that SoFA reduces position bias and enhances the reasoning performance of existing LVLMs.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. PeRL: Permutation-Enhanced Reinforcement Learning for Interleaved Vision-Language Reasoning

    cs.CV 2025-06 conditional novelty 6.0 of 10

    PeRL applies reinforcement learning to a vision-language model with image-order permutation and difficulty-based data filtering, improving multi-image reasoning while keeping single-image performance.

  2. Zooming from Context to Cue: Hierarchical Preference Optimization for Multi-Image MLLMs

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Context-to-Cue Direct Preference Optimization (CcDPO) reduces multi-image hallucinations in 7B multimodal LLMs by training on perturbed full-sequence captions and region-focused visual prompts, improving average multi...

Pith tools