Pith. sign in

REVIEW 6 cited by

Words or Vision: Do Vision-Language Models Have Blind Faith in Text?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.02199 v1 pith:4C24DVGP submitted 2025-03-04 cs.CV cs.AIcs.CLcs.LGcs.MM

classification cs.CVcs.AIcs.CLcs.LGcs.MM
keywords textdatatextualvlmsmodelsvisualbiasblind
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Vision-Language Models (VLMs) excel in integrating visual and textual information for vision-centric tasks, but their handling of inconsistencies between modalities is underexplored. We investigate VLMs' modality preferences when faced with visual data and varied textual inputs in vision-centered settings. By introducing textual variations to four vision-centric tasks and evaluating ten Vision-Language Models (VLMs), we discover a \emph{``blind faith in text''} phenomenon: VLMs disproportionately trust textual data over visual data when inconsistencies arise, leading to significant performance drops under corrupted text and raising safety concerns. We analyze factors influencing this text bias, including instruction prompts, language model size, text relevance, token order, and the interplay between visual and textual certainty. While certain factors, such as scaling up the language model size, slightly mitigate text bias, others like token order can exacerbate it due to positional biases inherited from language models. To address this issue, we explore supervised fine-tuning with text augmentation and demonstrate its effectiveness in reducing text bias. Additionally, we provide a theoretical analysis suggesting that the blind faith in text phenomenon may stem from an imbalance of pure text and multi-modal data during training. Our findings highlight the need for balanced training and careful consideration of modality interactions in VLMs to enhance their robustness and reliability in handling multi-modal data inconsistencies.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Can Large Multimodal Models Actively Recognize Faulty Inputs? A Systematic Evaluation Framework of Their Input Scrutiny Ability

    cs.CV 2025-08 unverdicted novelty 6.0 of 10

    Large multimodal models mostly fail to proactively detect flawed textual premises, and their performance depends on error type and on how they weight text versus images.

  2. How Do Vision-Language Models Process Conflicting Information Across Modalities?

    cs.CL 2025-07 conditional novelty 6.0 of 10

    Vision-language models answer from whichever modality is encoded more saliently in their final-layer representations, and specific attention heads can be manipulated to shift that preference.

  3. Devil in the Lens: Analyzing and Defending Physical Prompt Injection Against Vision-Language Models on Wearable Devices

    cs.CR 2026-07 conditional novelty 5.0 of 10

    Physical scene text can inject prompts into wearable VLMs, hijacking decisions and content with high success rates across six threat scenarios, partially mitigated by OCR masking and token-drift defenses.

  4. Innocence in the Crossfire: Roles of Skip Connections in Jailbreaking Visual Language Models

    cs.CL 2025-07 conditional novelty 5.0 of 10

    The paper reports higher harmful-output rates in three open-source VLMs from detailed image descriptions, in-context examples, and positive openings, and from a skip connection between internal layers, with memes riva...

  5. Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints

    cs.LG 2025-06 conditional novelty 5.0 of 10

    A 3B VLM trained with GRPO to call a zoom tool improves V*Bench accuracy by 5.7% over its base model but degrades TextVQA and HR-Bench performance.

  6. Challenges in Understanding Modality Conflict in Vision-Language Models

    cs.LG 2025-09 conditional novelty 4.0 of 10

    In LLaVA-OV-7B, a linearly decodable conflict signal appears in intermediate layers and detection-related attention shifts precede resolution-related ones, supporting a detection/resolution separation in the model.

Pith tools