Pith. sign in

REVIEW 6 cited by

Insight Over Sight: Exploring the Vision-Knowledge Conflicts in Multimodal LLMs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.08145 v2 pith:4U3XOM73 submitted 2024-10-10 cs.CL cs.CV

classification cs.CLcs.CV
keywords benchmarkconflictsframeworkmllmsvision-knowledgeconflictcommonsenseevaluate
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper explores the problem of commonsense level vision-knowledge conflict in Multimodal Large Language Models (MLLMs), where visual information contradicts model's internal commonsense knowledge. To study this issue, we introduce an automated framework, augmented with human-in-the-loop quality control, to generate inputs designed to simulate and evaluate these conflicts in MLLMs. Using this framework, we have crafted a diagnostic benchmark consisting of 374 original images and 1,122 high-quality question-answer (QA) pairs. The benchmark covers two aspects of conflict and three question types, providing a thorough assessment tool. We apply this benchmark to assess the conflict-resolution capabilities of nine representative MLLMs from various model families. Our results indicate an evident over-reliance on parametric knowledge for approximately 20% of all queries, especially among Yes-No and action-related problems. Based on these findings, we evaluate the effectiveness of existing approaches to mitigating the conflicts and compare them to our "Focus-on-Vision" prompting strategy. Despite some improvement, the vision-knowledge conflict remains unresolved and can be further scaled through our data construction framework. Our proposed framework, benchmark, and analysis contribute to the understanding and mitigation of vision-knowledge conflicts in MLLMs.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Prior Bias in Vision Language Models on UML Diagram Interpretation

    cs.CV 2026-07 conditional novelty 6.5 of 10

    Reversing only the UML relation arrow while keeping class names and layout fixed cuts open-source VLM relation accuracy by about 33%, revealing prior-over-vision bias.

  2. ENTRAP-VL: A Taxonomic Probe for Dual Contextual Entrainment in Vision-Language Models

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A new taxonomy and 1,500-item dataset, ENTRAP-VL, lets researchers measure whether vision-language models are entrained by textual and visual context separately.

  3. Benchmarking Multimodal Knowledge Conflict for Large Multimodal Models

    cs.LG 2025-05 conditional novelty 6.0 of 10

    MMKC-Bench provides a human-verified benchmark of multimodal knowledge conflicts and shows that current LMMs prefer internal parametric knowledge over external evidence.

  4. MissingBench-Verified: Probing Vision-Language Models' Inability to Detect Missing Object Parts

    cs.CV 2026-07 conditional novelty 5.0 of 10

    Ten leading VLMs mostly fail to report removed essential object parts as missing, and simulated detector evidence, image tools, longer reasoning, and an easier fine-tune barely improve accuracy.

  5. Robust Multimodal Large Language Models Against Modality Conflict

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A new benchmark, MMMC, triggers hallucinations in all tested multimodal models, and reinforcement learning on it reduces such hallucinations more than prompt engineering or supervised fine-tuning.

  6. MLLMs are Deeply Affected by Modality Bias

    cs.AI 2025-05 conditional novelty 4.0 of 10

    A position paper with a case study showing that multimodal LLMs rely on language priors and underuse visual input, together with a research roadmap and calls for balanced training.

Pith tools