Pith. sign in

REVIEW 8 cited by

Unraveling Cross-Modality Knowledge Conflicts in Large Vision-Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.03659 v2 pith:CGJTKNHI submitted 2024-10-04 cs.CV cs.CL

Unraveling Cross-Modality Knowledge Conflicts in Large Vision-Language Models

classification cs.CV cs.CL
keywords conflictsknowledgemodelscontrastiveaccuracycomponentsconflictcross-modality
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
abstract

Large Vision-Language Models (LVLMs) have demonstrated impressive capabilities for capturing and reasoning over multimodal inputs. However, these models are prone to parametric knowledge conflicts, which arise from inconsistencies of represented knowledge between their vision and language components. In this paper, we formally define the problem of $\textbf{cross-modality parametric knowledge conflict}$ and present a systematic approach to detect, interpret, and mitigate them. We introduce a pipeline that identifies conflicts between visual and textual answers, showing a persistently high conflict rate across modalities in recent LVLMs regardless of the model size. We further investigate how these conflicts interfere with the inference process and propose a contrastive metric to discern the conflicting samples from the others. Building on these insights, we develop a novel dynamic contrastive decoding method that removes undesirable logits inferred from the less confident modality components based on answer confidence. For models that do not provide logits, we also introduce two prompt-based strategies to mitigate the conflicts. Our methods achieve promising improvements in accuracy on both the ViQuAE and InfoSeek datasets. Specifically, using LLaVA-34B, our proposed dynamic contrastive decoding improves an average accuracy of 2.24%.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. When Correct Decisions Hide Internal Stress: Decision-State Probing in Multimodal Language Models

    cs.CL 2026-06 unverdicted novelty 7.0

    S³E framework finds excess decision-state displacement under semantic stress in multimodal models despite consistent correct forced-choice behavior.

  2. ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors?

    cs.CV 2026-06 unverdicted novelty 7.0

    ChronoPhyBench is a new benchmark and dataset for chronological physical dynamics reasoning that combines video-conditioned next-state prediction with VQA to reduce language bias in MLLM evaluation.

  3. ENTRAP-VL: A Taxonomic Probe for Dual Contextual Entrainment in Vision-Language Models

    cs.CV 2026-07 conditional novelty 6.0

    A new taxonomy and 1,500-item dataset, ENTRAP-VL, lets researchers measure whether vision-language models are entrained by textual and visual context separately.

  4. Linguistic Context Recodes Visual Representations in Vision-Language Models

    cs.AI 2026-07 conditional novelty 6.0

    Goal-directed language prompts make VLMs add a transferable goal-relevant marker to selected image objects and amplify those objects' queried attributes in later layers, and both effects causally influence answers.

  5. MLLMs Get It Right, Then Get It Wrong: Tracing and Correcting Late-Layer Textual Bias

    cs.CV 2026-06 unverdicted novelty 6.0

    MLLMs show late-layer textual override of correct visual predictions, with a directional signature enabling a simple inference-time recovery method that improves conflict benchmarks by up to 9.4%.

  6. RIHA: Report-Image Hierarchical Alignment for Radiology Report Generation

    cs.CV 2026-04 unverdicted novelty 6.0

    RIHA proposes a hierarchical alignment transformer that uses multi-scale visual and textual feature pyramids plus optimal transport to generate more accurate radiology reports from medical images.

  7. Causal Evidence for Attention Head Imbalance in Modality Conflict Hallucination

    cs.AI 2026-05 unverdicted novelty 5.0

    Causal path-patching analysis across five MLLMs identifies distributed hallucination-driving attention heads and localized resisting heads whose imbalance biases generation toward erroneous text over visual evidence; ...

  8. Foundation Models for Astrophysics

    astro-ph.IM 2026-08 conditional novelty 3.0

    Astronomical 'foundation models' largely reuse transformers and self-supervised pretraining, but evidence of transfer to new instruments, populations, or tasks remains rare; the paper argues such evidence, not archite...