Pith. sign in

REVIEW 10 cited by

LVLM-Interpret: An Interpretability Tool for Large Vision-Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.03118 v3 pith:K6YTCBSM submitted 2024-04-03 cs.CV

classification cs.CV
keywords largemechanismsmodelsapplicationmodelunderstandingimageinternal
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

In the rapidly evolving landscape of artificial intelligence, multi-modal large language models are emerging as a significant area of interest. These models, which combine various forms of data input, are becoming increasingly popular. However, understanding their internal mechanisms remains a complex task. Numerous advancements have been made in the field of explainability tools and mechanisms, yet there is still much to explore. In this work, we present a novel interactive application aimed towards understanding the internal mechanisms of large vision-language models. Our interface is designed to enhance the interpretability of the image patches, which are instrumental in generating an answer, and assess the efficacy of the language model in grounding its output in the image. With our application, a user can systematically investigate the model and uncover system limitations, paving the way for enhancements in system capabilities. Finally, we present a case study of how our application can aid in understanding failure mechanisms in a popular large multi-modal model: LLaVA.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMs

    cs.CV 2025-11 unverdicted novelty 8.0 of 10

    MVI-Bench supplies the first taxonomy and dataset focused on misleading visual inputs to measure LVLM robustness, with tests on 18 models revealing clear weaknesses.

  2. Multimodal Model Diffing for Feature Discovery and Control

    cs.CV 2026-08 conditional novelty 7.0 of 10

    By diffing base-language and multimodal sparse autoencoder features, MMDiff isolates causally relevant features that can be ablated or steered to control spatial, OCR, and safety behaviors in multimodal LLMs.

  3. Backdoor Attack on Vision Language Models with Stealthy Semantic Manipulation

    cs.CV 2025-06 conditional novelty 7.0 of 10

    BadSem shows that semantic mismatches between images and text can serve as stealthy backdoor triggers for VLMs, achieving near-perfect attack success with low poisoning rates.

  4. SurgXBench: Explainable Vision-Language Model Benchmark for Surgery

    cs.CV 2025-05 conditional novelty 6.0 of 10

    An explainability-based benchmark showing that surgical vision-language models often make correct predictions without attending to the relevant instruments or tissue.

  5. Cross-modal Information Flow in Multimodal Large Language Models

    cs.AI 2024-11 conditional novelty 6.0 of 10

    In LLaVA multimodal models, visual information flows into question token representations in two stages, global then object-specific, before propagating to the final answer position.

  6. Understanding Multimodal LLMs: the Mechanistic Interpretability of Llava in Visual Question Answering

    cs.CL 2024-11 conditional novelty 6.0 of 10

    The paper shows LLaVA's visual QA mechanism parallels textual QA: visual embeddings encode animal and color features, attention heads extract and match them, and visual instruction tuning refines existing Vicuna heads.

  7. GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A gradient-attention explainability method produces sequence-level visual and textual saliency maps for free-form answers from large vision-language models, with stronger human-attention alignment and faithfulness tha...

  8. On the Risk of Misleading Reports: Diagnosing Textual Biases in Multimodal Clinical AI

    cs.CV 2025-07 conditional novelty 4.0 of 10

    A new perturbation test shows that medical vision-language models rely more on clinical text than on images, with calibration errors growing when text conflicts with the image.

  9. Adapting Lightweight Vision Language Models for Radiological Visual Question Answering

    cs.CV 2025-06 reject novelty 4.0 of 10

    A 3B PaliGemma model fine-tuned with synthetic QA pairs and two-stage training reaches 41.5% accuracy on open-ended radiology VQA, about 15 points below LLaVA-Med.

  10. Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey

    cs.CL 2024-12 conditional novelty 4.0 of 10

    A survey maps the field of MLLM explainability and interpretability into data, model, and training and inference perspectives.

Pith tools