REVIEW 15 cited by
Towards Interpreting Visual Information Processing in Vision-Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Vision-Language Models (VLMs) are powerful tools for processing and understanding text and images. We study the processing of visual tokens in the language model component of LLaVA, a prominent VLM. Our approach focuses on analyzing the localization of object information, the evolution of visual token representations across layers, and the mechanism of integrating visual information for predictions. Through ablation studies, we demonstrated that object identification accuracy drops by over 70\% when object-specific tokens are removed. We observed that visual token representations become increasingly interpretable in the vocabulary space across layers, suggesting an alignment with textual tokens corresponding to image content. Finally, we found that the model extracts object information from these refined representations at the last token position for prediction, mirroring the process in text-only language models for factual association tasks. These findings provide crucial insights into how VLMs process and integrate visual information, bridging the gap between our understanding of language and vision models, and paving the way for more interpretable and controllable multimodal systems.
Forward citations
Cited by 15 Pith papers
-
MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMs
MVI-Bench supplies the first taxonomy and dataset focused on misleading visual inputs to measure LVLM robustness, with tests on 18 models revealing clear weaknesses.
-
Multimodal Model Diffing for Feature Discovery and Control
By diffing base-language and multimodal sparse autoencoder features, MMDiff isolates causally relevant features that can be ablated or steered to control spatial, OCR, and safety behaviors in multimodal LLMs.
-
What's in the Image? A Deep-Dive into the Vision of Vision Language Models
Vision-language models store a global image summary in the query text tokens, rely on the middle transformer layers for vision-to-text transfer, and fetch fine details from image tokens in a spatially localized way.
-
How Do VLMs Fail? Vision-Operation Misalignment in Compositional VQA
The paper proposes four operation-level VLM failure modes and a pathway dissociation, but the dissociation is not supported by the paper's own intervention statistics.
-
Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders
A lightweight Q-Former proxy trained on VLM hidden states reveals that localization signals peak in input-dependent intermediate layers, not the final layers used by standard editing pipelines.
-
Self-Aware Safety Augmentation: Leveraging Internal Semantic Understanding to Enhance Safety in Vision-Language Models
A tuning-free method that projects middle-layer semantic representations back onto early safety layers, improving vision-language model safety with minimal utility loss.
-
How Visual Representations Map to Language Feature Space in Multimodal LLMs
Visual tokens in a fully frozen-backbone VLM with a linear adapter only become well-represented by the LLM's sparse autoencoder features in middle-to-late layers, converging around layer 18.
-
AudioLens: A Closer Look at Auditory Attribute Perception of Large Audio-Language Models
By projecting hidden states to the vocabulary at every layer, the paper shows that failed attribute recognition in three LALMs is marked by mid-network information peaks followed by degradation, and that models rely o...
-
Cross-modal Information Flow in Multimodal Large Language Models
In LLaVA multimodal models, visual information flows into question token representations in two stages, global then object-specific, before propagating to the final answer position.
-
Ask Twice, Look Twice: Prompt Echoing Resolves the Question-First Paradox in Vision-Language Models
Repeating the question on both sides of the image (question echoing) closes the question-first accuracy gap in five open VLMs and beats standard single-pass orderings on several VQA benchmarks.
-
Enhancing Multi-Robot Exploration Using Probabilistic Frontier Prioritization with Dirichlet Process Gaussian Mixtures
DP-GMM-based probabilistic frontier prioritization improves two multi-agent frontier explorers by roughly 10–14% across clutter, team size, and communication settings.
-
Short-LVLM: Compressing and Accelerating Large Vision-Language Models by Pruning Redundant Layers
A training-free layer pruning framework for large vision-language models, combining token importance scoring with subspace-compensated weight projection, preserves most accuracy while speeding inference.
-
Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models
A new patch-aligned pretraining loss improves fine-grained vision-language alignment and grounding in multimodal LLMs.
-
TDSal: Task-Based Top-Down Saliency Prediction Model
A modular network fuses truncated YOLOv5 features with Sentence-BERT task tokens via a one-layer transformer to output task-conditioned saliency maps on a four-task eye-tracking set.
-
Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey
A survey maps the field of MLLM explainability and interpretability into data, model, and training and inference perspectives.
Discussion (0). Continue with ORCID to comment.