REVIEW 16 cited by
Interpreting and Editing Vision-Language Representations to Mitigate Hallucinations
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We investigate the internal representations of vision-language models (VLMs) to address hallucinations, a persistent challenge despite advances in model size and training. We project VLMs' internal image representations to their language vocabulary and observe more confident output probabilities on real objects than hallucinated objects. We additionally use these output probabilities to spatially localize real objects. Building on this approach, we introduce a knowledge erasure algorithm that removes hallucinations by linearly orthogonalizing image features with respect to hallucinated object features. We show that targeted edits to a model's latent representations can reduce hallucinations by up to 25.7% on the COCO2014 dataset while preserving performance. Our findings demonstrate how a deeper understanding of VLMs' latent representations can enhance reliability and enable novel capabilities, such as zero-shot segmentation.
Forward citations
Cited by 16 Pith papers
-
Multimodal Model Diffing for Feature Discovery and Control
By diffing base-language and multimodal sparse autoencoder features, MMDiff isolates causally relevant features that can be ablated or steered to control spatial, OCR, and safety behaviors in multimodal LLMs.
-
Verbalizable Representations Form a Global Workspace in Language Models
Language models represent their current reasoning in a small, readable set of verbalizable vectors (the J-space) that functions like a global workspace.
-
Image Tokens Matter: Mitigating Hallucination in Discrete Tokenizer-based Large Vision-Language Models via Latent Editing
CGC+VTD identifies co-occurring image token clusters as a source of hallucinated objects in discrete-token LVLMs and suppresses clusters' absent-token signals in latent space, cutting hallucination rates across Chamel...
-
What's in the Image? A Deep-Dive into the Vision of Vision Language Models
Vision-language models store a global image summary in the query text tokens, rely on the middle transformer layers for vision-to-text transfer, and fetch fine details from image tokens in a spatially localized way.
-
TruthLens: Object Hallucination Detection via Self-Evaluating Truthfulness Scores in LVLMs
TruthLens fine-tunes LVLMs so the log-probability of a special token at each object mention becomes a truthfulness score, detecting object hallucinations with state-of-the-art AUROC.
-
Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders
A lightweight Q-Former proxy trained on VLM hidden states reveals that localization signals peak in input-dependent intermediate layers, not the final layers used by standard editing pipelines.
-
SAEmnesia: Erasing Concepts in Diffusion Models with Supervised Sparse Autoencoders
A supervised sparse autoencoder binds each concept to a single neuron, letting Stable Diffusion erase a concept by steering one latent.
-
How Visual Representations Map to Language Feature Space in Multimodal LLMs
Visual tokens in a fully frozen-backbone VLM with a linear adapter only become well-represented by the LLM's sparse autoencoder features in middle-to-late layers, converging around layer 18.
-
Debiasing CLIP: Interpreting and Correcting Bias in Attention Heads
Using wrong/correct hard-sample head comparisons, LTC finds spurious CLIP attention heads and corrects them to raise worst-group accuracy on biased benchmarks.
-
Devils in Middle Layers of Large Vision-Language Models: Interpreting, Detecting and Mitigating Object Hallucinations via Attention Lens
The middle layers of LVLMs process visual information in two stages, and amplifying image attention in the first 'enrichment' stage reduces object hallucinations.
-
Same Attention, Different Truths: Put Logit-Lens over Visual Attention to Detect and Mitigate LVLM Object Hallucination
Object hallucination in LVLMs is not explained by attention strength but by a mismatch between attended visual semantics and generated tokens, which can be detected via Logit Lens and mitigated by targeted masking or ...
-
Evidence Recomposition and Predictive Context Residualization for Visual Attribution in Multimodal Large Language Models
MSEA+ARC, a multi-scale and ranking-based residualization method, claims consistent F1-IoU gains over TAM for token-level MLLM visual attribution.
-
Steering Conceptual Bias via Transformer Latent-Subspace Activation
G-ACT improves per-layer probes for steering LLMs toward C++ code generation, yet the paper's main evidence is probe accuracy rather than actual output-language statistics.
-
Mitigating Hallucinations via Inter-Layer Consistency Aggregation in Large Vision-Language Models
DCLA reduces hallucinations by correcting each transformer layer's hidden state toward an exponentially weighted average of earlier layers, gated by a cosine-similarity threshold.
-
Beyond Logit Lens: Contextual Embeddings for Robust Hallucination Detection & Grounding in VLMs
Middle-layer contextual embeddings, not logit-lens readings, improve hallucination detection in VLMs and enable bounding-box grounding for visual question answering.
-
Test-Time Hallucination Control in Large Vision-Language Models
At each decoding step, TTH blends the LVLM's logits for candidate object tokens with CLIP image-text similarity scores, weighted by the model's uncertainty, to suppress hallucinated objects.
Discussion (0). Continue with ORCID to comment.