Pith. sign in

REVIEW 16 cited by

Interpreting and Editing Vision-Language Representations to Mitigate Hallucinations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.02762 v2 pith:OHACGJCN submitted 2024-10-03 cs.CV cs.LG

classification cs.CVcs.LG
keywords representationshallucinationsobjectsvlmsfeatureshallucinatedimageinternal
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We investigate the internal representations of vision-language models (VLMs) to address hallucinations, a persistent challenge despite advances in model size and training. We project VLMs' internal image representations to their language vocabulary and observe more confident output probabilities on real objects than hallucinated objects. We additionally use these output probabilities to spatially localize real objects. Building on this approach, we introduce a knowledge erasure algorithm that removes hallucinations by linearly orthogonalizing image features with respect to hallucinated object features. We show that targeted edits to a model's latent representations can reduce hallucinations by up to 25.7% on the COCO2014 dataset while preserving performance. Our findings demonstrate how a deeper understanding of VLMs' latent representations can enhance reliability and enable novel capabilities, such as zero-shot segmentation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 16 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Multimodal Model Diffing for Feature Discovery and Control

    cs.CV 2026-08 conditional novelty 7.0 of 10

    By diffing base-language and multimodal sparse autoencoder features, MMDiff isolates causally relevant features that can be ablated or steered to control spatial, OCR, and safety behaviors in multimodal LLMs.

  2. Verbalizable Representations Form a Global Workspace in Language Models

    cs.CL 2026-07 conditional novelty 7.0 of 10

    Language models represent their current reasoning in a small, readable set of verbalizable vectors (the J-space) that functions like a global workspace.

  3. Image Tokens Matter: Mitigating Hallucination in Discrete Tokenizer-based Large Vision-Language Models via Latent Editing

    cs.CV 2025-05 conditional novelty 7.0 of 10

    CGC+VTD identifies co-occurring image token clusters as a source of hallucinated objects in discrete-token LVLMs and suppresses clusters' absent-token signals in latent space, cutting hallucination rates across Chamel...

  4. What's in the Image? A Deep-Dive into the Vision of Vision Language Models

    cs.CV 2024-11 conditional novelty 7.0 of 10

    Vision-language models store a global image summary in the query text tokens, rely on the middle transformer layers for vision-to-text transfer, and fetch fine details from image tokens in a spatially localized way.

  5. TruthLens: Object Hallucination Detection via Self-Evaluating Truthfulness Scores in LVLMs

    cs.CV 2026-08 conditional novelty 6.0 of 10

    TruthLens fine-tunes LVLMs so the log-probability of a special token at each object mention becomes a truthfulness score, detecting object hallucinations with state-of-the-art AUROC.

  6. Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A lightweight Q-Former proxy trained on VLM hidden states reveals that localization signals peak in input-dependent intermediate layers, not the final layers used by standard editing pipelines.

  7. SAEmnesia: Erasing Concepts in Diffusion Models with Supervised Sparse Autoencoders

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A supervised sparse autoencoder binds each concept to a single neuron, letting Stable Diffusion erase a concept by steering one latent.

  8. How Visual Representations Map to Language Feature Space in Multimodal LLMs

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Visual tokens in a fully frozen-backbone VLM with a linear adapter only become well-represented by the LLM's sparse autoencoder features in middle-to-late layers, converging around layer 18.

  9. Debiasing CLIP: Interpreting and Correcting Bias in Attention Heads

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Using wrong/correct hard-sample head comparisons, LTC finds spurious CLIP attention heads and corrects them to raise worst-group accuracy on biased benchmarks.

  10. Devils in Middle Layers of Large Vision-Language Models: Interpreting, Detecting and Mitigating Object Hallucinations via Attention Lens

    cs.CV 2024-11 conditional novelty 6.0 of 10

    The middle layers of LVLMs process visual information in two stages, and amplifying image attention in the first 'enrichment' stage reduces object hallucinations.

  11. Same Attention, Different Truths: Put Logit-Lens over Visual Attention to Detect and Mitigate LVLM Object Hallucination

    cs.CV 2026-08 conditional novelty 5.0 of 10

    Object hallucination in LVLMs is not explained by attention strength but by a mismatch between attended visual semantics and generated tokens, which can be detected via Logit Lens and mitigated by targeted masking or ...

  12. Evidence Recomposition and Predictive Context Residualization for Visual Attribution in Multimodal Large Language Models

    cs.CV 2025-09 conditional novelty 5.0 of 10

    MSEA+ARC, a multi-scale and ranking-based residualization method, claims consistent F1-IoU gains over TAM for token-level MLLM visual attribution.

  13. Steering Conceptual Bias via Transformer Latent-Subspace Activation

    cs.AI 2025-06 reject novelty 5.0 of 10

    G-ACT improves per-layer probes for steering LLMs toward C++ code generation, yet the paper's main evidence is probe accuracy rather than actual output-language statistics.

  14. Mitigating Hallucinations via Inter-Layer Consistency Aggregation in Large Vision-Language Models

    cs.LG 2025-05 conditional novelty 5.0 of 10

    DCLA reduces hallucinations by correcting each transformer layer's hidden state toward an exponentially weighted average of earlier layers, gated by a cosine-similarity threshold.

  15. Beyond Logit Lens: Contextual Embeddings for Robust Hallucination Detection & Grounding in VLMs

    cs.CL 2024-11 conditional novelty 5.0 of 10

    Middle-layer contextual embeddings, not logit-lens readings, improve hallucination detection in VLMs and enable bounding-box grounding for visual question answering.

  16. Test-Time Hallucination Control in Large Vision-Language Models

    cs.CV 2026-08 conditional novelty 3.0 of 10

    At each decoding step, TTH blends the LVLM's logits for candidate object tokens with CLIP image-text similarity scores, weighted by the model's uncertainty, to suppress hallucinated objects.

Pith tools