Pith. sign in

REVIEW 12 cited by

Reducing Hallucinations in Vision-Language Models via Latent Space Steering

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.15778 v2 pith:CK6ANY7M submitted 2024-10-21 cs.CV cs.AIcs.LGcs.MM

classification cs.CVcs.AIcs.LGcs.MM
keywords hallucinationslvlmsmodelshallucinationlargevisiondecodersinputs
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Hallucination poses a challenge to the deployment of large vision-language models (LVLMs) in applications. Unlike in large language models (LLMs), hallucination in LVLMs often arises from misalignments between visual inputs and textual outputs. This paper investigates the underlying mechanisms of hallucination, focusing on the unique structure of LVLMs that distinguishes them from large language models (LLMs). We identify that hallucinations often arise from the sensitivity of text decoders to vision inputs, a natural phenomenon when image encoders and text decoders are pre-trained separately. Inspired by this, we introduce Visual and Textual Intervention (VTI), a novel technique designed to reduce hallucinations by steering latent space representations during inference to enhance the stability of vision features. As a task-agnostic test-time intervention, VTI can be easily applied to any problem without additional cost. Extensive experiments demonstrate that it can effectively reduce hallucinations and outperform baseline methods across multiple metrics, highlighting the critical role of vision feature stability in LVLMs.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Beyond Global Editing: Per-Instance Disentangled Subspaces for Training-Free Hallucination Mitigation in LVLMs

    cs.CV 2026-08 conditional novelty 6.0 of 10

    Per-instance adaptive projection onto clustered hallucination subspaces reduces LVLM hallucination on CHAIR and POPE benchmarks without fine-tuning.

  2. Modality Bias in LVLMs: Analyzing and Mitigating Object Hallucination via Attention Lens

    cs.CV 2025-08 conditional novelty 6.0 of 10

    Modality bias, an imbalanced attention to text or image during hallucinated outputs, is shown to be mitigated by a training-free attention intervention plus contrastive decoding.

  3. GrAInS: Gradient-based Attribution for Inference-Time Steering of LLMs and VLMs

    cs.CL 2025-07 conditional novelty 6.0 of 10

    GrAInS uses Integrated Gradients to identify the most influential tokens, then builds layer-wise steering vectors that improve truthfulness, reduce hallucination, and preserve general capabilities in LLMs and VLMs.

  4. Not All Tokens and Heads Are Equally Important: Dual-Level Attention Intervention for Hallucination Mitigation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A dual-level attention intervention that boosts salient visual-token attention and suppresses text/system attention during decoding reduces hallucination rates in LLaVA, MiniGPT-4, and mPLUG-Owl2 on POPE and CHAIR.

  5. DriveSOTIF: Advancing Perception SOTIF Through Multimodal Large Language Models

    cs.RO 2025-05 conditional novelty 6.0 of 10

    Fine-tuning multimodal language models on a new SOTIF-focused driving dataset improves question answering and captioning, but the open-ended gains are measured by an LLM judge with no independent human scoring.

  6. The Hidden Life of Tokens: Reducing Hallucination of Large Vision-Language Models via Visual Information Steering

    cs.CV 2025-02 conditional novelty 6.0 of 10

    VISTA reduces hallucination in vision-language models by adding a per-image visual steering vector to hidden states and blending in early-layer logits, cutting CHAIR object hallucination by about 40%.

  7. The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook

    cs.AI 2026-04 accept novelty 5.0 of 10

    A large survey organizes latent-space work in language-based models by foundation, evolution, four mechanisms, seven abilities, and open challenges.

  8. CAI: Caption-Sensitive Attention Intervention for Mitigating Object Hallucination in Large Vision-Language Models

    cs.CV 2025-06 conditional novelty 5.0 of 10

    CAI reduces object hallucination in LVLMs by injecting caption-query attention patterns into selected attention heads at inference time.

  9. ReCo: Reminder Composition Mitigates Hallucinations in Vision-Language Models

    cs.CV 2025-06 conditional novelty 5.0 of 10

    ReCo, a lightweight DPO-trained linear head that re-injects pooled image embeddings at every step, reduces hallucination on five benchmarks across three VLMs and combines with existing mitigation methods.

  10. Can Generic LLMs Help Analyze Child-adult Interactions Involving Children with Autism in Clinical Observation?

    cs.CL 2024-11 conditional novelty 4.0 of 10

    Generic open-source LLMs can classify speakers, engaged activities, language skill levels, and age ranges in ASD child-adult clinical transcripts, and sometimes outperform non-expert human raters.

  11. Generative AI Act II: Test Time Scaling Drives Cognition Engineering

    cs.CL 2025-04 conditional novelty 3.0 of 10

    Test-time scaling techniques such as long chain-of-thought, tree search, and self-correction define the paper's 'cognition engineering' paradigm, which it surveys, taxonomizes, and tutorials.

  12. Empowering Multimodal LLMs with External Tools: A Comprehensive Survey

    cs.CV 2025-08 unverdicted novelty 2.0 of 10

    A survey paper maps how external tools are used to augment multimodal large language models across data, tasks, evaluation, and future directions.

Pith tools