Pith. sign in

REVIEW 13 cited by

Don't Miss the Forest for the Trees: Attentional Vision Calibration for Large Vision Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.17820 v2 pith:FXNAGOVH submitted 2024-05-28 cs.CV cs.AI

classification cs.CVcs.AI
keywords visionattentionaviscblindlvlmstokensattentionalcalibration
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Large Vision Language Models (LVLMs) demonstrate strong capabilities in visual understanding and description, yet often suffer from hallucinations, attributing incorrect or misleading features to images. We observe that LVLMs disproportionately focus on a small subset of image tokens--termed blind tokens--which are typically irrelevant to the query (e.g., background or non-object regions). We hypothesize that such attention misalignment plays a key role in generating hallucinated responses. To mitigate this issue, we propose Attentional Vision Calibration (AvisC), a test-time approach that dynamically recalibrates the influence of blind tokens without modifying the underlying attention mechanism. AvisC first identifies blind tokens by analyzing layer-wise attention distributions over image tokens, then employs a contrastive decoding strategy to balance the influence of original and blind-token-biased logits. Experiments on standard benchmarks, including POPE, MME, and AMBER, demonstrate that AvisC effectively reduces hallucinations in LVLMs.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. VADER: Adaptive Debiasing for Hallucination Mitigation in Video Large Language Models

    cs.CV 2026-08 conditional novelty 7.0 of 10

    A training-free decoding framework that adaptively reweights attention toward video tokens and erases key visual evidence per frame to suppress hallucinated predictions, achieving 72.60% accuracy on EventHallusion wit...

  2. Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models

    cs.CV 2025-02 conditional novelty 7.0 of 10

    SPARC selectively and progressively reinforces attention to relevant image tokens during decoding, improving both precision and recall in detailed image captioning compared to baselines and prior hallucination-mitigat...

  3. LookBack: Where and How to Score LVLM Responses via Visual Reference Usage

    cs.CV 2026-08 conditional novelty 6.0 of 10

    LookBack scores LVLM responses by calibrating token likelihood with an attention-based visual lookback score and weighting by visual relevance, improving Best-of-N selection over baselines.

  4. C-PTQ: Fisher-weighted Channel-wise Sensitivity for Post-training Quantization of MLLMs

    cs.CV 2026-07 conditional novelty 6.0 of 10

    C-PTQ weights quantization error by per-channel Fisher information of the task loss, improving low-bit accuracy of multimodal LLMs by small margins over existing channel-wise scaling methods.

  5. Focus on What Matters: Enhancing Medical Vision-Language Models with Automatic Attention Alignment Tuning

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A3Tune aligns the visual attention of medical LVLMs to prompt-relevant regions via SAM and BioMedCLIP weak labels plus a Mixture-of-Experts over LoRA, improving VQA and report generation accuracy.

  6. The Hidden Life of Tokens: Reducing Hallucination of Large Vision-Language Models via Visual Information Steering

    cs.CV 2025-02 conditional novelty 6.0 of 10

    VISTA reduces hallucination in vision-language models by adding a per-image visual steering vector to hidden states and blending in early-layer logits, cutting CHAIR object hallucination by about 40%.

  7. Cross-Modal Attention Calibration for LVLM Hallucination Mitigation

    cs.CV 2025-01 conditional novelty 6.0 of 10

    IMCCD combines value-vector masking in cross-modal attention with a position-normalizing decoding step and reports lower hallucination than VCD and ICD on POPE, CHAIR, and MME.

  8. SeeMe: Mitigating Hallucinations in Large Vision-Language Models through Effective Visual Token Engineering

    cs.CV 2026-07 conditional novelty 5.5 of 10

    Restructuring visual tokens via cross-modal prune–merge–refine consistently lowers hallucination rates on MME, POPE and AMBER across four 7B LVLMs without any training.

  9. ReCo: Reminder Composition Mitigates Hallucinations in Vision-Language Models

    cs.CV 2025-06 conditional novelty 5.0 of 10

    ReCo, a lightweight DPO-trained linear head that re-injects pooled image embeddings at every step, reduces hallucination on five benchmarks across three VLMs and combines with existing mitigation methods.

  10. ECG-Byte: A Tokenizer for End-to-End Generative Electrocardiogram Language Modeling

    cs.CL 2024-12 conditional novelty 5.0 of 10

    A BPE-based tokenizer lets an LLM generate clinical text directly from quantized ECG signals, matching two-stage encoder methods with roughly 3x faster training and 48% of the data.

  11. Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning

    cs.CV 2024-11 conditional novelty 5.0 of 10

    A DPO-trained VLM critic that critiques and iteratively refines a reasoning VLM improves accuracy on several multimodal benchmarks, with large gains on MathVista and RealWorldQA.

  12. Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features

    cs.CV 2025-09 conditional novelty 4.0 of 10

    Prompt specificity measurably affects counting accuracy and attention allocation in Qwen2.5-VL and Kimi-VL, and can partially overcome learned visual priors.

  13. Mixture of Decoding: An Attention-Inspired Adaptive Decoding Strategy to Mitigate Hallucinations in Large Vision-Language Models

    cs.CL 2025-05 conditional novelty 4.0 of 10

    MoD reduces hallucinations in large vision-language models by measuring the Jensen-Shannon divergence between outputs from full and attention-masked image tokens and switching between complementary and contrastive decoding.

Pith tools