REVIEW 13 cited by
Don't Miss the Forest for the Trees: Attentional Vision Calibration for Large Vision Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Large Vision Language Models (LVLMs) demonstrate strong capabilities in visual understanding and description, yet often suffer from hallucinations, attributing incorrect or misleading features to images. We observe that LVLMs disproportionately focus on a small subset of image tokens--termed blind tokens--which are typically irrelevant to the query (e.g., background or non-object regions). We hypothesize that such attention misalignment plays a key role in generating hallucinated responses. To mitigate this issue, we propose Attentional Vision Calibration (AvisC), a test-time approach that dynamically recalibrates the influence of blind tokens without modifying the underlying attention mechanism. AvisC first identifies blind tokens by analyzing layer-wise attention distributions over image tokens, then employs a contrastive decoding strategy to balance the influence of original and blind-token-biased logits. Experiments on standard benchmarks, including POPE, MME, and AMBER, demonstrate that AvisC effectively reduces hallucinations in LVLMs.
Forward citations
Cited by 13 Pith papers
-
VADER: Adaptive Debiasing for Hallucination Mitigation in Video Large Language Models
A training-free decoding framework that adaptively reweights attention toward video tokens and erases key visual evidence per frame to suppress hallucinated predictions, achieving 72.60% accuracy on EventHallusion wit...
-
Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models
SPARC selectively and progressively reinforces attention to relevant image tokens during decoding, improving both precision and recall in detailed image captioning compared to baselines and prior hallucination-mitigat...
-
LookBack: Where and How to Score LVLM Responses via Visual Reference Usage
LookBack scores LVLM responses by calibrating token likelihood with an attention-based visual lookback score and weighting by visual relevance, improving Best-of-N selection over baselines.
-
C-PTQ: Fisher-weighted Channel-wise Sensitivity for Post-training Quantization of MLLMs
C-PTQ weights quantization error by per-channel Fisher information of the task loss, improving low-bit accuracy of multimodal LLMs by small margins over existing channel-wise scaling methods.
-
Focus on What Matters: Enhancing Medical Vision-Language Models with Automatic Attention Alignment Tuning
A3Tune aligns the visual attention of medical LVLMs to prompt-relevant regions via SAM and BioMedCLIP weak labels plus a Mixture-of-Experts over LoRA, improving VQA and report generation accuracy.
-
The Hidden Life of Tokens: Reducing Hallucination of Large Vision-Language Models via Visual Information Steering
VISTA reduces hallucination in vision-language models by adding a per-image visual steering vector to hidden states and blending in early-layer logits, cutting CHAIR object hallucination by about 40%.
-
Cross-Modal Attention Calibration for LVLM Hallucination Mitigation
IMCCD combines value-vector masking in cross-modal attention with a position-normalizing decoding step and reports lower hallucination than VCD and ICD on POPE, CHAIR, and MME.
-
SeeMe: Mitigating Hallucinations in Large Vision-Language Models through Effective Visual Token Engineering
Restructuring visual tokens via cross-modal prune–merge–refine consistently lowers hallucination rates on MME, POPE and AMBER across four 7B LVLMs without any training.
-
ReCo: Reminder Composition Mitigates Hallucinations in Vision-Language Models
ReCo, a lightweight DPO-trained linear head that re-injects pooled image embeddings at every step, reduces hallucination on five benchmarks across three VLMs and combines with existing mitigation methods.
-
ECG-Byte: A Tokenizer for End-to-End Generative Electrocardiogram Language Modeling
A BPE-based tokenizer lets an LLM generate clinical text directly from quantized ECG signals, matching two-stage encoder methods with roughly 3x faster training and 48% of the data.
-
Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning
A DPO-trained VLM critic that critiques and iteratively refines a reasoning VLM improves accuracy on several multimodal benchmarks, with large gains on MathVista and RealWorldQA.
-
Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features
Prompt specificity measurably affects counting accuracy and attention allocation in Qwen2.5-VL and Kimi-VL, and can partially overcome learned visual priors.
-
Mixture of Decoding: An Attention-Inspired Adaptive Decoding Strategy to Mitigate Hallucinations in Large Vision-Language Models
MoD reduces hallucinations in large vision-language models by measuring the Jensen-Shannon divergence between outputs from full and attention-masked image tokens and switching between complementary and contrastive decoding.
Discussion (0). Continue with ORCID to comment.