REVIEW 22 cited by
MLLMs Know Where to Look: Training-free Perception of Small Visual Details with Multimodal LLMs
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Multimodal Large Language Models (MLLMs) have experienced rapid progress in visual recognition tasks in recent years. Given their potential integration into many critical applications, it is important to understand the limitations of their visual perception. In this work, we study whether MLLMs can perceive small visual details as effectively as large ones when answering questions about images. We observe that their performance is very sensitive to the size of the visual subject of the question, and further show that this effect is in fact causal by conducting an intervention study. Next, we study the attention patterns of MLLMs when answering visual questions, and intriguingly find that they consistently know where to look, even when they provide the wrong answer. Based on these findings, we then propose training-free visual intervention methods that leverage the internal knowledge of any MLLM itself, in the form of attention and gradient maps, to enhance its perception of small visual details. We evaluate our proposed methods on two widely-used MLLMs and seven visual question answering benchmarks and show that they can significantly improve MLLMs' accuracy without requiring any training. Our results elucidate the risk of applying MLLMs to visual recognition tasks concerning small details and indicate that visual intervention using the model's internal state is a promising direction to mitigate this risk.
Forward citations
Cited by 22 Pith papers
-
MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMs
MVI-Bench supplies the first taxonomy and dataset focused on misleading visual inputs to measure LVLM robustness, with tests on 18 models revealing clear weaknesses.
-
HAFI-VLM: A Frequency Perspective for Diagnosing and Enhancing Visual Perception in Vision-Language Models
Pretrained vision encoders show spectral response rigidity, and HAFI-VLM injects text-conditioned low/mid/high frequency evidence to improve VLM perception on VQA, text-rich understanding, and hallucination robustness.
-
An Exam for Active Observers
On a new 17-task benchmark of active visual observation, the best frontier multimodal model solves 10.6% of items and humans solve 96.1%.
-
eSkinHealth: A Multimodal Dataset for Neglected Tropical Skin Diseases
eSkinHealth is a new West African skin-disease dataset with 5,623 images, 47 conditions, and multimodal annotations including masks, captions, and clinical concepts for AI dermatology research.
-
Look Where It Matters: Adaptive Visual Refinement for Vision-Language-Action Models
Inserting register tokens into a VLA encoder plus uncertainty-gated, attention-guided cropping raises π0's success from 94.2% to 98.4% on LIBERO and 46.5% to 69.0% on a real-world benchmark, at 1.4–1.6× compute.
-
Remember-R1: Mitigating Long-Context Visual Forgetting through Reinforcement Learning
Remember-R1 adds three process-level rewards to GRPO, encouraging keyword coverage, sustained visual attention, and focus on question-relevant regions, improving MLLM reasoning across seven benchmarks.
-
Geo3R: Mitigating Spatial Reasoning Hallucination in Multimodal Large Language Models
Augmenting MLLMs with structured 3D geometric cards from monocular depth, camera calibration, and object orientation reduces spatial reasoning errors on 18 tasks, with gains up to 10.9 points over unaugmented baselines.
-
Scene Graph Thinking: Reinforcing Structured Visual Reasoning for Multimodal Large Language Models
Scene Graph Thinking (SaGe) converts images into hierarchical graphs, samples 120K node-articulated CoTs, and uses node-as-proxy GRPO rewards to improve MLLM fine-grained and relational visual reasoning.
-
Beyond Scene Priors: Fine-Grained Traffic Scene Reasoning with Benchmarking and Query-Guided Small-Object Focus
A query-guided sparse residual module (TG-SOF) plus a new distractor-heavy traffic MCQ benchmark lifts a 4B MLLM by 2.1 points on fine-grained local-evidence questions.
-
Token-Based Affordance Grounding with Large Vision-Language Models
TokAG selects the LVLM output token whose aggregated cross-attention is most concentrated on a CLIPSeg object mask, converting that map into a zero-shot affordance heatmap that outperforms weakly supervised baselines.
-
Fully Spiking Neural Networks with Target Awareness for Energy-Efficient UAV Tracking
A two-stage pure RL method with an information-gap global view and hierarchical grounding loss makes MLLMs truly rely on precise crops and sets SOTA on high-res VQA under tight token budgets.
-
LookWise: Knowing When and Where to Look for Fine-Grained Visual Reasoning in Multimodal Large Language Models
A confidence-based when-to-look gate plus semantic-guided attention where-to-look module improves training-free fine-grained visual reasoning accuracy and reduces inference cost versus search-based baselines.
-
Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding
Blink dynamically expands high-saliency visual tokens and drops them when attention shifts, improving LLaVA-1.5 and LLaVA-NeXT across seven multimodal benchmarks.
-
Capturing Gaze Shifts for Guidance: Cross-Modal Fusion Enhancement for VLM Hallucination Mitigation
Tracking positive shifts in visual attention over information-rich query words yields a saliency map that, when used to boost visual and query attention during decoding, reduces object hallucination on CHAIR, POPE, an...
-
CARES: Context-Aware Resolution Selector for VLMs
A lightweight selector trained on target-VLM rollouts routes each image-query pair to the minimal sufficient resolution, cutting prefill compute by 63-80% across five benchmarks at roughly matched accuracy.
-
SheetDesigner: MLLM-Powered Spreadsheet Layout Generation with Rule-Based and Vision-Based Reflection
SheetDesigner uses zero-shot multimodal LLMs with rule- and vision-based reflection to generate spreadsheet layouts, and claims a 22.6% gain over baselines on a new seven-criterion benchmark.
-
ZoomEye: Enhancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image Exploration
ZoomEye uses tree-based zooming, guided by an MLLM's own confidence scores, to improve high-resolution visual question answering without retraining the model.
-
Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning
CARVE contrasts attention maps from a general prompt and a specific question to mask out visual noise, then re-asks the question on the cropped and enlarged image, improving VQA accuracy by up to 75% on some benchmarks.
-
Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning
A new micro-edit dataset and fine-tuning recipe appear to help multimodal LLMs notice small visual changes, but the central 'feature consistency loss' claim is not present in the method.
-
Visually Interpretable Subtask Reasoning for Visual Question Answering
Training a visual question answering model on LLM-generated subtask rationales with bounding boxes improves GQA accuracy from 64.0 to 65.1 percent and adds grounded explanations.
-
Demystifying the Visual Quality Paradox in Multimodal Large Language Models
Multimodal LLM accuracy can improve on visually degraded images, and a lightweight test-time tuning module that modulates input quality yields small accuracy gains on some benchmarks.
-
Enhancing Adversarial Transferability via Component-Wise Transformation
A block-wise interpolation and selective rotation attack, CWT, improves adversarial transferability across CNN and transformer models on ImageNet.
Discussion (0). Continue with ORCID to comment.