Pith. sign in

REVIEW 22 cited by

MLLMs Know Where to Look: Training-free Perception of Small Visual Details with Multimodal LLMs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.17422 v1 pith:QAVTUYJH submitted 2025-02-24 cs.CV cs.AIcs.CL

classification cs.CVcs.AIcs.CL
keywords visualmllmsdetailssmallansweringinterventionperceptionthey
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Multimodal Large Language Models (MLLMs) have experienced rapid progress in visual recognition tasks in recent years. Given their potential integration into many critical applications, it is important to understand the limitations of their visual perception. In this work, we study whether MLLMs can perceive small visual details as effectively as large ones when answering questions about images. We observe that their performance is very sensitive to the size of the visual subject of the question, and further show that this effect is in fact causal by conducting an intervention study. Next, we study the attention patterns of MLLMs when answering visual questions, and intriguingly find that they consistently know where to look, even when they provide the wrong answer. Based on these findings, we then propose training-free visual intervention methods that leverage the internal knowledge of any MLLM itself, in the form of attention and gradient maps, to enhance its perception of small visual details. We evaluate our proposed methods on two widely-used MLLMs and seven visual question answering benchmarks and show that they can significantly improve MLLMs' accuracy without requiring any training. Our results elucidate the risk of applying MLLMs to visual recognition tasks concerning small details and indicate that visual intervention using the model's internal state is a promising direction to mitigate this risk.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 22 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMs

    cs.CV 2025-11 unverdicted novelty 8.0 of 10

    MVI-Bench supplies the first taxonomy and dataset focused on misleading visual inputs to measure LVLM robustness, with tests on 18 models revealing clear weaknesses.

  2. HAFI-VLM: A Frequency Perspective for Diagnosing and Enhancing Visual Perception in Vision-Language Models

    cs.CV 2026-08 conditional novelty 7.0 of 10

    Pretrained vision encoders show spectral response rigidity, and HAFI-VLM injects text-conditioned low/mid/high frequency evidence to improve VLM perception on VQA, text-rich understanding, and hallucination robustness.

  3. An Exam for Active Observers

    cs.CV 2026-07 conditional novelty 7.0 of 10

    On a new 17-task benchmark of active visual observation, the best frontier multimodal model solves 10.6% of items and humans solve 96.1%.

  4. eSkinHealth: A Multimodal Dataset for Neglected Tropical Skin Diseases

    cs.AI 2025-08 conditional novelty 7.0 of 10

    eSkinHealth is a new West African skin-disease dataset with 5,623 images, 47 conditions, and multimodal annotations including masks, captions, and clinical concepts for AI dermatology research.

  5. Look Where It Matters: Adaptive Visual Refinement for Vision-Language-Action Models

    cs.RO 2026-08 conditional novelty 6.0 of 10

    Inserting register tokens into a VLA encoder plus uncertainty-gated, attention-guided cropping raises π0's success from 94.2% to 98.4% on LIBERO and 46.5% to 69.0% on a real-world benchmark, at 1.4–1.6× compute.

  6. Remember-R1: Mitigating Long-Context Visual Forgetting through Reinforcement Learning

    cs.CV 2026-08 conditional novelty 6.0 of 10

    Remember-R1 adds three process-level rewards to GRPO, encouraging keyword coverage, sustained visual attention, and focus on question-relevant regions, improving MLLM reasoning across seven benchmarks.

  7. Geo3R: Mitigating Spatial Reasoning Hallucination in Multimodal Large Language Models

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Augmenting MLLMs with structured 3D geometric cards from monocular depth, camera calibration, and object orientation reduces spatial reasoning errors on 18 tasks, with gains up to 10.9 points over unaugmented baselines.

  8. Scene Graph Thinking: Reinforcing Structured Visual Reasoning for Multimodal Large Language Models

    cs.CV 2026-07 accept novelty 6.0 of 10

    Scene Graph Thinking (SaGe) converts images into hierarchical graphs, samples 120K node-articulated CoTs, and uses node-as-proxy GRPO rewards to improve MLLM fine-grained and relational visual reasoning.

  9. Beyond Scene Priors: Fine-Grained Traffic Scene Reasoning with Benchmarking and Query-Guided Small-Object Focus

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A query-guided sparse residual module (TG-SOF) plus a new distractor-heavy traffic MCQ benchmark lifts a 4B MLLM by 2.1 points on fine-grained local-evidence questions.

  10. Token-Based Affordance Grounding with Large Vision-Language Models

    cs.CV 2026-07 accept novelty 6.0 of 10

    TokAG selects the LVLM output token whose aggregated cross-attention is most concentrated on a CLIPSeg object mask, converting that map into a zero-shot affordance heatmap that outperforms weakly supervised baselines.

  11. Fully Spiking Neural Networks with Target Awareness for Energy-Efficient UAV Tracking

    cs.CV 2026-03 conditional novelty 6.0 of 10

    A two-stage pure RL method with an information-gap global view and hierarchical grounding loss makes MLLMs truly rely on precise crops and sets SOTA on high-res VQA under tight token budgets.

  12. LookWise: Knowing When and Where to Look for Fine-Grained Visual Reasoning in Multimodal Large Language Models

    cs.CV 2026-02 conditional novelty 6.0 of 10

    A confidence-based when-to-look gate plus semantic-guided attention where-to-look module improves training-free fine-grained visual reasoning accuracy and reduces inference cost versus search-based baselines.

  13. Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding

    cs.CV 2025-12 conditional novelty 6.0 of 10

    Blink dynamically expands high-saliency visual tokens and drops them when attention shifts, improving LLaVA-1.5 and LLaVA-NeXT across seven multimodal benchmarks.

  14. Capturing Gaze Shifts for Guidance: Cross-Modal Fusion Enhancement for VLM Hallucination Mitigation

    cs.CV 2025-10 conditional novelty 6.0 of 10

    Tracking positive shifts in visual attention over information-rich query words yields a saliency map that, when used to boost visual and query attention during decoding, reduces object hallucination on CHAIR, POPE, an...

  15. CARES: Context-Aware Resolution Selector for VLMs

    cs.CV 2025-10 conditional novelty 6.0 of 10

    A lightweight selector trained on target-VLM rollouts routes each image-query pair to the minimal sufficient resolution, cutting prefill compute by 63-80% across five benchmarks at roughly matched accuracy.

  16. SheetDesigner: MLLM-Powered Spreadsheet Layout Generation with Rule-Based and Vision-Based Reflection

    cs.AI 2025-09 conditional novelty 6.0 of 10

    SheetDesigner uses zero-shot multimodal LLMs with rule- and vision-based reflection to generate spreadsheet layouts, and claims a 22.6% gain over baselines on a new seven-criterion benchmark.

  17. ZoomEye: Enhancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image Exploration

    cs.CV 2024-11 conditional novelty 6.0 of 10

    ZoomEye uses tree-based zooming, guided by an MLLM's own confidence scores, to improve high-resolution visual question answering without retraining the model.

  18. Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning

    cs.CV 2025-09 conditional novelty 5.0 of 10

    CARVE contrasts attention maps from a general prompt and a specific question to mask out visual noise, then re-asks the question on the cropped and enlarged image, improving VQA accuracy by up to 75% on some benchmarks.

  19. Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning

    cs.CV 2025-06 reject novelty 5.0 of 10

    A new micro-edit dataset and fine-tuning recipe appear to help multimodal LLMs notice small visual changes, but the central 'feature consistency loss' claim is not present in the method.

  20. Visually Interpretable Subtask Reasoning for Visual Question Answering

    cs.CV 2025-05 conditional novelty 5.0 of 10

    Training a visual question answering model on LLM-generated subtask rationales with bounding boxes improves GQA accuracy from 64.0 to 65.1 percent and adds grounded explanations.

  21. Demystifying the Visual Quality Paradox in Multimodal Large Language Models

    cs.CV 2025-06 reject novelty 4.0 of 10

    Multimodal LLM accuracy can improve on visually degraded images, and a lightweight test-time tuning module that modulates input quality yields small accuracy gains on some benchmarks.

  22. Enhancing Adversarial Transferability via Component-Wise Transformation

    cs.CV 2025-01 conditional novelty 4.0 of 10

    A block-wise interpolation and selective rotation attack, CWT, improves adversarial transferability across CNN and transformer models on ImageNet.

Pith tools