ViRGo adaptively routes visual retrieval decisions in VLMs by estimating object scale from intrinsic localization heads combined with token confidence, matching patch retrieval on small objects, attention retrieval on large ones, and global processing when unnecessary.
Exploring perceptual limitation of multimodal large language models
2 Pith papers cite this work. Polarity classification is still indexing.
fields
cs.CV 2years
2026 2verdicts
UNVERDICTED 2representative citing papers
RISE proposes a self-evolving VLM framework with three designs to address challenges in question generation and solver adaptation, reporting consistent gains on seven benchmarks across two backbones.
citing papers explorer
-
Look Before You Zoom: Adaptive Routing for the Resolution-Context Trade-off in Visual RAG
ViRGo adaptively routes visual retrieval decisions in VLMs by estimating object scale from intrinsic localization heads combined with token confidence, matching patch retrieval on small objects, attention retrieval on large ones, and global processing when unnecessary.
-
RISE: Reliable Improvement in Self-Evolving Vision-Language Models
RISE proposes a self-evolving VLM framework with three designs to address challenges in question generation and solver adaptation, reporting consistent gains on seven benchmarks across two backbones.