CAVIA, a training-free loop where an LLM directs a VLM to inspect specific video frames and repeats until confident, reports accuracy gains on EgoSchema, NExT-QA, and IntentQA, but those gains are not shown to come from the loop itself.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
See What You Need: Query-Aware Visual Intelligence through Reasoning-Perception Loops
CAVIA, a training-free loop where an LLM directs a VLM to inspect specific video frames and repeats until confident, reports accuracy gains on EgoSchema, NExT-QA, and IntentQA, but those gains are not shown to come from the loop itself.