A new verifiable benchmark shows VLMs struggle to recover complete decision problems from images, with large sensitivity to presentation and only modest gains from agent scaffolds.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making
A new verifiable benchmark shows VLMs struggle to recover complete decision problems from images, with large sensitivity to presentation and only modest gains from agent scaffolds.