A three-paradigm evaluation framework shows that reasoning over text descriptions of images (Componential Analysis) outperforms direct visual reasoning on Bongard and Winoground benchmarks, and that many open-source VLMs are limited more by perception than by reasoning.
Bongard-openworld: Few-shot reasoning for free-form visual concepts in the real world
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
dataset 1
citation-polarity summary
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1roles
dataset 1polarities
use dataset 1representative citing papers
citing papers explorer
-
A Cognitive Paradigm Approach to Probe the Perception-Reasoning Interface in VLMs
A three-paradigm evaluation framework shows that reasoning over text descriptions of images (Componential Analysis) outperforms direct visual reasoning on Bongard and Winoground benchmarks, and that many open-source VLMs are limited more by perception than by reasoning.