A new video benchmark asks models to reorder shuffled clips from everyday tasks; most large video language models perform near chance, and a recognize-then-reason prompt decomposition gives large gains.
Spatialvlm: Endowing vision-language models with spatial reasoning capabilities
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models
A new video benchmark asks models to reorder shuffled clips from everyday tasks; most large video language models perform near chance, and a recognize-then-reason prompt decomposition gives large gains.