VLMs score near random on new causal-order benchmarks (VQA-Causal, VCR-Causal) despite strong object and activity recognition, and targeted hard-negative fine-tuning only partially closes the gap.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
What's Missing in Vision-Language Models? Probing Their Struggles with Causal Order Reasoning
VLMs score near random on new causal-order benchmarks (VQA-Causal, VCR-Causal) despite strong object and activity recognition, and targeted hard-negative fine-tuning only partially closes the gap.