A new 1,000-video benchmark with adversarial question variants shows video LLMs remain far below human accuracy and robustness on short real-world videos.
Collecting highly paral- lel data for paraphrase evaluation
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding
A new 1,000-video benchmark with adversarial question variants shows video LLMs remain far below human accuracy and robustness on short real-world videos.