A new dataset of 420 math video-question pairs with step-by-step reasoning annotations shows that current multimodal AI models, including the best proprietary system, answer fewer than half of the multi-binary questions correctly.
Mmbench-video: A long-form multi-shot benchmark for holistic video understanding
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
dataset 1
citation-polarity summary
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1roles
dataset 1polarities
baseline 1representative citing papers
citing papers explorer
-
VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos
A new dataset of 420 math video-question pairs with step-by-step reasoning annotations shows that current multimodal AI models, including the best proprietary system, answer fewer than half of the multi-binary questions correctly.