VideoRewardBench, a 1,563-sample benchmark, finds all 28 tested video reward models score below 64% accuracy, exposing a large gap in video-based reward modeling.
Mllm-as-a-judge: Assessing multimodal llm-as-a-judge with vision-language benchmark
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding
VideoRewardBench, a 1,563-sample benchmark, finds all 28 tested video reward models score below 64% accuracy, exposing a large gap in video-based reward modeling.