AutoCaption uses MCTS to generate fine-grained video key points, forming the MCTS-VCB benchmark that ranks MLLMs and yields training data improving a fine-tuned model's captioning.
The defination and examples of each category is described in Figure 10
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search
AutoCaption uses MCTS to generate fine-grained video key points, forming the MCTS-VCB benchmark that ranks MLLMs and yields training data improving a fine-tuned model's captioning.