AutoCaption uses MCTS to generate fine-grained video key points, forming the MCTS-VCB benchmark that ranks MLLMs and yields training data improving a fine-tuned model's captioning.
As shown in Figure 6, the Nature and Wildlife category has the highest number of videos, reaching 382, while the Arts and Creativity category has the fewest, with 57 videos
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search
AutoCaption uses MCTS to generate fine-grained video key points, forming the MCTS-VCB benchmark that ranks MLLMs and yields training data improving a fine-tuned model's captioning.