Pith. sign in

Zero-Shot Long-Form Video Understanding through Screenplay

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

The Long-form Video Question-Answering task requires the comprehension and analysis of extended video content to respond accurately to questions by utilizing both temporal and contextual information. In this paper, we present MM-Screenplayer, an advanced video understanding system with multi-modal perception capabilities that can convert any video into textual screenplay representations. Unlike previous storytelling methods, we organize video content into scenes as the basic unit, rather than just visually continuous shots. Additionally, we developed a ``Look Back'' strategy to reassess and validate uncertain information, particularly targeting breakpoint mode. MM-Screenplayer achieved highest score in the CVPR'2024 LOng-form VidEo Understanding (LOVEU) Track 1 Challenge, with a global accuracy of 87.5% and a breakpoint accuracy of 68.8%.

fields

cs.CV 1

years

2025 1

verdicts

CONDITIONAL 1

representative citing papers

VEU-Bench: Towards Comprehensive Understanding of Video Editing

cs.CV · 2025-04-24 · conditional · novelty 6.0

A new 19-task video editing benchmark shows that current video LLMs struggle to understand editing concepts, and a model fine-tuned on the benchmark improves both editing and general video reasoning.

citing papers explorer

Showing 1 of 1 citing paper.

  • VEU-Bench: Towards Comprehensive Understanding of Video Editing cs.CV · 2025-04-24 · conditional · none · ref 44 · internal anchor

    A new 19-task video editing benchmark shows that current video LLMs struggle to understand editing concepts, and a model fine-tuned on the benchmark improves both editing and general video reasoning.