Combining video frames with transcript text in an RL-trained summarizer improves rank-based highlight metrics and summary fidelity on Mr. HiSum, while slightly hurting F1 and top-5% highlight detection.
UMT: Un ified multi-modal transformers for joint video moment retrieval and highlight detection,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Unsupervised Transcript-assisted Video Summarization and Highlight Detection
Combining video frames with transcript text in an RL-trained summarizer improves rank-based highlight metrics and summary fidelity on Mr. HiSum, while slightly hurting F1 and top-5% highlight detection.