Introduces VG-GUIBench benchmark and TASKER keyframe extraction algorithm that improves performance on VideoQA and video-guided agentic tasks.
Temporal preference optimization for long-form video understanding
3 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
fields
cs.CV 3roles
background 1polarities
background 1representative citing papers
UnifiedReward is the first unified reward model that jointly assesses multimodal understanding and generation to provide better preference signals for aligning vision models via DPO.
CREST selects video frames using curvature-adaptive non-maximum suppression on CLIP relevance scores, beating AKS by ~0.5% on two long-video QA benchmarks while using a fraction of MIRA's preprocessing cost.
citing papers explorer
-
Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe Extraction
Introduces VG-GUIBench benchmark and TASKER keyframe extraction algorithm that improves performance on VideoQA and video-guided agentic tasks.
-
Unified Reward Model for Multimodal Understanding and Generation
UnifiedReward is the first unified reward model that jointly assesses multimodal understanding and generation to provide better preference signals for aligning vision models via DPO.
-
CREST: Curvature-Regulated Event-Centric Sampling for Efficient Long-Video Understanding
CREST selects video frames using curvature-adaptive non-maximum suppression on CLIP relevance scores, beating AKS by ~0.5% on two long-video QA benchmarks while using a fraction of MIRA's preprocessing cost.