GCR is a training-free frame-selection pipeline that binds subtitle text to its aligned frame, fills the budget with diverse context, then replaces weak context frames with medoids from omitted regions, improving long-video QA by 2 to 3 points.
MovieChat+: Question-Aware Sparse Memory for Long Video Question Answering , year=
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Ground, Cover, and Refine: Evidence-Centric Frame Selection for Long-Video Question Answering
GCR is a training-free frame-selection pipeline that binds subtitle text to its aligned frame, fills the budget with diverse context, then replaces weak context frames with medoids from omitted regions, improving long-video QA by 2 to 3 points.