Combining cell-level local captions from grid-overlaid frames with global frame captions improves zero-shot video question answering from text-only transcripts.
VLAAD: Vision and language assistant for au- tonomous driving,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering
Combining cell-level local captions from grid-overlaid frames with global frame captions improves zero-shot video question answering from text-only transcripts.