Combining cell-level local captions from grid-overlaid frames with global frame captions improves zero-shot video question answering from text-only transcripts.
Baseline models We leveraged the LLaV A-1.6-7B VLM [23] to generate local and global transcriptions and the Llama-3.1-8B LLM [24] for the VideoQA task
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering
Combining cell-level local captions from grid-overlaid frames with global frame captions improves zero-shot video question answering from text-only transcripts.