Combining cell-level local captions from grid-overlaid frames with global frame captions improves zero-shot video question answering from text-only transcripts.
The system is inherently modular and built on open-source VLM and LLM
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering
Combining cell-level local captions from grid-overlaid frames with global frame captions improves zero-shot video question answering from text-only transcripts.