REVIEW 6 cited by
Revealing Single Frame Bias for Video-and-Language Learning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Training an effective video-and-language model intuitively requires multiple frames as model inputs. However, it is unclear whether using multiple frames is beneficial to downstream tasks, and if yes, whether the performance gain is worth the drastically-increased computation and memory costs resulting from using more frames. In this work, we explore single-frame models for video-and-language learning. On a diverse set of video-and-language tasks (including text-to-video retrieval and video question answering), we show the surprising result that, with large-scale pre-training and a proper frame ensemble strategy at inference time, a single-frame trained model that does not consider temporal information can achieve better performance than existing methods that use multiple frames for training. This result reveals the existence of a strong "static appearance bias" in popular video-and-language datasets. Therefore, to allow for a more comprehensive evaluation of video-and-language models, we propose two new retrieval tasks based on existing fine-grained action recognition datasets that encourage temporal modeling. Our code is available at https://github.com/jayleicn/singularity
Forward citations
Cited by 6 Pith papers
-
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering
MUPA combines three ordering-based reasoning paths with a reflection agent that verifies and fuses answer-evidence pairs, reaching 30.3% and 47.4% grounded QA accuracy on NExT-GQA and DeVE-QA with a 7B model.
-
RTime-QA: A Benchmark for Atomic Temporal Event Understanding in Large Multi-modal Models
RTime-QA is a video-question benchmark where models choose between temporally opposite descriptions of the same event, and current AI models score far below humans.
-
Temporal-Oriented Recipe for Transferring Large Vision-Language Model to Video Understanding
A four-step recipe (Q-Former interface, temporal training, memory bank, MoE) improves large vision-language models on standard video QA and captioning benchmarks.
-
Progress-Aware Video Frame Captioning
A two-stage model trained on VLM-generated, critic-filtered captions produces frame-level captions that track action progression and outperforms existing VLMs on the new FrameCapEval benchmark.
-
Motion-Grounded Video Reasoning: Understanding and Perceiving Motion at Pixel Level
A new benchmark and baseline for motion-grounded video reasoning, where the answer to a motion question is a spatiotemporal segmentation mask.
-
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering
FocusChat uses prompt-guided spatial and temporal filtering to cut visual tokens to as few as 16 while matching or beating a larger Video-LLaMA baseline.
Discussion (0). Continue with ORCID to comment.