Pith. sign in

REVIEW 6 cited by

Revealing Single Frame Bias for Video-and-Language Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2206.03428 v1 pith:6YQPEPUD submitted 2022-06-07 cs.CV cs.AIcs.CL

classification cs.CVcs.AIcs.CL
keywords video-and-languageframesmodelmultipletasksbiasdatasetsexisting
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Training an effective video-and-language model intuitively requires multiple frames as model inputs. However, it is unclear whether using multiple frames is beneficial to downstream tasks, and if yes, whether the performance gain is worth the drastically-increased computation and memory costs resulting from using more frames. In this work, we explore single-frame models for video-and-language learning. On a diverse set of video-and-language tasks (including text-to-video retrieval and video question answering), we show the surprising result that, with large-scale pre-training and a proper frame ensemble strategy at inference time, a single-frame trained model that does not consider temporal information can achieve better performance than existing methods that use multiple frames for training. This result reveals the existence of a strong "static appearance bias" in popular video-and-language datasets. Therefore, to allow for a more comprehensive evaluation of video-and-language models, we propose two new retrieval tasks based on existing fine-grained action recognition datasets that encourage temporal modeling. Our code is available at https://github.com/jayleicn/singularity

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering

    cs.CV 2025-06 conditional novelty 6.0 of 10

    MUPA combines three ordering-based reasoning paths with a reflection agent that verifies and fuses answer-evidence pairs, reaching 30.3% and 47.4% grounded QA accuracy on NExT-GQA and DeVE-QA with a 7B model.

  2. RTime-QA: A Benchmark for Atomic Temporal Event Understanding in Large Multi-modal Models

    cs.CV 2025-05 conditional novelty 6.0 of 10

    RTime-QA is a video-question benchmark where models choose between temporally opposite descriptions of the same event, and current AI models score far below humans.

  3. Temporal-Oriented Recipe for Transferring Large Vision-Language Model to Video Understanding

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A four-step recipe (Q-Former interface, temporal training, memory bank, MoE) improves large vision-language models on standard video QA and captioning benchmarks.

  4. Progress-Aware Video Frame Captioning

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A two-stage model trained on VLM-generated, critic-filtered captions produces frame-level captions that track action progression and outperforms existing VLMs on the new FrameCapEval benchmark.

  5. Motion-Grounded Video Reasoning: Understanding and Perceiving Motion at Pixel Level

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A new benchmark and baseline for motion-grounded video reasoning, where the answer to a motion question is a spatiotemporal segmentation mask.

  6. FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering

    cs.CV 2024-12 conditional novelty 5.0 of 10

    FocusChat uses prompt-guided spatial and temporal filtering to cut visual tokens to as few as 16 while matching or beating a larger Video-LLaMA baseline.

Pith tools