Pith. sign in

REVIEW 26 cited by

TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.10818 v2 pith:DEIB5PUN submitted 2024-10-14 cs.CV cs.AIcs.CLcs.LG

classification cs.CVcs.AIcs.CLcs.LG
keywords temporalvideomodelsunderstandingtemporalbenchfine-grainedevaluatingmultimodal
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Understanding fine-grained temporal dynamics is crucial for multimodal video comprehension and generation. Due to the lack of fine-grained temporal annotations, existing video benchmarks mostly resemble static image benchmarks and are incompetent at evaluating models for temporal understanding. In this paper, we introduce TemporalBench, a new benchmark dedicated to evaluating fine-grained temporal understanding in videos. TemporalBench consists of ~10K video question-answer pairs, derived from ~2K high-quality human annotations detailing the temporal dynamics in video clips. As a result, our benchmark provides a unique testbed for evaluating various temporal understanding and reasoning abilities such as action frequency, motion magnitude, event order, etc. Moreover, it enables evaluations on various tasks like both video question answering and captioning, both short and long video understanding, as well as different models such as multimodal video embedding models and text generation models. Results show that state-of-the-art models like GPT-4o achieve only 38.5% question answering accuracy on TemporalBench, demonstrating a significant gap (~30%) between humans and AI in temporal understanding. Furthermore, we notice a critical pitfall for multi-choice QA where LLMs can detect the subtle changes in negative captions and find a centralized description as a cue for its prediction, where we propose Multiple Binary Accuracy (MBA) to correct such bias. We hope that TemporalBench can foster research on improving models' temporal reasoning capabilities. Both dataset and evaluation code will be made available.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 26 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. VADER: Adaptive Debiasing for Hallucination Mitigation in Video Large Language Models

    cs.CV 2026-08 conditional novelty 7.0 of 10

    A training-free decoding framework that adaptively reweights attention toward video tokens and erases key visual evidence per frame to suppress hallucinated predictions, achieving 72.60% accuracy on EventHallusion wit...

  2. The TIME Machine: On The Power of Motion for Efficient Perception

    cs.CV 2026-05 unverdicted novelty 7.0 of 10

    TIME is a motion-based embedding from point tracks, trained only on synthetic data via masked autoencoding, that matches state-of-the-art video model performance with up to 10,000x less training data.

  3. PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding

    cs.CV 2025-04 conditional novelty 7.0 of 10

    PerceptionLM releases 2.8M human-labeled fine-grained video QA pairs and spatio-temporal captions, plus models and a new benchmark, arguing that human data, not just synthetic data, is needed for detailed video understanding.

  4. HD-EPIC: A Highly-Detailed Egocentric Video Dataset

    cs.CV 2025-02 conditional novelty 7.0 of 10

    A new densely annotated, 3D-grounded egocentric kitchen dataset with a 26K-question VQA benchmark that current video-language models mostly fail.

  5. MMVU: Measuring Expert-Level Multi-Discipline Video Understanding

    cs.CV 2025-01 conditional novelty 7.0 of 10

    MMVU is a new expert-annotated video benchmark across 27 subjects, and top AI models remain far behind human open-book performance.

  6. FADE: From Passive Verification to Active Discovery in Counterfactual Video Understanding

    cs.CV 2026-08 conditional novelty 6.0 of 10

    FADE trains a video MLLM with evidence-internalized SFT plus fading-anchor RL, preserving counterfactual judgment accuracy when MCQ guidance is removed, with 90.4% and 67.4% retention on OQA and captioning on DualityV...

  7. From Segments to Scenes: Temporal Understanding for Agentic Autonomous Driving via Vision-Language Models

    cs.CV 2025-12 conditional novelty 6.0 of 10

    TAD, a 5,861-question benchmark, shows VLMs score far below humans on temporal understanding of driving videos, and an ego-trajectory text summary (TCogMap) substantially boosts their scores.

  8. AdsQA: Towards Advertisement Video Understanding

    cs.CV 2025-09 conditional novelty 6.0 of 10

    AdsQA adds an ad-video question-answering benchmark and ReAd-R, a GRPO-trained model that beats 7B baselines but not larger closed models.

  9. CausalStep: A Benchmark for Explicit Stepwise Causal Reasoning in Videos

    cs.CV 2025-07 conditional novelty 6.0 of 10

    CausalStep introduces a stepwise video QA protocol and reports that top multimodal models (chain success rate 51%) remain far below human performance (79%) on explicit causal chains.

  10. Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A new 1,000-video benchmark with adversarial question variants shows video LLMs remain far below human accuracy and robustness on short real-world videos.

  11. GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them?

    cs.CV 2025-07 conditional novelty 6.0 of 10

    Introduces GLIMPSE, a video-QA benchmark whose questions cannot be answered from single frames; best model GPT-o3 scores 66.43% vs 94.82% human accuracy.

  12. ProactiveVideoQA: A Comprehensive Benchmark Evaluating Proactive Interactions in Video Large Language Models

    cs.CV 2025-07 conditional novelty 6.0 of 10

    ProactiveVideoQA is a benchmark for proactive video question answering, and the proposed PAUC metric jointly scores response timing and content, claiming better alignment with human preferences.

  13. Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Video-language models perform far below humans on a new quadrilingual benchmark that tests understanding of action completion and duration through grammatical aspect.

  14. Fostering Video Reasoning via Next-Event Prediction

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Next-event prediction, training video language models to caption unseen future frames, improves their scores on several temporal benchmarks while roughly preserving general video understanding.

  15. TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos

    cs.CV 2025-05 conditional novelty 6.0 of 10

    TUNA introduces a 1,000-video benchmark with dense temporal captions and 1,432 multiple-choice questions, and finds that current video LMMs are weakest at camera motion, action sequences, and multi-subject scenes.

  16. RTime-QA: A Benchmark for Atomic Temporal Event Understanding in Large Multi-modal Models

    cs.CV 2025-05 conditional novelty 6.0 of 10

    RTime-QA is a video-question benchmark where models choose between temporally opposite descriptions of the same event, and current AI models score far below humans.

  17. VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A new video benchmark asks models to reorder shuffled clips from everyday tasks; most large video language models perform near chance, and a recognize-then-reason prompt decomposition gives large gains.

  18. MINERVA: Evaluating Complex Video Reasoning

    cs.LG 2025-05 conditional novelty 6.0 of 10

    MINERVA provides 1,515 multi-step video QA questions with human reasoning traces; frontier models score far below humans and fail mainly on temporal localization and perception.

  19. Is Your World Simulator a Good Story Presenter? A Consecutive Events-Based Benchmark for Future Long Video Generation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A new story-completion benchmark, StoryEval, shows that 11 current text-to-video models complete fewer than half of the consecutive events in short story prompts.

  20. Apollo: An Exploration of Video Understanding in Large Multimodal Models

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Design decisions for video-LMMs can be made on 2-4B models and datasets and transfer to larger models, yielding efficient Apollo models, though some SOTA claims are contradicted by the paper's own table.

  21. Progress-Aware Video Frame Captioning

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A two-stage model trained on VLM-generated, critic-filtered captions produces frame-level captions that track action progression and outperforms existing VLMs on the new FrameCapEval benchmark.

  22. VideoOrion: Tokenizing Object Dynamics in Videos

    cs.CV 2024-11 conditional novelty 6.0 of 10

    Encoding video as a small set of object tokens, produced by off-the-shelf detection, segmentation, and tracking models, improves video QA accuracy and enables video-based referring in a 7B video-LLM.

  23. NeMo: Needle in a Montage for Video-Language Understanding

    cs.CV 2025-09 conditional novelty 5.0 of 10

    NeMoBench, an automatically generated benchmark with 31,378 QA pairs, shows that video LLMs struggle with temporal grounding of relevant clips hidden in long montages.

  24. Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No!

    cs.CV 2025-01 conditional novelty 5.0 of 10

    A new benchmark shows that GPT-4o and other multimodal LLMs perform near chance on ordering image events and far below humans on estimating time lapses.

  25. VideoCogQA: A Controllable Benchmark for Evaluating Cognitive Abilities in Video-Language Models

    cs.CV 2024-11 conditional novelty 5.0 of 10

    A new controllable synthetic-video benchmark shows that even state-of-the-art video-language models struggle with abstract and symbolic video cognition, with accuracy falling as task difficulty rises.

  26. Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models

    cs.CV 2024-12 reject novelty 4.0 of 10

    The paper claims that a video-language model trained with dynamic temporal prompts and temporal contrastive learning beats four published models on three self-defined VidSitu temporal reasoning tasks.

Pith tools