Pith. sign in

REVIEW 19 cited by

VideoVista: A Versatile Benchmark for Video Understanding and Reasoning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.11303 v1 pith:ESAUWQAY submitted 2024-06-17 cs.CV cs.AIcs.CL

classification cs.CVcs.AIcs.CL
keywords reasoningvideounderstandingvideo-lmmsvideovistabenchmarklmmsmodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Despite significant breakthroughs in video analysis driven by the rapid development of large multimodal models (LMMs), there remains a lack of a versatile evaluation benchmark to comprehensively assess these models' performance in video understanding and reasoning. To address this, we present VideoVista, a video QA benchmark that integrates challenges across diverse content categories, durations, and abilities. Specifically, VideoVista comprises 25,000 questions derived from 3,400 videos spanning 14 categories (e.g., Howto, Film, and Entertainment) with durations ranging from a few seconds to over 10 minutes. Besides, it encompasses 19 types of understanding tasks (e.g., anomaly detection, interaction understanding) and 8 reasoning tasks (e.g., logical reasoning, causal reasoning). To achieve this, we present an automatic data construction framework, leveraging powerful GPT-4o alongside advanced analysis tools (e.g., video splitting, object segmenting, and tracking). We also utilize this framework to construct training data to enhance the capabilities of video-related LMMs (Video-LMMs). Through a comprehensive and quantitative evaluation of cutting-edge models, we reveal that: 1) Video-LMMs face difficulties in fine-grained video tasks involving temporal location, object tracking, and anomaly detection; 2) Video-LMMs present inferior logical and relation reasoning abilities; 3) Open-source Video-LMMs' performance is significantly lower than GPT-4o and Gemini-1.5, lagging by 20 points. This highlights the crucial role VideoVista will play in advancing LMMs that can accurately understand videos and perform precise reasoning.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 19 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding

    cs.CV 2025-05 conditional novelty 7.0 of 10

    ScaleLong embeds four timescale question types into the same long videos, and evaluation of 23 MLLMs reveals a U-shaped accuracy curve across timescales.

  2. The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping

    cs.AI 2026-08 conditional novelty 6.0 of 10

    Trace-grounded parametric profiling of three synthetic counting tasks shows current video-language models only count reliably at low event counts and low rates, and final-answer accuracy masks poor timestamp-level eve...

  3. VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A person-anchored tree plus multi-agent LLM pipeline lets a system answer cross-video queries about the same person, and it beats single-video models on the authors' new CrossVideoQA benchmark.

  4. SmartHome-Bench: A Comprehensive Benchmark for Video Anomaly Detection in Smart Homes Using Multi-Modal Large Language Models

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A new smart-home video anomaly benchmark and a taxonomy-driven reflective LLM chain that improves MLLM anomaly detection accuracy by 11.62 percentage points over zero-shot prompting.

  5. VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos

    cs.CV 2025-06 conditional novelty 6.0 of 10

    VRBench is a benchmark of 960 long narrative videos with 8,243 human-written multi-step questions, plus a two-level evaluation of answer accuracy and reasoning quality for 31 large models.

  6. VUDG: A Dataset for Video Understanding Domain Generalization

    cs.CV 2025-05 conditional novelty 6.0 of 10

    VUDG is a domain-generalization benchmark for video understanding with 11 domains and 36,388 QA pairs, and it shows that current large video-language models lose accuracy across visual domains.

  7. VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation

    cs.CV 2025-05 conditional novelty 6.0 of 10

    VCapsBench is a video caption quality benchmark with 109,796 QA pairs across 21 fine-grained dimensions on 5,677 videos, evaluating caption accuracy, inconsistency, and coverage.

  8. TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos

    cs.CV 2025-05 conditional novelty 6.0 of 10

    TUNA introduces a 1,000-video benchmark with dense temporal captions and 1,432 multiple-choice questions, and finds that current video LMMs are weakest at camera motion, action sequences, and multi-subject scenes.

  9. HLV-1K: A Large-scale Hour-Long Video Benchmark for Time-Specific Long Video Understanding

    cs.CV 2025-01 conditional novelty 6.0 of 10

    HLV-1K is a new benchmark of over one thousand long videos with time-specific question-answer pairs, used to measure how well AI models understand hour-scale video.

  10. Apollo: An Exploration of Video Understanding in Large Multimodal Models

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Design decisions for video-LMMs can be made on 2-4B models and datasets and transfer to larger models, yielding efficient Apollo models, though some SOTA claims are contradicted by the paper's own table.

  11. Neptune: The Long Orbit to Benchmarking Long Video Understanding

    cs.LG 2024-12 conditional novelty 6.0 of 10

    Neptune is a 3,268-question, 2,405-video benchmark for long video understanding with a scalable LLM-based generation pipeline and an open-source answer-equivalence metric, GEM.

  12. VISTA: Enhancing Long-Duration and High-Resolution Video Understanding by Video Spatiotemporal Augmentation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    VISTA creates synthetic long and high-resolution video instruction data by combining existing clips, and finetuning video LMMs on it improves their accuracy on long-video and high-resolution benchmarks.

  13. TimeMarker: A Versatile Video-LLM for Long and Short Video Understanding with Superior Temporal Localization Ability

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A video-LLM using text-form absolute timestamp tokens and length-adaptive frame/pooling settings achieves competitive-to-leading results on several short and long video benchmarks, with its cleanest temporal-grounding...

  14. NeMo: Needle in a Montage for Video-Language Understanding

    cs.CV 2025-09 conditional novelty 5.0 of 10

    NeMoBench, an automatically generated benchmark with 31,378 QA pairs, shows that video LLMs struggle with temporal grounding of relevant clips hidden in long montages.

  15. VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos

    cs.CV 2025-05 conditional novelty 5.0 of 10

    A new benchmark, VF-Eval, measures how well multimodal LLMs check, detect, and reason about errors in AI-generated videos, and shows frontier models remain far below human performance.

  16. VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation

    cs.CV 2025-05 conditional novelty 5.0 of 10

    An open-ended short-answer long-video benchmark, built by converting MCQ questions from four existing tests, shows large accuracy drops and different model rankings versus multiple-choice evaluation.

  17. LinVT: Empower Your Image-level Large Language Model to Understand Videos

    cs.CV 2024-12 conditional novelty 5.0 of 10

    A plug-and-play linear video tokenizer converts existing image-based LLMs into video-understanding LLMs by condensing frames into weighted-average tokens while preserving image capabilities.

  18. VidHalluc: Evaluating Temporal Hallucinations in Multimodal Large Language Models for Video Understanding

    cs.CV 2024-12 conditional novelty 5.0 of 10

    VidHalluc is a 5,002-video paired benchmark for action, temporal sequence, and scene transition hallucinations in video MLLMs, and DINO-HEAL is a training-free saliency reweighting method that improves hallucination s...

  19. VideoCogQA: A Controllable Benchmark for Evaluating Cognitive Abilities in Video-Language Models

    cs.CV 2024-11 conditional novelty 5.0 of 10

    A new controllable synthetic-video benchmark shows that even state-of-the-art video-language models struggle with abstract and symbolic video cognition, with accuracy falling as task difficulty rises.

Pith tools