Pith. sign in

REVIEW 12 cited by

EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2308.09126 v1 pith:W7KJSUVV submitted 2023-08-17 cs.CV cs.AIcs.CL

classification cs.CVcs.AIcs.CL
keywords videoegoschemaunderstandingtemporalclipdatasetevaluationintrinsic
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We introduce EgoSchema, a very long-form video question-answering dataset, and benchmark to evaluate long video understanding capabilities of modern vision and language systems. Derived from Ego4D, EgoSchema consists of over 5000 human curated multiple choice question answer pairs, spanning over 250 hours of real video data, covering a very broad range of natural human activity and behavior. For each question, EgoSchema requires the correct answer to be selected between five given options based on a three-minute-long video clip. While some prior works have proposed video datasets with long clip lengths, we posit that merely the length of the video clip does not truly capture the temporal difficulty of the video task that is being considered. To remedy this, we introduce temporal certificate sets, a general notion for capturing the intrinsic temporal understanding length associated with a broad range of video understanding tasks & datasets. Based on this metric, we find EgoSchema to have intrinsic temporal lengths over 5.7x longer than the second closest dataset and 10x to 100x longer than any other video understanding dataset. Further, our evaluation of several current state-of-the-art video and language models shows them to be severely lacking in long-term video understanding capabilities. Even models with several billions of parameters achieve QA accuracy less than 33% (random is 20%) on the EgoSchema multi-choice question answering task, while humans achieve about 76% accuracy. We posit that \name{}{}, with its long intrinsic temporal structures and diverse complexity, would serve as a valuable evaluation probe for developing effective long-term video understanding systems in the future. Data and Zero-shot model evaluation code are open-sourced for both public and commercial use under the Ego4D license at http://egoschema.github.io

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding

    cs.CV 2025-06 conditional novelty 7.0 of 10

    MF2 evaluates long-movie understanding by asking models to classify fact/fib claim pairs; the best model trails humans by 23.5 points in pairwise accuracy.

  2. Online Video Understanding: OVBench and VideoChat-Online

    cs.CV 2024-12 conditional novelty 7.0 of 10

    A new benchmark and a memory-bank architecture let a 4B video AI outperform larger offline and streaming models on online video understanding tasks.

  3. CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models

    cs.CV 2026-08 conditional novelty 6.0 of 10

    CRAFT recursively merges video tokens with training-free similarity selection plus learnable gated fusion, retaining ~97% of average accuracy at 8x compression across six benchmarks.

  4. ObjectStream: Latent Objects as Memory Anchors for Streaming Video Understanding

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Training-free latent-object memory anchors let frozen Video-LLMs retain object histories under a tight token budget and improve streaming and long-video QA.

  5. MEMORA: Embodied Action Memory from Egocentric Videos for Reasoning and Planning

    cs.RO 2026-07 conditional novelty 6.0 of 10

    A typed, editable memory built from egocentric video improves memory-grounded question answering and out-of-distribution robot planning over flat-text and graph baselines.

  6. CausalVQA: A Physically Grounded Causal Reasoning Benchmark for Video Models

    cs.CV 2025-06 conditional novelty 6.0 of 10

    CausalVQA provides 793 paired real-video causal reasoning questions on which the best multimodal model scores 61.66% versus 84.78% for humans, with the largest gaps on anticipation and hypothetical questions.

  7. VCA: Video Curious Agent for Long Video Understanding

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A training-free video agent that combines segment-level tree search, self-generated intrinsic rewards, and a fixed memory buffer achieves higher long-video QA accuracy with fewer observed frames than prior agent baselines.

  8. LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models

    cs.CV 2025-07 conditional novelty 5.0 of 10

    Inserting a pixel-shuffle plus residual patch-merge layer inside the vision encoder compresses visual tokens more efficiently than post-encoder compression, at modest accuracy cost.

  9. MASR: Self-Reflective Reasoning through Multimodal Hierarchical Attention Focusing for Agent-based Video Understanding

    cs.CV 2025-04 conditional novelty 5.0 of 10

    MASR combines coarse-to-fine relevance selection, dilated temporal expansion, and confidence-driven self-reflection to improve agent-based video question answering, and reports strong benchmark results.

  10. $\infty$-Video: A Training-Free Approach to Long Video Understanding via Continuous-Time Memory Consolidation

    cs.CV 2025-01 conditional novelty 5.0 of 10

    ∞-Video extends ∞-former continuous-attention ideas to video Q-formers, adding a training-free long-term memory that improves long-video question answering in several benchmarks.

  11. Infinite Video Understanding

    cs.CV 2025-07 conditional novelty 3.0 of 10

    The paper argues that video understanding research should aim at processing streams of arbitrary, unbounded duration and outlines the challenges, directions, and metrics needed.

  12. VideoLLM Benchmarks and Evaluation: A Survey

    cs.CV 2025-05 unverdicted novelty 1.0 of 10

    A survey of VideoLLM benchmarks and evaluation protocols that organizes known datasets and metrics, with proposed future benchmark designs.

Pith tools