Pith. sign in

REVIEW 22 cited by

EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2308.09126 v1 pith:W7KJSUVV submitted 2023-08-17 cs.CV cs.AIcs.CL

EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

classification cs.CV cs.AIcs.CL
keywords videoegoschemaunderstandingtemporalclipdatasetevaluationintrinsic
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

We introduce EgoSchema, a very long-form video question-answering dataset, and benchmark to evaluate long video understanding capabilities of modern vision and language systems. Derived from Ego4D, EgoSchema consists of over 5000 human curated multiple choice question answer pairs, spanning over 250 hours of real video data, covering a very broad range of natural human activity and behavior. For each question, EgoSchema requires the correct answer to be selected between five given options based on a three-minute-long video clip. While some prior works have proposed video datasets with long clip lengths, we posit that merely the length of the video clip does not truly capture the temporal difficulty of the video task that is being considered. To remedy this, we introduce temporal certificate sets, a general notion for capturing the intrinsic temporal understanding length associated with a broad range of video understanding tasks & datasets. Based on this metric, we find EgoSchema to have intrinsic temporal lengths over 5.7x longer than the second closest dataset and 10x to 100x longer than any other video understanding dataset. Further, our evaluation of several current state-of-the-art video and language models shows them to be severely lacking in long-term video understanding capabilities. Even models with several billions of parameters achieve QA accuracy less than 33% (random is 20%) on the EgoSchema multi-choice question answering task, while humans achieve about 76% accuracy. We posit that \name{}{}, with its long intrinsic temporal structures and diverse complexity, would serve as a valuable evaluation probe for developing effective long-term video understanding systems in the future. Data and Zero-shot model evaluation code are open-sourced for both public and commercial use under the Ego4D license at http://egoschema.github.io

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 22 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. EgoSAT: A Comprehensive Benchmark of Egocentric Streaming Interaction Understanding

    cs.CV 2026-06 unverdicted novelty 7.0

    EgoSAT is the first benchmark unifying retrospective, online, and prospective reasoning tasks in egocentric streaming video to evaluate VLMs, revealing struggles with temporal modeling and mis-calibration.

  2. NextMotionQA: Benchmarking and Judging Human Motion Understanding with Vision-Language Models

    cs.CV 2026-06 unverdicted novelty 7.0

    NextMotionQA benchmark reveals VLMs have critical gaps in fine-grained human motion understanding and align with experts on coarse judgment (κ=0.70) but not fine-grained (κ=0.10).

  3. SuperMemory-VQA: An Egocentric Visual Question-Answering Benchmark for Long-Horizon Memory

    cs.CV 2026-05 unverdicted novelty 7.0

    SuperMemory-VQA provides 4,853 human-verified QA pairs from 52.9 hours of egocentric AI glasses recordings to benchmark AI systems on realistic long-horizon memory tasks including an unanswerable option.

  4. Pause and Think: A Dataset and Benchmark for Video-Grounded Assistive Action Suggestion

    cs.CV 2026-05 unverdicted novelty 7.0

    Introduces pause-and-think-T dataset and pause-and-think-B benchmark; fine-tunes 4B VLM to 58% accuracy matching 235B model while generalizing out-of-distribution.

  5. CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models

    cs.CV 2026-08 conditional novelty 6.0

    CRAFT recursively merges video tokens with training-free similarity selection plus learnable gated fusion, retaining ~97% of average accuracy at 8x compression across six benchmarks.

  6. ObjectStream: Latent Objects as Memory Anchors for Streaming Video Understanding

    cs.CV 2026-07 conditional novelty 6.0

    Training-free latent-object memory anchors let frozen Video-LLMs retain object histories under a tight token budget and improve streaming and long-video QA.

  7. ObjectStream: Latent Objects as Memory Anchors for Streaming Video Understanding

    cs.CV 2026-07 conditional novelty 6.0

    A training-free memory framework that anchors streaming video memory to latent objects discovered from frozen Video-LLM features, improving streaming QA accuracy while cutting memory and latency.

  8. MEMORA: Embodied Action Memory from Egocentric Videos for Reasoning and Planning

    cs.RO 2026-07 conditional novelty 6.0

    A typed, editable memory built from egocentric video improves memory-grounded question answering and out-of-distribution robot planning over flat-text and graph baselines.

  9. From Accuracy to Visual Dependence: Auditing and Filtering Modality Collapse in Traffic VideoQA

    cs.CV 2026-06 unverdicted novelty 6.0

    Audit of four VideoQA benchmarks reveals text-only shortcuts in VLMs; new diagnostics Blind Gap, Visual Gain, and Shortcut Score quantify and filter visual dependence.

  10. MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention

    cs.CV 2026-06 unverdicted novelty 6.0

    MOSS-Video-Preview introduces a cross-attention architecture and synthesized real-time QA data to enable continuous perception, answer revision, and faster inference in video-language models compared to decoder-only designs.

  11. Pause and Think: A Dataset and Benchmark for Video-Grounded Assistive Action Suggestion

    cs.CV 2026-05 unverdicted novelty 6.0

    Introduces pause-and-think-T dataset and pause-and-think-B benchmark for video-grounded assistive action suggestion, enabling a 4B VLM to match larger models on reasoning tasks and generalize to EgoThink and TempCompass.

  12. TeachObs: A Human-Validated Benchmark for Multimodal Teaching Observation and Model Evaluation

    cs.CL 2026-05 unverdicted novelty 6.0

    TeachObs is a new human-validated benchmark dataset and evaluation protocol for multimodal AI on classroom teaching observation, showing no model dominates across tracks and that models over-rate procedurally clear lessons.

  13. TeachObs: A Human-Validated Benchmark for Multimodal Teaching Observation and Model Evaluation

    cs.CL 2026-05 unverdicted novelty 6.0

    TeachObs provides human-validated segment- and lesson-level labels on 30 classroom videos and finds frontier VLMs inconsistent, frame-inflated, and over-generous versus expert raters.

  14. EgoProx: Evaluating MLLMs on Egocentric 3D Proximity Reasoning Across a Cognitive Hierarchy

    cs.CV 2026-05 unverdicted novelty 6.0

    EgoProx benchmark shows MLLMs have some spatial knowledge but struggle to leverage it for egocentric 3D proximity reasoning VQA.

  15. HumanNet: Scaling Human-centric Video Learning to One Million Hours

    cs.CV 2026-05 unverdicted novelty 6.0

    HumanNet is a 1M-hour human-centric video dataset with interaction annotations that enables better vision-language-action model performance than equivalent robot data in a controlled test.

  16. MACS: Modality-Aware Capacity Scaling for Efficient Multimodal MoE Inference

    cs.LG 2026-04 unverdicted novelty 6.0

    MACS improves inference speed in multimodal MoE models by entropy-weighted balancing of visual tokens and real-time modality-adaptive expert capacity allocation.

  17. MACS: Modality-Aware Capacity Scaling for Efficient Multimodal MoE Inference

    cs.LG 2026-04 unverdicted novelty 6.0

    MACS improves MoE MLLM inference efficiency via entropy-weighted token loads and dynamic modality-adaptive expert capacity allocation.

  18. Adaptive Greedy Frame Selection for Long Video Understanding

    cs.CV 2026-03 unverdicted novelty 6.0

    A question-adaptive greedy frame selector combines SigLIP relevance and DINOv2 coverage under a submodular objective with a text classifier routing to preset trade-offs, yielding accuracy gains on MLVU especially at l...

  19. HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding

    cs.CV 2026-01 unverdicted novelty 6.0

    HERMES organizes the KV cache into a hierarchical memory to enable real-time streaming video understanding in MLLMs, achieving 10x faster TTFT and up to 11.4% accuracy gains on streaming benchmarks with 68% fewer tokens.

  20. MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

    cs.CV 2023-11 accept novelty 6.0

    MVBench is a benchmark of 20 temporal video understanding tasks built by transforming static tasks into dynamic ones, with VideoChat2 outperforming prior MLLMs by over 15%.

  21. EgoCoT-Bench: Benchmarking Grounded and Verifiable Operation-Centric Chain of Thought Reasoning for MLLMs

    cs.CV 2026-05 unverdicted novelty 5.0

    EgoCoT-Bench provides 3,172 verifiable QA pairs across perception, anticipation, and reasoning tasks on egocentric videos, revealing that many MLLMs give answer-correct but evidence-inconsistent explanations.

  22. MACS: Modality-Aware Capacity Scaling for Efficient Multimodal MoE Inference

    cs.LG 2026-04 unverdicted novelty 5.0

    MACS reduces Expert Parallelism stragglers in MoE MLLMs via entropy-weighted visual token load and dynamic modality-adaptive expert capacity, without retraining.