Pith. sign in

REVIEW 9 cited by

HourVideo: 1-Hour Video-Language Understanding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.04998 v1 pith:74CX4QQK submitted 2024-11-07 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords hourvideodatasetmultimodalbenchmarkunderstandingvideo-languageachieveavailable
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We present HourVideo, a benchmark dataset for hour-long video-language understanding. Our dataset consists of a novel task suite comprising summarization, perception (recall, tracking), visual reasoning (spatial, temporal, predictive, causal, counterfactual), and navigation (room-to-room, object retrieval) tasks. HourVideo includes 500 manually curated egocentric videos from the Ego4D dataset, spanning durations of 20 to 120 minutes, and features 12,976 high-quality, five-way multiple-choice questions. Benchmarking results reveal that multimodal models, including GPT-4 and LLaVA-NeXT, achieve marginal improvements over random chance. In stark contrast, human experts significantly outperform the state-of-the-art long-context multimodal model, Gemini Pro 1.5 (85.0% vs. 37.3%), highlighting a substantial gap in multimodal capabilities. Our benchmark, evaluation toolkit, prompts, and documentation are available at https://hourvideo.stanford.edu

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models

    cs.CV 2025-01 conditional novelty 7.0 of 10

    Consistent object IDs across video frames and a reconstructed bird's-eye view let large VLMs, and a fine-tuned 7B model, reach state-of-the-art results on ScanNet 3D question answering, captioning, and grounding.

  2. EgoEverything: A Benchmark for Human Behavior Inspired Long Context Egocentric Video Understanding in AR Environment

    cs.LG 2026-04 unverdicted novelty 6.0 of 10

    EgoEverything is a new benchmark for long-context egocentric video understanding that uses human gaze-based attention signals to generate questions reflecting natural behavior.

  3. CausalVQA: A Physically Grounded Causal Reasoning Benchmark for Video Models

    cs.CV 2025-06 conditional novelty 6.0 of 10

    CausalVQA provides 793 paired real-video causal reasoning questions on which the best multimodal model scores 61.66% versus 84.78% for humans, with the largest gaps on anticipation and hypothetical questions.

  4. VideoConviction: A Multimodal Benchmark for Human Conviction and Stock Market Recommendations

    cs.MM 2025-06 conditional novelty 6.0 of 10

    VideoConviction provides the first expert-annotated multimodal benchmark of financial influencer video recommendations, showing MLLMs extract tickers better but struggle with actions and conviction, and that an invers...

  5. MINERVA: Evaluating Complex Video Reasoning

    cs.LG 2025-05 conditional novelty 6.0 of 10

    MINERVA provides 1,515 multi-step video QA questions with human reasoning traces; frontier models score far below humans and fail mainly on temporal localization and perception.

  6. HLV-1K: A Large-scale Hour-Long Video Benchmark for Time-Specific Long Video Understanding

    cs.CV 2025-01 conditional novelty 6.0 of 10

    HLV-1K is a new benchmark of over one thousand long videos with time-specific question-answer pairs, used to measure how well AI models understand hour-scale video.

  7. VideoMultiAgents: A Multi-Agent Framework for Video Question Answering

    cs.CV 2025-04 conditional novelty 5.0 of 10

    A multi-agent video QA framework with independent text, video, and scene-graph agents plus an organizer achieves state-of-the-art zero-shot scores on Intent-QA, the EgoSchema subset, and NExT-QA.

  8. VideoRAG: Retrieval-Augmented Generation with Extreme Long-Context Videos

    cs.IR 2025-02 conditional novelty 5.0 of 10

    VideoRAG combines graph-based text indexing with multimodal visual embeddings to answer questions across multi-hour video collections, supported by a new 134-hour benchmark.

  9. Infinite Video Understanding

    cs.CV 2025-07 conditional novelty 3.0 of 10

    The paper argues that video understanding research should aim at processing streams of arbitrary, unbounded duration and outlines the challenges, directions, and metrics needed.

Pith tools