Pith. sign in

REVIEW 2 cited by

VideoGraph: Recognizing Minutes-Long Human Activities in Videos

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1905.05143 v2 pith:AFPQFCGK submitted 2019-05-13 cs.CV

classification cs.CV
keywords activitieshumantemporalvideographminutes-longstructurelearnrelated
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Many human activities take minutes to unfold. To represent them, related works opt for statistical pooling, which neglects the temporal structure. Others opt for convolutional methods, as CNN and Non-Local. While successful in learning temporal concepts, they are short of modeling minutes-long temporal dependencies. We propose VideoGraph, a method to achieve the best of two worlds: represent minutes-long human activities and learn their underlying temporal structure. VideoGraph learns a graph-based representation for human activities. The graph, its nodes and edges are learned entirely from video datasets, making VideoGraph applicable to problems without node-level annotation. The result is improvements over related works on benchmarks: Epic-Kitchen and Breakfast. Besides, we demonstrate that VideoGraph is able to learn the temporal structure of human activities in minutes-long videos.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Cross-Modal Dual-Causal Learning for Long-Term Action Recognition

    cs.CV 2025-07 reject novelty 6.0 of 10

    CMDCL debiases text embeddings by back-door adjustment and deconfounds video features by front-door adjustment, achieving state-of-the-art long-term action recognition on three benchmarks.

  2. MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition

    cs.CV 2025-06 conditional novelty 5.0 of 10

    MoMa adapts frozen CLIP to video by injecting Mamba-computed scale and bias into each layer, improving accuracy and efficiency on multiple action recognition benchmarks.

Pith tools