Pith. sign in

REVIEW 1 cited by

Activity Graph Transformer for Temporal Action Localization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2101.08540 v2 pith:BHWCDJO5 submitted 2021-01-21 cs.CV cs.AI

classification cs.CVcs.AI
keywords actioninstancestemporalvideovideosmodelactivitydirectly
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce Activity Graph Transformer, an end-to-end learnable model for temporal action localization, that receives a video as input and directly predicts a set of action instances that appear in the video. Detecting and localizing action instances in untrimmed videos requires reasoning over multiple action instances in a video. The dominant paradigms in the literature process videos temporally to either propose action regions or directly produce frame-level detections. However, sequential processing of videos is problematic when the action instances have non-sequential dependencies and/or non-linear temporal ordering, such as overlapping action instances or re-occurrence of action instances over the course of the video. In this work, we capture this non-linear temporal structure by reasoning over the videos as non-sequential entities in the form of graphs. We evaluate our model on challenging datasets: THUMOS14, Charades, and EPIC-Kitchens-100. Our results show that our proposed model outperforms the state-of-the-art by a considerable margin.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. EVAS: Efficient Multimodal Temporal Forgery Localization via Audio-Visual Synergy and Steered Boundary Calibration

    cs.CV 2026-07 conditional novelty 6.0 of 10

    EVAS localizes sparse multimodal forgeries via multi-stage audio-visual synergy and decoupled boundary-aware refinement, reporting SOTA AP and AR on LAV-DF, AV-Deepfake1M, and TVIL.

Pith tools