Pith. sign in

Video action detection by learning graph-based spatio-temporal interactions

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Action Detection is a complex task that aims to detect and classify human actions in video clips. Typically, it has been addressed by processing fine-grained features extracted from a video classification backbone. Recently, thanks to the robustness of object and people detectors, a deeper focus has been added on relationship modelling. Following this line, we propose a graph-based framework to learn high-level interactions between people and objects, in both space and time. In our formulation, spatio-temporal relationships are learned through self-attention on a multi-layer graph structure which can connect entities from consecutive clips, thus considering long-range spatial and temporal dependencies. The proposed module is backbone independent by design and does not require end-to-end training. Extensive experiments are conducted on the AVA dataset, where our model demonstrates state-of-the-art results and consistent improvements over baselines built with different backbones. Code is publicly available at https://github.com/aimagelab/STAGE_action_detection.

citation-role summary

background 1

citation-polarity summary

fields

cs.CV 1

years

2025 1

verdicts

CONDITIONAL 1

roles

background 1

polarities

background 1

representative citing papers

Dual Guidance Semi-Supervised Action Detection

cs.CV · 2025-07-28 · conditional · novelty 5.0

Dual guidance, combining frame-level action classification with box-level prediction, improves pseudo-box selection for semi-supervised spatio-temporal action localization.

citing papers explorer

Showing 1 of 1 citing paper.

  • Dual Guidance Semi-Supervised Action Detection cs.CV · 2025-07-28 · conditional · none · ref 43 · internal anchor

    Dual guidance, combining frame-level action classification with box-level prediction, improves pseudo-box selection for semi-supervised spatio-temporal action localization.