MOOSE fuses frozen DINOv2 image features with optical flow features via cross-attention and causal aggregation, reporting 70.84% Kinetics-400 and 65.23% SSv2 top-1 accuracy, alongside interpretable attention maps.
Action recognition for surveillance applications using optic flow and svm
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.CV 1years
2025 1verdicts
REJECT 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
MOOSE: Pay Attention to Temporal Dynamics for Video Understanding via Optical Flows
MOOSE fuses frozen DINOv2 image features with optical flow features via cross-attention and causal aggregation, reporting 70.84% Kinetics-400 and 65.23% SSv2 top-1 accuracy, alongside interpretable attention maps.