ASEA is a skeleton-based interaction recognition network that selects active joints via temporal-weighted L2 norms and applies cross-attention between individuals, achieving state-of-the-art accuracy on NTU-26, SBU, and Kinetics-10.
Skeleton-based Action Recognition via Temporal-Channel Aggregation
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Skeleton-based action recognition methods are limited by the semantic extraction of spatio-temporal skeletal maps. However, current methods have difficulty in effectively combining features from both temporal and spatial graph dimensions and tend to be thick on one side and thin on the other. In this paper, we propose a Temporal-Channel Aggregation Graph Convolutional Networks (TCA-GCN) to learn spatial and temporal topologies dynamically and efficiently aggregate topological features in different temporal and channel dimensions for skeleton-based action recognition. We use the Temporal Aggregation module to learn temporal dimensional features and the Channel Aggregation module to efficiently combine spatial dynamic channel-wise topological features with temporal dynamic topological features. In addition, we extract multi-scale skeletal features on temporal modeling and fuse them with an attention mechanism. Extensive experiments show that our model results outperform state-of-the-art methods on the NTU RGB+D, NTU RGB+D 120, and NW-UCLA datasets.
citation-role summary
citation-polarity summary
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Learning Adaptive Node Selection with External Attention for Human Interaction Recognition
ASEA is a skeleton-based interaction recognition network that selects active joints via temporal-weighted L2 norms and applies cross-attention between individuals, achieving state-of-the-art accuracy on NTU-26, SBU, and Kinetics-10.