REVIEW 3 major objections 3 minor 133 references
Evolving Skeletons: Motion Dynamics in Action Recognition
T0 review · 3 major / 3 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read Motion-injected Taylor skeletons improve ST-GCN accuracy on NTU-60 and NTU-120 but slightly reduce Hyperformer accuracy, showing that the value of motion-enriched skeletons depends on the model architecture.
desk verdict A useful but under-powered evaluation study: the reported architecture-dependent effect of Taylor skeletons is plausible and honestly discussed, yet the missing error bars and confounded Taylor configuration keep the headline interaction from being fully established. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The Taylor-transformed skeleton sequence is the load-bearing input representation: joint positions are augmented with temporal derivatives—zeroth-order positions, first-order velocity from frame-to-frame differences, and second-order acceleration—combined in the displacement concept with one term, four frames per temporal block, and step size one. This object carries the comparison: both models receive identical static or motion-enriched inputs, so any accuracy difference is attributed to how the architecture uses the added dynamics.
What would settle it
Re-run the ST-GCN and Hyperformer evaluations on NTU-60 Cross-Subject with a different Taylor block size (for example, eight frames) or with first- and second-order terms included, and check whether Hyperformer's top-1 accuracy still falls below its original-skeleton baseline; if the drop shrinks, reverses, or moves to other classes, the architecture-dependence claim is tied to the untuned transform configuration rather than to Taylor skeletons as such.
Extended reading notes
Core claim
On the paper's own terms, the central finding is that injecting motion dynamics into skeleton sequences through the Taylor transform helps one architecture and hurts another. With ST-GCN, top-1 accuracy rises on all four benchmarks (NTU-60 X-Sub 81.5 to 83.5, X-View 88.3 to 89.4; NTU-120 X-Sub 70.7 to 74.1, X-Set 73.2 to 75.8). With Hyperformer, the same transformed input lowers accuracy on all four benchmarks (90.7 to 86.7, 95.1 to 92.1, 86.6 to 79.1, 88.0 to 81.9), and Hyperformer still beats ST-GCN in every configuration. The authors interpret this as evidence that Taylor skeletons supply motion-sensitive features that graph convolutions can exploit but that obscure spatial joint arrangement that the hypergraph-transformer still depends on.
Load-bearing premise
The comparison rests on the assumption that one fixed, untuned Taylor configuration is a fair test of motion-injected input for both ST-GCN and Hyperformer; if that configuration were tuned separately for each model, the reported gains and losses could change or disappear.
Editorial extensions
If this is right
- For ST-GCN, motion-injected skeletons act as implicit temporal feature engineering: adding them lifts accuracy by roughly 1 to 4 points across benchmarks without changing the network.
- For Hyperformer, Taylor skeletons consistently cost 3 to 7 points, meaning the model sacrifices spatial joint-arrangement information that its hypergraph self-attention still depends on.
- The benefit is action-specific: dynamic actions such as using a fan, wearing a shoe, and hopping improve, while spatially or fine-motor actions such as pointing, writing, and cutting with scissors degrade.
- Because Hyperformer beats ST-GCN even with the transformed input, higher-order hypergraph modeling is the stronger baseline, but it is not the right home for this motion representation.
- A hybrid that preserves spatial structure while adding motion derivatives is the paper's stated next step, and the confusion-matrix analysis gives a per-class map of where such a hybrid would help.
Reading between the lines
- A per-model search over the Taylor configuration (block size, step size, or including second-order terms) might shrink or reverse the Hyperformer loss, since the paper uses one fixed, untuned configuration for both architectures.
- A natural testable extension is a two-stream input that concatenates or fuses original and Taylor skeletons; the confusion matrices suggest this would recover Hyperformer's spatially reliant classes while keeping ST-GCN's motion gains.
- The same interaction may generalize beyond these two models: architectures with fixed anatomical topology benefit from explicit derivatives, while attention-based models that can already infer dynamics from raw positions may only be hurt by the loss of spatial detail.
- Per-class gains and losses are concentrated in actions with fine hand and finger motion, so a motion representation that encodes local joint-group derivatives rather than global displacement might serve both architectures.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper reports a comparative empirical study of two skeleton-based action recognition models, ST-GCN and Hyperformer, on the NTU-60 and NTU-120 benchmarks. For each model, the authors compare the original skeleton sequences against 'Taylor-transformed' skeletons, an input encoding borrowed from the prior Taylor Videos work (ref [95]) that is intended to emphasize motion dynamics. The headline finding is that Taylor-transformed skeletons improve ST-GCN accuracy on all four evaluated benchmarks (e.g., NTU-60 X-Sub 81.5% to 83.5%) but decrease Hyperformer accuracy (e.g., NTU-60 X-Sub 90.7% to 86.7%). The paper also provides per-class confusion matrices and tables of the action classes with the largest gains and losses under the transformation, and it discusses the trade-off between motion sensitivity and spatial detail.
Significance. If the reported interaction is reliable, the paper provides a useful empirical data point: the benefit of motion-injected skeleton representations is not uniform across architectures, and models with spatial-distance-based attention may be harmed by displacement-only inputs. The study's strengths are its breadth of benchmarks (four standard NTU splits), the detailed per-class analysis in Tables 2-9, and the honest disclosure that no hyperparameter search or denoising was performed. However, the paper's central contribution is purely empirical, and its credibility depends on the stability and confound-control of the reported accuracy numbers, both of which are currently insufficient for a journal-level claim about architecture-dependent input representations.
major comments (3)
- [Section 4.1, Table 1] All accuracy values are from single runs with no error bars, multiple seeds, or significance tests. The central interaction claim (Taylor skeletons help ST-GCN and harm Hyperformer) rests on differences ranging from 1.1 to 7.5 percentage points across the four NTU benchmarks. In skeleton action recognition, run-to-run variation of this magnitude is common, so the reader cannot tell whether the observed interaction is a stable property or a seed artifact. For an evaluation-only paper, the authors should report mean and standard deviation over at least three seeds, or apply a paired statistical test over the test set (e.g., bootstrap per-sample accuracies) to support the claim.
- [Section 4.1, Models] The Taylor encoding is fixed to 'the displacement concept with a single term, four frames per temporal block and a step size of one' with 'no hyperparameter search,' and each model uses its own standard training recipe. This confounds the input representation with the model's ability to consume that representation. Hyperformer's attention mechanism explicitly includes joint-distance attention computed on input coordinates (Section 3.3, Figure 2); a displacement-only representation removes the static spatial structure that this attention was designed to process. The 4-7 point drops for Hyperformer may therefore reflect an encoding/model mismatch rather than a general incompatibility between motion-enhanced skeletons and hypergraph-transformer models. The paper's own explanation in Section 4.2 ('they lack detailed spatial information... which the Hyperformer may still rely on') concedes this confound. To support the architecture-dependent conclusion, the authors should test at least one Taylor variant that retains static pose information, or tune the Taylor hyperparameters for Hyperformer, or explicitly re-frame the claim as being about this particular encoding and pipeline.
- [Sections 3.2 and 4.1] The description of the Taylor transform is internally inconsistent, and the exact input representation used in the experiments is not fully specified. Section 3.2 describes Taylor-transformed skeletons as combining zeroth-, first-, and second-order temporal derivatives, while Section 4.1 states that the experiments use 'the displacement concept with a single term' without defining that concept. The reader cannot determine whether the input fed to Hyperformer contains static joint coordinates, only velocities, or some combination. Since the paper's central interpretation hinges on the loss of spatial information in the Taylor input, the exact formula (and whether static positions are present) must be stated unambiguously, ideally in Section 4.1 or in a short appendix.
minor comments (3)
- [Table 1] Top-5 accuracy is reported only for the Taylor-transformed runs; the original-skeleton rows show '–' in the Top-5 columns. This prevents the reader from comparing top-5 performance between the two input conditions, which would be informative given the large top-1 differences for Hyperformer.
- [Contributions and Section 3.1] The paper frames the comparison as 'skeletal graphs versus hypergraphs,' but only one graph model (ST-GCN) and one hypergraph model (Hyperformer) are evaluated. The conclusions about representation families are model-specific and should be explicitly qualified as such, both in the contribution list and in the conclusion.
- [Figures 3 and 4] The captions state that predictions below 5% are filtered out for clarity, but several rows of the displayed matrices do not sum to approximately 100% even after accounting for this filtering (e.g., Figure 4(a), 'drink water' row). The paper should clarify whether the remaining discrepancy is due to rounding, filtering of additional values, or entries not shown, so that readers can trust the numerical values in the figures.
Circularity Check
No significant circularity: the paper is an empirical comparison with measured accuracies, not a derivation.
full rationale
The paper's central claims are empirical: Taylor-transformed skeletons improve ST-GCN and decrease Hyperformer on NTU-60 and NTU-120. These are measured recognition accuracies reported in Table 1, not quantities derived from the input transformation. The Taylor transformation configuration is taken from the authors' prior work [95], which is a self-citation, but it is not load-bearing in a circular sense: the paper does not claim the improvement follows by definition, and the reported results include a negative effect for Hyperformer that is not forced by the citation. The fixed configuration ('displacement concept with a single term, four frames per temporal block and a step size of one', with 'no hyperparameter search') may limit external validity or fair comparison across architectures, but that is a correctness or generalization concern, not a circularity one. No fitted parameter is renamed as a prediction, no uniqueness theorem is imported, and no result is equivalent to its input by construction. The paper is self-contained as an evaluation against standard benchmarks, so the circularity score is 0.
Assumptions & free parameters
free parameters (1)
- Taylor transformation hyperparameters =
single displacement term, 4 frames per temporal block, step size 1
assumptions (3)
- domain assumption NTU-60 and NTU-120 benchmark protocols provide a valid measure of action recognition performance.
- domain assumption The Taylor transformation as implemented from [95] correctly captures motion dynamics and is compatible with both ST-GCN and Hyperformer input layers.
- ad hoc to paper Single-run accuracy values are treated as deterministic estimates of model performance.
Cite this review
Pith. "Pith review of Evolving Skeletons: Motion Dynamics in Action Recognition." pith.science (2026). https://pith.science/paper/N3KX4RES
@misc{pith2026250102593,
author = {Pith},
title = {Pith review of: Evolving Skeletons: Motion Dynamics in Action Recognition},
year = {2026},
howpublished = {\url{https://pith.science/paper/N3KX4RES}},
note = {Machine review of arXiv:2501.02593}
}
read the original abstract
Skeleton-based action recognition has gained significant attention for its ability to efficiently represent spatiotemporal information in a lightweight format. Most existing approaches use graph-based models to process skeleton sequences, where each pose is represented as a skeletal graph structured around human physical connectivity. Among these, the Spatiotemporal Graph Convolutional Network (ST-GCN) has become a widely used framework. Alternatively, hypergraph-based models, such as the Hyperformer, capture higher-order correlations, offering a more expressive representation of complex joint interactions. A recent advancement, termed Taylor Videos, introduces motion-enhanced skeleton sequences by embedding motion concepts, providing a fresh perspective on interpreting human actions in skeleton-based action recognition. In this paper, we conduct a comprehensive evaluation of both traditional skeleton sequences and Taylor-transformed skeletons using ST-GCN and Hyperformer models on the NTU-60 and NTU-120 datasets. We compare skeletal graph and hypergraph representations, analyzing static poses against motion-injected poses. Our findings highlight the strengths and limitations of Taylor-transformed skeletons, demonstrating their potential to enhance motion dynamics while exposing current challenges in fully using their benefits. This study underscores the need for innovative skeletal modelling techniques to effectively handle motion-rich data and advance the field of action recognition.
Figures
Figures from the paper (10 more)
Reference graph
Works this paper leans on
-
[95]
Lei Wang, Xiuyuan Yuan, Tom Gedeon, and Liang Zheng. 2024. Taylor videos for action recognition. arXiv preprint arXiv:2402.03019 (2024)
arXiv 2024
-
[1]
Tasweer Ahmad, Lianwen Jin, Xin Zhang, Songxuan Lai, Guozhi Tang, and Luojun Lin. 2021. Graph convolutional neural network for human action recog- nition: A comprehensive survey. IEEE Transactions on Artificial Intelligence 2, 2 (2021), 128–145
2021
-
[2]
Tamam Alsarhan, Syed Sadaf Ali, Ayoub Alsarhan, Iyyakutti Iyappan Ganapathi, and Naoufel Werghi. 2024. Human Action Recognition with Multi-Level Granu- larity and Pair-Wise Hyper GCN. In 2024 IEEE 18th International Conference on Automatic Face and Gesture Recognition (FG) . IEEE, 1–10
2024
-
[3]
Tamam Alsarhan, Syed Sadaf Ali, Iyyakutti Iyappan Ganapathi, Ahmad Ali, and Naoufel Werghi. 2024. PH-GCN: Boosting Human Action Recognition through Multi-Level Granularity with Pair-wise Hyper GCN. IEEE Access (2024)
2024
-
[4]
Vu Ho Tran Anh and Thi-Oanh Nguyen. 2024. Enhanced Topology Repre- sentation Learning for Skeleton-Based Human Action Recognition. Procedia Computer Science 246 (2024), 3093–3102
2024
-
[5]
Hamza Bouzid and Lahoucine Ballihi. 2024. SpATr: MoCap 3D human action recognition based on spiral auto-encoder and transformer network. Computer Vision and Image Understanding 241 (2024), 103974
2024
-
[6]
Joao Carreira and Andrew Zisserman. 2017. Quo vadis, action recognition? a new model and the kinetics dataset. In proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . 6299–6308
2017
-
[7]
Huilin Chen, Lei Wang, Yifan Chen, Tom Gedeon, and Piotr Koniusz. 2024. When spatial meets temporal in action recognition. arXiv preprint arXiv:2411.15284 (2024)
work page Pith review arXiv 2024
Show all 133 references
-
[8]
Qixiang Chen, Lei Wang, Piotr Koniusz, and Tom Gedeon. [n. d.]. Motion meets attention: Video motion prompts. In The 16th Asian Conference on Machine Learning (Conference Track)
-
[9]
Wenshuo Chen, Hongru Xiao, Erhang Zhang, Lijie Hu, Lei Wang, Mengyuan Liu, and Chen Chen. 2024. SATO: Stable Text-to-Motion Framework. In Proceedings of the 32nd ACM International Conference on Multimedia . 6989–6997
2024
-
[10]
Xi Chen and Markus Koskela. 2015. Skeleton-based action recognition with extreme learning machines. Neurocomputing 149 (2015), 387–396
2015
-
[11]
Yanjun Chen, Ying Li, Chongyang Zhang, Hao Zhou, Yan Luo, and Chuanping Hu. 2022. Informed Patch Enhanced HyperGCN for skeleton-based action recognition. Information Processing & Management 59, 4 (2022), 102950
2022
-
[12]
Zefang Chen, Yang Gao, and Qiuyan Yan. 2024. A Key Skeleton Points Guided Classroom Action Recognition Method Based on Multimodal Symmetry Fusion. IEEE Access (2024)
2024
-
[13]
Zengzhao Chen, Wenkai Huang, Hai Liu, Zhuo Wang, Yuqun Wen, and Sheng- ming Wang. 2024. ST-TGR: Spatio-Temporal Representation Learning for Skeleton-Based Teaching Gesture Recognition. Sensors 24, 8 (2024), 2589
2024
-
[14]
Yan Cheng, Chengxing Fang, Jiawen Huang, et al. [n. d.]. Spatiotemporal Action Detection Based on Fine-Grained. Chengxing and Huang, Jiawen, Spatiotemporal Action Detection Based on Fine-Grained ([n. d.])
-
[15]
Sangwoo Cho, Muhammad Maqbool, Fei Liu, and Hassan Foroosh. 2020. Self- attention network for skeleton-based human action recognition. In Proceedings of the IEEE/CVF winter conference on applications of computer vision . 635–644
2020
-
[16]
Haigang Deng, Guocheng Lin, Chengwei Li, and Chuanxu Wang. 2024. Research on decoupled adaptive graph convolution networks based on skeleton data for action recognition. Pattern Analysis and Applications 27, 4 (2024), 118
2024
-
[17]
Chongyang Ding, Shan Wen, Wenwen Ding, Kai Liu, and Evgeny Belyaev
-
[18]
Dexuan Ding, Lei Wang, Liyun Zhu, Tom Gedeon, and Piotr Koniusz. 2024. Lego: Learnable expansion of graph operators for multi-modal feature fusion. arXiv preprint arXiv:2410.01506 (2024)
2024 arXiv
-
[19]
Xi Ding and Lei Wang. 2024. Do language models understand time? arXiv preprint arXiv:2412.13845 (2024)
2024 arXiv
-
[20]
Xi Ding and Lei Wang. 2024. Quo Vadis, Anomaly Detection? LLMs and VLMs in the Spotlight. arXiv preprint arXiv:2412.18298 (2024)
2024 arXiv
-
[21]
Jeonghyeok Do and Munchurl Kim. 2025. Skateformer: skeletal-temporal trans- former for human action recognition. In European Conference on Computer Vision. Springer, 401–420
2025
-
[22]
Yong Du, Yun Fu, and Liang Wang. 2015. Skeleton based action recognition with convolutional neural network. In 2015 3rd IAPR Asian conference on pattern recognition (ACPR). IEEE, 579–583
2015
-
[23]
Haodong Duan, Jiaqi Wang, Kai Chen, and Dahua Lin. 2022. Dg-stgcn: Dynamic spatial-temporal modeling for skeleton-based action recognition. arXiv preprint arXiv:2210.05895 (2022)
2022 arXiv
-
[24]
Haodong Duan, Jiaqi Wang, Kai Chen, and Dahua Lin. 2022. Pyskl: Towards good practices for skeleton action recognition. In Proceedings of the 30th ACM International Conference on Multimedia . 7351–7354
2022
-
[25]
Michael Duhme, Raphael Memmesheimer, and Dietrich Paulus. 2021. Fusion- gcn: Multimodal action recognition using graph convolutional networks. In DAGM German conference on pattern recognition . Springer, 265–281
2021
-
[26]
Zheng Fang, Xiongwei Zhang, Tieyong Cao, Yunfei Zheng, and Meng Sun. 2022. A new adjacency matrix configuration in GCN-based models for skeleton-based action recognition. arXiv preprint arXiv:2206.14344 (2022)
2022 arXiv
-
[27]
Mesafint Fanuel, Xiaohong Yuan, Hyung Nam Kim, Letu Qingge, and Kaushik Roy. 2021. A survey on skeleton-based activity recognition using graph convo- lutional networks (GCN). In 2021 12th International Symposium on Image and Signal Processing and Analysis (ISPA). IEEE, 177–182
2021
-
[28]
Liqi Feng, Yaqin Zhao, Wenxuan Zhao, and Jiaxi Tang. 2022. A comparative review of graph convolutional networks for human skeleton-based action recog- nition. Artificial Intelligence Review (2022), 1–31
2022
-
[29]
Yifan Feng, Haoxuan You, Zizhao Zhang, Rongrong Ji, and Yue Gao. 2019. Hypergraph neural networks. In Proceedings of the AAAI conference on artificial intelligence, Vol. 33. 3558–3565
2019
-
[30]
Benjamin Filtjens, Bart Vanrumste, and Peter Slaets. 2022. Skeleton-based action segmentation with multi-stage spatial-temporal graph convolutional neural networks. IEEE Transactions on Emerging Topics in Computing 12, 1 (2022), 202–212
2022
-
[31]
Annalisa Franco, Antonio Magnani, and Dario Maio. 2020. A multimodal ap- proach for human activity recognition based on skeleton and RGB data. Pattern Recognition Letters 131 (2020), 293–299
2020
-
[32]
Bing-Kun Gao, Le Dong, Hong-Bo Bi, and Yun-Ze Bi. 2022. Focus on temporal graph convolutional networks with unified attention for skeleton-based action recognition. Applied Intelligence 52, 5 (2022), 5608–5616
2022
-
[33]
Yue Gao, Jiaxuan Lu, Siqi Li, Yipeng Li, and Shaoyi Du. 2024. Hypergraph-Based Multi-View Action Recognition Using Event Cameras. IEEE Transactions on Pattern Analysis and Machine Intelligence (2024)
2024
-
[34]
Yue Gao, Meng Wang, Dacheng Tao, Rongrong Ji, and Qionghai Dai. 2012. 3-D object retrieval and recognition with hypergraph analysis. IEEE transactions on image processing 21, 9 (2012), 4290–4303
2012
-
[35]
Wen Ge, Guanyi Mou, Emmanuel O Agu, and Kyumin Lee. 2024. Deep Het- erogeneous Contrastive Hyper-Graph Learning for In-the-Wild Context-Aware Human Activity Recognition. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 7, 4 (2024), 1–23
2024
-
[36]
Rui Hang and MinXian Li. 2022. Spatial-temporal adaptive graph convolutional network for skeleton-based action recognition. In Proceedings of the Asian Conference on Computer Vision . 1265–1281
2022
-
[37]
Xiaoke Hao, Jie Li, Yingchun Guo, Tao Jiang, and Ming Yu. 2021. Hypergraph Neural Network for Skeleton-Based Action Recognition. IEEE Transactions on Image Processing 30 (2021), 2263–2275. https://doi.org/10.1109/TIP.2021.3051495
2021
-
[38]
Lianyu Hu, Shenglan Liu, and Wei Feng. 2023. Skeleton-based action recog- nition with local dynamic spatial–temporal aggregation. Expert Systems with Applications 232 (2023), 120683
2023
-
[39]
Hongbo Huang, Longfei Xu, Yaolin Zheng, and Xiaoxu Yan. 2024. MAFormer: A cross-channel spatio-temporal feature aggregation method for human action recognition. AI Communications Preprint (2024), 1–15
2024
-
[40]
Junhao Huang, Ziming Wang, Jian Peng, and Feihu Huang. 2023. Feature recon- struction graph convolutional network for skeleton-based action recognition. Engineering Applications of Artificial Intelligence 126 (2023), 106855
2023
-
[41]
Yuchi Huang, Qingshan Liu, and Dimitris Metaxas. 2009. ] Video object seg- mentation by hypergraph cut. In 2009 IEEE conference on computer vision and pattern recognition. IEEE, 1738–1745
2009
-
[42]
Zengxi Huang, Yusong Qin, Xiaobing Lin, Tianlin Liu, Zhenhua Feng, and Yiguang Liu. 2022. Motion-driven spatial and temporal adaptive high-resolution graph convolutional networks for skeleton-based action recognition. IEEE Transactions on Circuits and Systems for Video Technol...
2022
-
[43]
Shengqin Jiang, Haokui Zhang, Yuankai Qi, and Qingshan Liu. 2024. Spatial- temporal interleaved network for efficient action recognition. IEEE Transactions on Industrial Informatics (2024)
2024
-
[44]
Misha Karim, Shah Khalid, Aliya Aleryani, Jawad Khan, Irfan Ullah, and Zafar Ali. 2024. Human action recognition systems: A review of the trends and state-of-the-art. IEEE Access (2024)
2024
-
[45]
Manjin Kim, Paul Hongsuck Seo, Cordelia Schmid, and Minsu Cho. 2024. Learn- ing correlation structures for vision transformers. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 18941–18951
2024
-
[46]
Sunwoo Kim, Soo Yong Lee, Yue Gao, Alessia Antelmi, Mirko Polato, and Kijung Shin. 2024. A survey on hypergraph neural networks: An in-depth and step-by- step guide. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining . 6534–6544
2024
-
[47]
Piotr Koniusz, Lei Wang, and Anoop Cherian. 2021. Tensor representations for action recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence 44, 2 (2021), 648–665
2021
-
[48]
Ce Li, Chunyu Xie, Baochang Zhang, Jungong Han, Xiantong Zhen, and Jie Chen
-
[49]
Maosen Li, Siheng Chen, Xu Chen, Ya Zhang, Yanfeng Wang, and Qi Tian. 2019. Actional-structural graph convolutional networks for skeleton-based action recognition. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 3595–3603. WWW Companion ’2...
2019
-
[50]
Tianchen Li, Pei Geng, Xuequan Lu, Wanqing Li, and Lei Lyu. 2024. Skeleton- based action recognition through attention guided heterogeneous graph neural network. Knowledge-Based Systems (2024), 112868
2024
-
[51]
Xiaolong Li, Yang Dong, Yunfei Yi, Zhixun Liang, and Shuqi Yan. 2024. Hyper- graph Neural Network for Multimodal Depression Recognition. Electronics 13, 22 (2024), 4544
2024
-
[52]
Xuanfeng Li, Jian Lu, Jian Zhou, Wei Liu, and Kaibing Zhang. 2024. Multi- temporal scale aggregation refinement graph convolutional network for skeleton-based action recognition. Computer Animation and Virtual Worlds 35, 1 (2024), e2221
2024
-
[53]
Fenglin Liu, Chenyu Wang, Zhiqiang Tian, Shaoyi Du, and Wei Zeng. 2025. Advancing skeleton-based human behavior recognition: multi-stream fusion spatiotemporal graph convolutional networks. Complex & Intelligent Systems 11, 1 (2025), 94
2025
-
[54]
Guiyu Liu, Jiuchao Qian, Fei Wen, Xiaoguang Zhu, Rendong Ying, and Peilin Liu. 2019. Action recognition based on 3d skeleton and rgb frame fusion. In 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . IEEE, 258–264
2019
-
[55]
Jinfu Liu, Chen Chen, and Mengyuan Liu. 2024. Multi-modality co-learning for efficient skeleton-based action recognition. In Proceedings of the 32nd ACM International Conference on Multimedia . 4909–4918
2024
-
[56]
Jun Liu, Amir Shahroudy, Mauricio Perez, Gang Wang, Ling-Yu Duan, and Alex C Kot. 2019. Ntu rgb+ d 120: A large-scale benchmark for 3d human activity understanding. IEEE transactions on pattern analysis and machine intelligence 42, 10 (2019), 2684–2701
2019
-
[57]
Shengyuan Liu, Pei Lv, Yuzhen Zhang, Jie Fu, Junjin Cheng, Wanqing Li, Bing Zhou, and Mingliang Xu. 2020. Semi-Dynamic Hypergraph Neural Network for 3D Pose Estimation. In Proceedings of the Twenty-Ninth International Joint Conference on Artificial Intelligence, IJCAI-20 , Chr...
2020 doi
-
[58]
Shaocan Liu, Xingtao Wang, Ruiqin Xiong, and Xiaopeng Fan. 2024. GCN-based Multi-modality Fusion Network for Action Recognition. IEEE Transactions on Multimedia (2024)
2024
-
[59]
Yi Liu, Ruyi Liu, Yuzhi Hu, Mengyao Wu, Wentian Xin, Qiguang Miao, Shuai Wu, and Long Li. [n. d.]. A Systematic Review of Skeleton-Based Action Recognition: Recent Advances, Challenges, and Future Directions. Challenges, and Future Directions ([n. d.])
-
[60]
Yanan Liu, Hao Zhang, Yanqiu Li, Kangjian He, and Dan Xu. 2023. Skeleton- based human action recognition via large-kernel attention graph convolutional network. IEEE Transactions on Visualization and Computer Graphics 29, 5 (2023), 2575–2585
2023
-
[61]
Ziyu Liu, Hongwen Zhang, Zhenghao Chen, Zhiyong Wang, and Wanli Ouyang
-
[62]
Mayank Lovanshi and Vivek Tiwari. 2024. Human skeleton pose and spatio- temporal feature-based activity recognition using ST-GCN. Multimedia Tools and Applications 83, 5 (2024), 12705–12730
2024
-
[63]
Nan Ma, Zhixuan Wu, Yifan Feng, Cheng Wang, and Yue Gao. 2024. Multi- View Time-Series Hypergraph Neural Network for Action Recognition. IEEE Transactions on Image Processing (2024)
2024
-
[64]
Woomin Myung, Nan Su, Jing-Hao Xue, and Guijin Wang. 2024. DeGCN: De- formable Graph Convolutional Networks for Skeleton-Based Action Recognition. IEEE Transactions on Image Processing 33 (2024), 2477–2490
2024
-
[65]
Zhenyue Qin, Yang Liu, Pan Ji, Dongwoo Kim, Lei Wang, Saeed Anwar, and Tom Gedeon. 2022. Fusing higher-order features in graph neural networks for skeleton-based action recognition. IEEE Transactions on Neural Networks and Learning Systems 35, 4 (2022), 4783–4797
2022
-
[66]
Mrugendrasinh Rahevar, Amit Ganatra, Tanzila Saba, Amjad Rehman, and Saeed Ali Bahaj. 2023. Spatial–temporal dynamic graph attention network for skeleton-based action recognition. IEEE Access 11 (2023), 21546–21553
2023
-
[67]
Arjun Raj, Lei Wang, and Tom Gedeon. 2024. Tracknetv4: Enhancing fast sports object tracking with motion attention maps. arXiv preprint arXiv:2409.14543 (2024)
2024 arXiv
-
[68]
Bin Ren, Mengyuan Liu, Runwei Ding, and Hong Liu. 2024. A survey on 3d skeleton-based action recognition using learning method. Cyborg and Bionic Systems 5 (2024), 0100
2024
-
[69]
Zilaing Ren, Li Luo, Yong Qin, Xiangyang Gao, and Qieshi Zhang. [n. d.]. Skeleton-Guided and Supervised Learning of Hybrid Network for Multi-Modal Action Recognition. A vailable at SSRN 4970121([n. d.])
-
[70]
ZiLiang Ren, QieShi Zhang, Qin Cheng, ZhenYu Xu, Shuai Yuan, and Delin Luo. 2024. Segment differential aggregation representation and supervised compensation learning of ConvNets for human action recognition. Science China Technological Sciences 67, 1 (2024), 197–208
2024
-
[71]
Amir Shahroudy, Jun Liu, Tian-Tsong Ng, and Gang Wang. 2016. Ntu rgb+ d: A large scale dataset for 3d human activity analysis. In Proceedings of the IEEE conference on computer vision and pattern recognition . 1010–1019
2016
-
[72]
Muhammad Bilal Shaikh, Douglas Chai, Syed Muhammad Shamsul Islam, and Naveed Akhtar. 2024. From CNNs to Transformers in Multimodal Human Action Recognition: A Survey. ACM Transactions on Multimedia Computing, Communications and Applications (2024)
2024
-
[73]
Lei Shi, Yifan Zhang, Jian Cheng, and Hanqing Lu. 2019. Skeleton-based action recognition with directed graph neural networks. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition . 7912–7921
2019
-
[74]
Lei Shi, Yifan Zhang, Jian Cheng, and Hanqing Lu. 2019. Two-stream adaptive graph convolutional networks for skeleton-based action recognition. In Pro- ceedings of the IEEE/CVF conference on computer vision and pattern recognition . 12026–12035
2019
-
[75]
Chenyang Si, Wentao Chen, Wei Wang, Liang Wang, and Tieniu Tan. 2019. An attention enhanced graph convolutional lstm network for skeleton-based action recognition. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 1227–1236
2019
-
[76]
Yi-Fan Song, Zhang Zhang, Caifeng Shan, and Liang Wang. 2022. Construct- ing stronger and faster baselines for skeleton-based action recognition. IEEE transactions on pattern analysis and machine intelligence 45, 2 (2022), 1474–1488
2022
-
[77]
Yaohui Sun, Weiyao Xu, Xiaoyi Yu, and Ju Gao. 2024. VT-BPAN: vision transformer-based bilinear pooling and attention network fusion of RGB and skeleton features for human action recognition. Multimedia Tools and Applica- tions 83, 29 (2024), 73391–73405
2024
-
[78]
Haoyu Tian, Xin Ma, Xiang Li, and Yibin Li. 2023. Skeleton-based action recognition with select-assemble-normalize graph convolutional networks.IEEE Transactions on Multimedia (2023)
2023
-
[79]
Xiaoyan Tian, Ye Jin, Zhao Zhang, Peng Liu, and Xianglong Tang. 2024. Spatial- temporal graph transformer network for skeleton-based temporal action seg- mentation. Multimedia Tools and Applications 83, 15 (2024), 44273–44297
2024
-
[80]
A Vaswani. 2017. Attention is all you need. Advances in Neural Information Processing Systems (2017)
2017
-
[81]
Cheng Wang, Nan Ma, Zhixuan Wu, Jin Zhang, and Yongqiang Yao. 2022. Survey of Hypergraph Neural Networks and Its Application to Action Recognition. In CAAI International Conference on Artificial Intelligence . Springer, 387–398
2022
-
[82]
Lei Wang. 2017. Analysis and Evaluation of Kinect-based Action Recognition Algorithms. Master’s thesis. School of the Computer Science and Software Engineering, The University of Western Australia
2017
-
[83]
Lei Wang. 2023. Robust human action modelling . Ph. D. Dissertation. The Australian National University (Australia)
2023
-
[84]
Lei Wang, Du Q Huynh, and Piotr Koniusz. 2019. A comparative review of recent kinect-based action recognition algorithms. IEEE Transactions on Image Processing 29 (2019), 15–28
2019
-
[85]
Lei Wang, Du Q Huynh, and Moussa Reda Mansour. 2019. Loss switching fusion with similarity search for video classification. In 2019 IEEE international conference on image processing (ICIP) . IEEE, 974–978
2019
-
[86]
Lei Wang and Piotr Koniusz. 2021. Self-supervising action recognition by statistical moment and subspace descriptors. In Proceedings of the 29th ACM international conference on multimedia . 4324–4333
2021
-
[87]
Lei Wang and Piotr Koniusz. 2022. Temporal-viewpoint transportation plan for skeletal few-shot action recognition. In Proceedings of the Asian Conference on Computer Vision. 4176–4193
2022
-
[88]
Lei Wang and Piotr Koniusz. 2022. Uncertainty-dtw for time series and se- quences. In European Conference on Computer Vision . Springer, 176–195
2022
-
[89]
Lei Wang and Piotr Koniusz. 2023. 3mformer: Multi-order multi-mode trans- former for skeletal action recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 5620–5631
2023
-
[90]
Lei Wang and Piotr Koniusz. 2024. Flow dynamics correction for action recog- nition. In ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 3795–3799
2024
-
[91]
Lei Wang, Piotr Koniusz, and Du Q Huynh. 2019. Hallucinating idt descriptors and i3d optical flow features for action recognition with cnns. In Proceedings of the IEEE/CVF international conference on computer vision . 8698–8708
2019
-
[92]
Lei Wang, Jun Liu, and Piotr Koniusz. 2021. 3D Skeleton-based Few-shot Action Recognition with JEANIE is not so Naïve.arXiv preprint arXiv:2112.12668 (2021)
2021 arXiv
-
[93]
Lei Wang, Jun Liu, Liang Zheng, Tom Gedeon, and Piotr Koniusz. 2024. Meet JEANIE: a Similarity Measure for 3D Skeleton Sequences via Temporal- Viewpoint Alignment. International Journal of Computer Vision (2024), 1–32
2024
-
[94]
Lei Wang, Ke Sun, and Piotr Koniusz. 2024. High-order tensor pooling with at- tention for action recognition. InICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 3885–3889
2024
-
[96]
Xinghan Wang and Yadong Mu. 2024. Localized Linear Temporal Dynamics for Self-supervised Skeleton Action Recognition. IEEE Transactions on Multimedia (2024)
2024
-
[97]
Ying Wang, Lu Zhang, Jingliang Peng, and Na Lv. [n. d.]. Motion-Centric Retrieval of 3d Human Skeleton and 2d Human Image Sequences. A vailable at SSRN 4925586 ([n. d.]). Evolving Skeletons: Motion Dynamics in Action Recognition WWW Companion ’25, April 28-May 2, 2025, Sydney,...
2025
-
[98]
Zhengjie Wang, Mingjing Ma, Xiaoxue Feng, Xue Li, Fei Liu, Yinjing Guo, and Da Chen. 2022. Skeleton-based human pose recognition using channel state information: A survey. Sensors 22, 22 (2022), 8738
2022
-
[99]
Xu Weiyao, Wu Muqing, Zhao Min, and Xia Ting. 2021. Fusion of skeleton and RGB features for RGB-D human action recognition. IEEE sensors journal 21, 17 (2021), 19157–19164
2021
-
[100]
Qianhan Wu, Qian Huang, and Xing Li. 2023. Multimodal human action recog- nition based on spatio-temporal action representation recognition model. Mul- timedia Tools and Applications 82, 11 (2023), 16409–16430
2023
-
[101]
Weiwei Wu, Fengbin Tu, Mengqi Niu, Zhiheng Yue, Leibo Liu, Shaojun Wei, Xiangyu Li, Yang Hu, and Shouyi Yin. 2023. STAR: An STGCN ARchitecture for Skeleton-Based Human Action Recognition. IEEE Transactions on Circuits and Systems I: Regular Papers 70, 6 (2023), 2370–2383
2023
-
[102]
Zhize Wu, Yue Ding, Long Wan, Teng Li, and Fudong Nian. 2025. Local and global self-attention enhanced graph convolutional network for skeleton-based action recognition. Pattern Recognition 159 (2025), 111106
2025
-
[103]
Zhixuan Wu, Nan Ma, Cheng Wang, Cheng Xu, Genbao Xu, and Mingxing Li
-
[104]
Limin Xia and Xin Wen. 2024. Multi-stream network with key frame sampling for human action recognition. The Journal of Supercomputing (2024), 1–31
2024
-
[105]
Zhenggui Xie, Gengzhong Zheng, Liming Miao, and Wei Huang. 2023. STGL- GCN: Spatial–temporal mixing of global and local self-attention graph convolu- tional networks for human action recognition. IEEE Access 11 (2023), 16526– 16532
2023
-
[106]
Wentian Xin, Ruyi Liu, Yi Liu, Yu Chen, Wenxin Yu, and Qiguang Miao. 2023. Transformer for skeleton-based action recognition: A review of recent advances. Neurocomputing 537 (2023), 164–186
2023
-
[107]
Yuling Xing, Jia Zhu, Yu Li, Jin Huang, and Jinlong Song. 2023. An improved spatial temporal graph convolutional network for robust skeleton-based action recognition. Applied Intelligence 53, 4 (2023), 4592–4608
2023
-
[108]
Zhuoyan Xu and Jingke Xu. 2024. GR-Former: Graph-reinforcement transformer for skeleton-based driver action recognition. IET Computer Vision 18, 7 (2024), 982–991
2024
-
[109]
Naganand Yadati, Madhav Nimishakavi, Prateek Yadav, Vikram Nitin, Anand Louis, and Partha Talukdar. 2019. Hypergcn: A new method for training graph convolutional networks on hypergraphs. Advances in neural information pro- cessing systems 32 (2019)
2019
-
[110]
Sijie Yan, Yuanjun Xiong, and Dahua Lin. 2018. Spatial temporal graph convo- lutional networks for skeleton-based action recognition. In Proceedings of the AAAI conference on artificial intelligence , Vol. 32
2018
-
[111]
Fan Yang, Yang Wu, Sakriani Sakti, and Satoshi Nakamura. 2019. Make skeleton- based action recognition model smaller, faster and better. In Proceedings of the 1st ACM International Conference on Multimedia in Asia . 1–6
2019
-
[112]
Hao Yang, Dan Yan, Li Zhang, Yunda Sun, Dong Li, and Stephen J Maybank. 2021. Feedback graph convolutional network for skeleton-based action recognition. IEEE Transactions on Image Processing 31 (2021), 164–175
2021
-
[113]
Shiqiang YANG, Zhuo LI, Jinhua WANG, Duo HE, Qi LI, and Dexin LI. 2023. ST-GCN human action recognition based on new partition strategy. Computer Integrated Manufacturing System 29, 12 (2023), 4040
2023
-
[114]
Yijie Yang, Jinlu Zhang, Jiaxu Zhang, and Zhigang Tu. 2024. Expressive Key- points for Skeleton-based Action Recognition via Skeleton Transformation. arXiv preprint arXiv:2406.18011 (2024)
2024 arXiv
-
[115]
Bruce XB Yu, Yan Liu, and Keith CC Chan. 2020. Skeleton focused human activity recognition in rgb video. arXiv preprint arXiv:2004.13979 (2020)
2020 arXiv
-
[116]
Rujing Yue, Zhiqiang Tian, and Shaoyi Du. 2022. Action recognition based on RGB and skeleton data sets: A survey. Neurocomputing 512 (2022), 287–306
2022
-
[117]
Jiaxu Zhang, Gaoxiang Ye, Zhigang Tu, Yongtao Qin, Qianqing Qin, Jinlu Zhang, and Jun Liu. 2022. A spatial attentive and temporal dilated (SATD) GCN for skeleton-based action recognition. CAAI Transactions on Intelligence Technology 7, 1 (2022), 46–55
2022
-
[118]
Xiantong Zhen, Ling Shao, Dacheng Tao, and Xuelong Li. 2013. Embedding motion and structure features for action recognition. IEEE Transactions on Circuits and Systems for Video Technology 23, 7 (2013), 1182–1190
2013
-
[119]
Dengyong Zhou, Jiayuan Huang, and Bernhard Schölkopf. 2006. Learning with hypergraphs: Clustering, classification, and embedding. Advances in neural information processing systems 19 (2006)
2006
-
[120]
Yuxuan Zhou, Zhi-Qi Cheng, Chao Li, Yanwen Fang, Yifeng Geng, Xuansong Xie, and Margret Keuper. 2022. Hypergraph transformer for skeleton-based action recognition. arXiv preprint arXiv:2211.09590 (2022)
2022 arXiv
-
[121]
Yuxuan Zhou, Xudong Yan, Zhi-Qi Cheng, Yan Yan, Qi Dai, and Xian-Sheng Hua. 2024. BlockGCN: Redefine Topology Awareness for Skeleton-Based Action Recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2049–2058
2024
-
[122]
Anlei Zhu, Yinghui Wang, Jinlong Yang, Tao Yan, Haomiao Ma, and Wei Li
-
[123]
Liyun Zhu, Lei Wang, Arjun Raj, Tom Gedeon, and Chen Chen. [n. d.]. Advanc- ing Video Anomaly Detection: A Concise Review and a New Dataset. In The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track
-
[124]
Xingyu Zhu, Xiangbo Shu, and Jinhui Tang. 2024. Motion-Aware Mask Feature Reconstruction for Skeleton-Based Action Recognition. IEEE Transactions on Circuits and Systems for Video Technology (2024)
2024
-
[125]
Xiaolin Zhu, Dongli Wang, Jianxun Li, Rui Su, Qin Wan, and Yan Zhou. 2024. Dynamical Attention Hypergraph Convolutional Network for Group Activity Recognition. IEEE Transactions on Neural Networks and Learning Systems (2024)
2024
-
[126]
Yiran Zhu, Guangji Huang, Xing Xu, Yanli Ji, and Fumin Shen. 2022. Selective hypergraph convolutional networks for skeleton-based action recognition. In Proceedings of the 2022 international conference on multimedia retrieval. 518–526
2022
-
[127]
IEEE Transactions on Circuits and Systems for Video Technology(2024)
YOWOv3: A Lightweight Spatio-Temporal Joint Network for Video Action Detection. IEEE Transactions on Circuits and Systems for Video Technology(2024)
2024
-
[128]
take off a shoe
Tianming Zhuang, Zhen Qin, Yi Ding, Zhiguang Qin, Ji Geng, Yi Liu, and Kim-Kwang Raymond Choo. 2024. DSDC-GCN: Decoupled Static-Dynamic Co-occurrence Graph Convolutional Networks for Skeleton-Based Action Recog- nition. IEEE Transactions on Circuits and Systems for Video Techn...
2024
-
[132]
Yisheng Zhu, Hui Shuai, Guangcan Liu, and Qingshan Liu. 2022. Multilevel spatial–temporal excited graph network for skeleton-based action recognition. IEEE Transactions on Image Processing 32 (2022), 496–508
2022
-
[2020]
In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition
Disentangling and unifying graph convolutions for skeleton-based action recognition. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 143–152
-
[2021]
IEEE Transactions on Neural Networks and Learning Systems 33, 9 (2021), 4800–4814
Memory attention networks for skeleton-based action recognition. IEEE Transactions on Neural Networks and Learning Systems 33, 9 (2021), 4800–4814
2021
-
[2022]
Engineering Applications of Artificial Intelligence 110 (2022), 104675
Temporal segment graph convolutional networks for skeleton-based action recognition. Engineering Applications of Artificial Intelligence 110 (2022), 104675
2022
-
[2024]
Pattern Recognition 151 (2024), 110427
Spatial–temporal hypergraph based on dual-stage attention network for multi-view data lightweight action recognition. Pattern Recognition 151 (2024), 110427
2024
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.