REVIEW 1 major objections 7 minor 118 references
The paper claims that a skeleton-only interaction recognizer can learn from RGB video during training—via contrastive alignment of skeleton and visual interaction features—and then outperform previous methods at inference using only skeleto
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 18:15 UTC pith:H4LVPAA2
load-bearing objection Competent extension of ISTA-Net with multi-modal alignment; the visual-alignment claim is plausible but only weakly evidenced, so the revision needs statistical rigor. the 1 major comments →
STAR: Skeletal Token Alignment and Rearrangement for Interaction Recognition
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
STAR is claimed to be the first skeleton-based interaction recognition method to use multi-modal alignment with visual interaction features. In training, an object detector finds both entities in each RGB frame; a maximum square box covering all detections at all sampled times crops an interaction Region of Interest, which a pretrained video encoder embeds. The skeleton encoder's intermediate feature vector is aligned to that visual embedding by a contrastive objective in a shared latent space, and a separate refinement head learns to predict the label from the visual embedding, so at test time it can be fed with the skeleton feature instead. On the Chico and HARPER human-robot datasets and
What carries the argument
Two mechanisms carry the argument. Entity Rearrangement randomly permutes the order of the two entities during training; because interaction labels are invariant under entity swap, the symmetric group reduces variance in the estimator and stabilizes optimization. Interactive Spatiotemporal Tokens are 3D windows sliding over time, joints, and entities, so each token bundles a local spatiotemporal patch of the interaction; stacked Token Self-Attention blocks then model interdependencies without relying on adjacency matrices, which matters because human bone structure differs from a quadruped robot's. The third mechanism is the multimodal alignment: a contrastive loss ties the skeleton token fe
Load-bearing premise
The load-bearing premise is that the Focus on Interactions crops—obtained from a pretrained object detector and a maximum square box—contain the interaction-discriminative visual information that survives temporal sampling, and that the detector does not fail; the paper itself notes (Fig. 8) failures when entities are partially out of frame or too far apart, and the conclusion acknowledges the assumption of known actor types.
What would settle it
Run STAR on a test set where the two interacting entities are frequently partially out of frame or far apart; if accuracy drops to the no-alignment baseline (about 91.5% in the paper's Table III) while the alignment loss stays high, the visual teacher is injecting noise rather than signal, and random-crop controls would confirm whether FoI localization is the source of gains.
If this is right
- If STAR is correct, a skeleton-only model can learn from video during training and match or beat methods that need video at inference, enabling privacy-sensitive and low-light deployments.
- The Entity Rearrangement perspective implies that two-entity interaction modeling can treat entity order as a symmetry, reducing reliance on subject-specific adjacency priors in graph-based skeleton models.
- Because the visual target comes from Focus on Interactions cropping, alignment quality depends on detector localization; the paper shows success cases under occlusion and identifies partially out-of-frame or far-apart entities as failure cases.
- The refinement head adds a 'think-twice' step at test time with negligible overhead (about 0.41M additional parameters reported for the alignment components).
- On fine-grained categories such as 'hit with object' and 'punch/slap', visual alignment yields large reported category-level accuracy gains, suggesting video cues specifically resolve ambiguities in contact point and manipulated object.
Where Pith is reading between the lines
- A natural extension of the training-time-video, inference-time-skeleton recipe is to distill knowledge from larger video foundation models without increasing deployment cost; the paper's encoder benchmark suggests video models transfer more useful cues than image models.
- The known-actor-type assumption flagged in the conclusion could be tested by clustering roles from spatiotemporal movement patterns; if that works, STAR-style alignment could apply to open-world human-robot interaction instead of pre-specified entity pairs.
- The FoI failure cases suggest a specific stress test: when the detector misses entities or crops too large a region, the contrastive objective may pull skeleton features toward noise; probing performance under progressively larger entity distances would reveal how much robustness margin remains.
- Because the visual branch is a training-time teacher only, the same alignment objective could be applied to other privacy-sensitive modalities (e.g., depth or thermal) to enrich skeleton features without changing the skeleton-only inference pipeline.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents STAR, a framework for skeleton-based human-robot and human-human interaction recognition. STAR consists of a skeleton encoder built on Entity Rearrangement (ER) and Interactive Spatiotemporal Tokens (ISTs), a visual branch that extracts interaction-focused RGB features via Focus on Interactions (FoI), and a contrastive alignment loss that trains skeleton and visual features in a shared latent space. At inference, only the skeleton branch is used, with an auxiliary refinement head that estimates visual-informed logits from skeleton features. The method is evaluated on Chico, HARPER, NTU Mutual 11, and NTU Mutual 26, reporting state-of-the-art accuracy. The paper also provides ablations for each component, robustness tests under noise and masking, and a public code release.
Significance. If the central claim holds, STAR would be a useful contribution: it demonstrates that visual cues can be distilled into a skeleton-only model through training-time alignment, preserving the efficiency and privacy advantages of skeleton-based inference while improving accuracy. The paper's strengths are its comprehensive evaluation across four benchmarks, including two recent HRI datasets, its component-wise ablations, and the release of code. The skeleton encoder design (ISTs and ER) appears well motivated and the authors provide quantitative and qualitative evidence that the learned representations are more discriminative. However, the paper's headline contribution—the multi-modal alignment—is supported by a small accuracy gain (about 0.7 percentage points on NTU Mutual 26) and lacks statistical validation. The alignment loss itself, as written, also raises a technical concern about whether it actually pulls positive pairs together. These issues must be addressed before the central claim can be considered established.
major comments (1)
- [§II-B] The claim that STAR is 'the first to leverage multi-modal alignment to learn both human-robot and human-human interactions' is stated in the introduction and related work, but the comparison to prior multi-modal alignment works (e.g., GAP, MMCL, C2VL) is made at the level of task scope. The paper does not discuss whether any of those methods could be adapted to the interaction setting with minimal changes. A more careful positioning, perhaps with an adapted-baseline experiment, would strengthen the novelty claim.
minor comments (7)
- [Abstract / Intro] The sentence 'with a refinement head further refines predictions' has a grammatical error ('head further refines'). Please revise.
- [§I, Fig. 1] The figure caption for Fig. 1 uses 'Ours can infer the interactions in dark or privacy-sensitive workspaces' but the figure shows only an illustrative example; consider clarifying that this is a schematic and not an actual experimental result.
- [§III-A] The use of Chen et al. [107] to argue variance reduction from random permutation is appropriate, but the notation O = d πO and the approximate-invariance case are introduced briefly. A short intuitive explanation of why approximate invariance also yields variance reduction would improve readability.
- [§III-C] In Eq. (8), the operator ∩ is defined as 'intersection' with the original video, but the actual operation is a spatial crop using the maximum covering box. The notation is confusing; please use a clearer operator name, e.g., 'crop'.
- [§IV-E, Table III(a)] The table title 'PRETRAINED VISION ENCODERS' includes 'No Alignment' as a row, which is not a vision encoder. Consider moving that row to Table III(b) or renaming the table.
- [§IV-E, Table III(e)] The parameter counts in Table III(e) are given for different encoder layers, but the 'No Alignment' row reports a different parameter count (6.22M) that is not directly comparable to the 6.63M used for the default model. Please clarify whether the parameter count includes the alignment MLP and refinement head, and note that removing alignment reduces parameters.
- [§IV-E, Table III(h)] The robustness experiment applies noise with σ=0.01 and masking with p=0.01, but the choice of these values is not justified. A small sensitivity analysis over noise/mask levels would make the robustness claim more convincing.
Circularity Check
No circularity found: STAR's alignment and refinement are empirical training objectives, not derivations that reduce to their inputs.
full rationale
STAR's derivation chain is empirical rather than algebraic. The skeleton encoder (Eqs. 3-7), Focus on Interactions (Eq. 8), contrastive alignment objective (Eqs. 9-10), and refinement head (Eqs. 11-12) define training procedures; the claimed skeleton-only benefit is measured by test-set accuracy and ablations, not derived from the equations. I could not exhibit any step where an output reduces by construction to a fitted value or to a definitional identity. The self-citations to the authors' ISTA-Net [15] and CHASE [85] are used for motivation, baselines, and design references; these are externally published results and the multi-modal alignment contribution does not depend on their correctness. The 'first to introduce multi-modal alignment' statement is a novelty claim, not a circular derivation. Skeptical concerns about the small alignment gain (0.73%), lack of significance testing, and FoI failure cases are evidence-strength and robustness issues, not circularity. Therefore the appropriate score is 0.
Axiom & Free-Parameter Ledger
free parameters (7)
- λ1 (alignment loss weight) =
0.4
- λ2 (refinement loss weight) =
0.7
- β (refinement fusion coefficient) =
0.8
- τ (contrastive temperature) =
not stated
- Window shape W=(Tw,Jw,Ew) =
(20,1,2)
- Alignment block index l_hat =
6
- Encoder layers L, downsampling layers, heads =
L=8, LD={3,5}, H=3
axioms (6)
- domain assumption Interaction labels are exactly permutation-invariant along the entity dimension (O =d πO, Eq. 2).
- domain assumption Visual RoI features extracted by FoI contain interaction-discriminative cues absent from skeletons.
- domain assumption The object detector Ω provides correct entity bounding boxes for the max covering box.
- standard math Chen et al. [107] group-theoretic variance-reduction result applies to ERM for skeleton interaction learning.
- domain assumption Actor types are known in advance for training and inference.
- domain assumption Skeleton and RGB sequences are correctly paired and synchronized in all datasets.
read the original abstract
Understanding physical human-robot and human-human interactions is a challenging yet emerging topic in 3D vision. While most existing methods rely on skeleton sequences--effective in low-light and privacy-sensitive environment--they face two major challenges: 1) learning and effectively exploiting interaction cues from skeletal data, and 2) compensating for the lack of visual information absent in skeletons alone. To address these challenges, we propose skeletal token alignment and rearrangement (STAR) for human-robot and human-human interaction recognition. It learns interaction-specific skeleton features and enriches them using visual cues by aligning skeleton and RGB video representations in a shared latent space. Specifically, STAR consists of three key components. First, we design a skeleton encoder that captures fine-grained interdependencies using Entity Rearrangement (ER) and Interactive Spatiotemporal Tokens (ISTs). Second, we present Visual Interaction Encoding that introduces a Focus on Interactions (FoI) strategy to attend to spatiotemporal regions relevant to interactions in RGB videos. Finally, these representations are aligned via a contrastive learning objective, with a refinement head further refines predictions. During training, STAR leverages both skeleton and RGB video data to learn robust, discriminative interaction representations. At inference time, it operates on skeletons alone, retaining visual-informed benefits while preserving skeleton-only efficiency. Extensive experiments on Chico, HARPER, NTU Mutual 11 and 26 datasets consistently validate our approach by demonstrating superior performance over state-of-the-art methods. Our code is publicly available at https://github.com/Necolizer/STAR.
Figures
Reference graph
Works this paper leans on
-
[1]
Hierarchical aggregated graph neu- ral network for skeleton-based action recognition,
P. Geng, X. Lu, W. Li, and L. Lyu, “Hierarchical aggregated graph neu- ral network for skeleton-based action recognition,”IEEE Transactions on Multimedia, pp. 1–16, 2024
2024
-
[2]
Joints-centered spatial-temporal features fused skeleton convolution network for action recognition,
W. Song, T. Chu, S. Li, N. Li, A. Hao, and H. Qin, “Joints-centered spatial-temporal features fused skeleton convolution network for action recognition,”IEEE Transactions on Multimedia, vol. 26, pp. 4602– 4616, 2024
2024
-
[3]
Noise- tolerant learning for audio-visual action recognition,
H. Han, Q. Zheng, M. Luo, K. Miao, F. Tian, and Y . Chen, “Noise- tolerant learning for audio-visual action recognition,”IEEE Transac- tions on Multimedia, vol. 26, pp. 7761–7774, 2024
2024
-
[4]
Commonsense knowledge prompt- ing for few-shot action recognition in videos,
Y . Shi, X. Wu, H. Lin, and J. Luo, “Commonsense knowledge prompt- ing for few-shot action recognition in videos,”IEEE Transactions on Multimedia, vol. 26, pp. 8395–8405, 2024
2024
-
[5]
Exploring rich semantics for open-set action recognition,
Y . Hu, J. Gao, J. Dong, B. Fan, and H. Liu, “Exploring rich semantics for open-set action recognition,”IEEE Transactions on Multimedia, vol. 26, pp. 5410–5421, 2024
2024
-
[6]
Dear-net: Learning diver- sities for skeleton-based early action recognition,
R. Wang, J. Liu, Q. Ke, D. Peng, and Y . Lei, “Dear-net: Learning diver- sities for skeleton-based early action recognition,”IEEE Transactions on Multimedia, vol. 25, pp. 1175–1189, 2023
2023
-
[7]
Just addπ! pose induced video transformers for understanding activities of daily living,
D. Reilly and S. Das, “Just addπ! pose induced video transformers for understanding activities of daily living,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2024, pp. 18 340–18 350
2024
-
[8]
Mmnet: A model-based multimodal network for human action recognition in rgb-d videos,
B. X. Yu, Y . Liu, X. Zhang, S.-h. Zhong, and K. C. Chan, “Mmnet: A model-based multimodal network for human action recognition in rgb-d videos,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 45, no. 3, pp. 3522–3538, 2023
2023
-
[9]
A unified multi- modal de- and re-coupling framework for rgb-d motion recognition,
B. Zhou, P. Wang, J. Wan, Y . Liang, and F. Wang, “A unified multi- modal de- and re-coupling framework for rgb-d motion recognition,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 45, no. 10, pp. 11 428–11 442, 2023
2023
-
[10]
Selective, interpretable and motion consistent privacy attribute obfuscation for action recognition,
F. Ilic, H. Zhao, T. Pock, and R. P. Wildes, “Selective, interpretable and motion consistent privacy attribute obfuscation for action recognition,” inConference on Computer Vision and Pattern Recognition (CVPR), 2024
2024
-
[11]
On the benefits of 3d pose and tracking for human action recognition,
J. Rajasegaran, G. Pavlakos, A. Kanazawa, C. Feichtenhofer, and J. Malik, “On the benefits of 3d pose and tracking for human action recognition,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2023, pp. 640–649
2023
-
[12]
Smam: Self and mutual adaptive matching for skeleton-based few-shot action recognition,
Z. Li, X. Gong, R. Song, P. Duan, J. Liu, and W. Zhang, “Smam: Self and mutual adaptive matching for skeleton-based few-shot action recognition,”IEEE Transactions on Image Processing, vol. 32, pp. 392– 402, 2023
2023
-
[13]
Integrating image and textual information in human–robot interactions for children with autism spectrum disorder,
X. Yang, M.-L. Shyu, H.-Q. Yu, S.-M. Sun, N.-S. Yin, and W. Chen, “Integrating image and textual information in human–robot interactions for children with autism spectrum disorder,”IEEE Transactions on Multimedia, vol. 21, no. 3, pp. 746–759, 2019
2019
-
[14]
Jrdb-social: A multifaceted robotic dataset for understanding of context and dynamics of human interactions within social groups,
S. Jahangard, Z. Cai, S. Wen, and H. Rezatofighi, “Jrdb-social: A multifaceted robotic dataset for understanding of context and dynamics of human interactions within social groups,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024
2024
-
[15]
Interactive spatiotem- poral token attention network for skeleton-based general interactive ac- tion recognition,
Y . Wen, Z. Tang, Y . Pang, B. Ding, and M. Liu, “Interactive spatiotem- poral token attention network for skeleton-based general interactive ac- tion recognition,” inIEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2023, pp. 7886–7892
2023
-
[16]
Pose forecasting in industrial human-robot collaboration,
A. Sampieri, G. M. D. di Melendugno, A. Avogaro, F. Cunico, F. Setti, G. Skenderi, M. Cristani, and F. Galasso, “Pose forecasting in industrial human-robot collaboration,” inProceedings of the 17th European Conference on Computer Vision (ECCV), 2022, pp. 51–69
2022
-
[17]
Exploring 3d human pose estimation and forecasting from the robot’s perspective: The harper dataset,
A. Avogaro, A. Toaiari, F. Cunico, X. Xu, H. Dafas, A. Vinciarelli, E. Li, and M. Cristani, “Exploring 3d human pose estimation and forecasting from the robot’s perspective: The harper dataset,” in2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2024
2024
-
[18]
Pedestrian trajectory prediction based on social interactions learning with random weights,
J. Xie, S. Zhang, B. Xia, Z. Xiao, H. Jiang, S. Zhou, Z. Qin, and H. Chen, “Pedestrian trajectory prediction based on social interactions learning with random weights,”IEEE Transactions on Multimedia, vol. 26, pp. 7503–7515, 2024
2024
-
[19]
Inter- action transformer for human reaction generation,
B. Chopin, H. Tang, N. Otberdout, M. Daoudi, and N. Sebe, “Inter- action transformer for human reaction generation,”IEEE Transactions on Multimedia, vol. 25, pp. 8842–8854, 2023
2023
-
[20]
Spikepoint: An efficient point-based spiking neural network for event cameras action recognition,
H. Ren, Y . ZHOU, X. LIN, Y . Huang, H. FU, J. Song, and B. Cheng, “Spikepoint: An efficient point-based spiking neural network for event cameras action recognition,” inInternational Conference on Learning Representations, vol. 2024, 2024, pp. 27 827–27 846
2024
-
[21]
Scaling up dynamic human-scene interaction modeling,
N. Jiang, Z. Zhang, H. Li, X. Ma, Z. Wang, Y . Chen, T. Liu, Y . Zhu, and S. Huang, “Scaling up dynamic human-scene interaction modeling,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 1737–1747
2024
-
[22]
An outlook into the future of egocentric vision,
C. Plizzari, G. Goletto, A. Furnari, S. Bansal, F. Ragusa, G. M. Farinella, D. Damen, and T. Tommasi, “An outlook into the future of egocentric vision,”International Journal of Computer Vision, May 2024
2024
-
[23]
Locllm: Exploiting generalizable human keypoint localization via large language model,
D. Wang, S. Xuan, and S. Zhang, “Locllm: Exploiting generalizable human keypoint localization via large language model,” inProceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2024, pp. 614–623
2024
-
[24]
Intergen: Diffusion- based multi-human motion generation under complex interactions,
H. Liang, W. Zhang, W. Li, J. Yu, and L. Xu, “Intergen: Diffusion- based multi-human motion generation under complex interactions,” International Journal of Computer Vision, pp. 1–21, 2024
2024
-
[25]
Multi-modal & multi-view & interactive benchmark dataset for human action recognition,
N. Xu, A. Liu, W. Nie, Y . Wong, F. Li, and Y . Su, “Multi-modal & multi-view & interactive benchmark dataset for human action recognition,” inProceedings of the 23rd ACM International Conference on Multimedia (ACMMM), ser. MM ’15, 2015, p. 1195–1198
2015
-
[26]
Motionbert: A unified perspective on learning human motion representations,
W. Zhu, X. Ma, Z. Liu, L. Liu, W. Wu, and Y . Wang, “Motionbert: A unified perspective on learning human motion representations,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), October 2023, pp. 15 085–15 099
2023
-
[27]
Versatile multi-modal pre-training for human-centric perception,
F. Hong, L. Pan, Z. Cai, and Z. Liu, “Versatile multi-modal pre-training for human-centric perception,” inProceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition (CVPR), June 2022, pp. 16 156–16 166
2022
-
[28]
Beyond appearance: A semantic controllable self-supervised learning framework for human-centric visual tasks,
W. Chen, X. Xu, J. Jia, H. Luo, Y . Wang, F. Wang, R. Jin, and X. Sun, “Beyond appearance: A semantic controllable self-supervised learning framework for human-centric visual tasks,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2023, pp. 15 050–15 061
2023
-
[29]
Humanbench: Towards general human-centric perception with projector assisted pretraining,
S. Tang, C. Chen, Q. Xie, M. Chen, Y . Wang, Y . Ci, L. Bai, F. Zhu, H. Yang, L. Yi, R. Zhao, and W. Ouyang, “Humanbench: Towards general human-centric perception with projector assisted pretraining,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2023, pp. 21 970–21 982
2023
-
[30]
HAP: Structure-aware masked image modeling for human- centric perception,
J. Yuan, X. Zhang, H. Zhou, J. Wang, Z. Qiu, Z. Shao, S. Zhang, S. Long, K. Kuang, K. Yao, J. Han, E. Ding, L. Lin, F. Wu, and J. Wang, “HAP: Structure-aware masked image modeling for human- centric perception,” inThirty-seventh Conference on Neural Informa- tion Processing Systems (NeurIPS), 2023
2023
-
[31]
Unihcp: A unified model for human-centric JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 12 perceptions,
Y . Ci, Y . Wang, M. Chen, S. Tang, L. Bai, F. Zhu, R. Zhao, F. Yu, D. Qi, and W. Ouyang, “Unihcp: A unified model for human-centric JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 12 perceptions,” inProceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition (CVPR), June 2023, pp. 17 840– 17 852
2021
-
[32]
Hulk: A universal knowledge translator for human-centric tasks,
Y . Wang, Y . Wu, W. He, X. Guo, F. Zhu, L. Bai, R. Zhao, J. Wu, T. He, W. Ouyang, and S. Tang, “Hulk: A universal knowledge translator for human-centric tasks,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 47, no. 7, pp. 5672–5689, 2025
2025
-
[33]
Sapiens: Foundation for human vision models,
R. Khirodkar, T. Bagautdinov, J. Martinez, S. Zhaoen, A. James, P. Selednik, S. Anderson, and S. Saito, “Sapiens: Foundation for human vision models,” inProceedings of the 18th European Conference on Computer Vision (ECCV), 2024
2024
-
[34]
Learning composite latent structures for 3d human action representation and recognition,
P. Wei, H. Sun, and N. Zheng, “Learning composite latent structures for 3d human action representation and recognition,”IEEE Transactions on Multimedia, vol. 21, no. 9, pp. 2195–2208, 2019
2019
-
[35]
Navigating open set scenarios for skeleton-based action recognition,
K. Peng, C. Yin, J. Zheng, R. Liu, D. Schneider, J. Zhang, K. Yang, M. S. Sarfraz, R. Stiefelhagen, and A. Roitberg, “Navigating open set scenarios for skeleton-based action recognition,”Proceedings of the AAAI Conference on Artificial Intelligence, vol. 38, no. 5, pp. 4487– 4496, Mar. 2024
2024
-
[36]
One-shot action recognition via multi-scale spatial-temporal skeleton matching,
S. Yang, J. Liu, S. Lu, E. M. Hwa, and A. C. Kot, “One-shot action recognition via multi-scale spatial-temporal skeleton matching,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 46, no. 7, pp. 5149–5156, 2024
2024
-
[37]
Self- supervised 3d action representation learning with skeleton cloud colorization,
S. Yang, J. Liu, S. Lu, E. M. Hwa, Y . Hu, and A. C. Kot, “Self- supervised 3d action representation learning with skeleton cloud colorization,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 46, no. 1, pp. 509–524, 2024
2024
-
[38]
Meet jeanie: A similarity measure for 3d skeleton sequences via temporal-viewpoint alignment,
L. Wang, J. Liu, L. Zheng, T. Gedeon, and P. Koniusz, “Meet jeanie: A similarity measure for 3d skeleton sequences via temporal-viewpoint alignment,”Int. J. Comput. Vision, vol. 132, no. 9, p. 4091–4122, may 2024
2024
-
[39]
Neural koopman pooling: Control- inspired temporal dynamics encoding for skeleton-based action recog- nition,
X. Wang, X. Xu, and Y . Mu, “Neural koopman pooling: Control- inspired temporal dynamics encoding for skeleton-based action recog- nition,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2023, pp. 10 597–10 607
2023
-
[40]
Learning discriminative representations for skeleton based action recognition,
H. Zhou, Q. Liu, and Y . Wang, “Learning discriminative representations for skeleton based action recognition,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2023, pp. 10 608–10 617
2023
-
[41]
You2me: Inferring body pose in egocentric video via first and second person interactions,
E. Ng, D. Xiang, H. Joo, and K. Grauman, “You2me: Inferring body pose in egocentric video via first and second person interactions,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2020
2020
-
[42]
Skeleton- based online action prediction using scale selection network,
J. Liu, A. Shahroudy, G. Wang, L.-Y . Duan, and A. C. Kot, “Skeleton- based online action prediction using scale selection network,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 42, no. 6, pp. 1453–1467, 2020
2020
-
[43]
Interaction relational network for mutual action recognition,
M. Perez, J. Liu, and A. C. Kot, “Interaction relational network for mutual action recognition,”IEEE Transactions on Multimedia, vol. 24, pp. 366–376, 2022
2022
-
[44]
Igformer: Interaction graph transformer for skeleton-based human interaction recognition,
Y . Pang, Q. Ke, H. Rahmani, J. Bailey, and J. Liu, “Igformer: Interaction graph transformer for skeleton-based human interaction recognition,” inProceedings of the 17th European Conference on Computer Vision (ECCV), 2022, pp. 605–622
2022
-
[45]
Graph diffusion convolutional network for skeleton based semantic recognition of two- person actions,
S. Li, X. He, W. Song, A. Hao, and H. Qin, “Graph diffusion convolutional network for skeleton based semantic recognition of two- person actions,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 45, no. 7, pp. 8477–8493, 2023
2023
-
[46]
Learning mutual exci- tation for hand-to-hand and human-to-human interaction recognition,
M. Liu, C. Chen, S. Wu, F. Meng, and H. Liu, “Learning mutual exci- tation for hand-to-hand and human-to-human interaction recognition,” IEEE Transactions on Human-Machine Systems, pp. 1–10, 2025
2025
-
[47]
Multi-modal enhancement transformer network for skeleton-based human interaction recognition,
Q. Hu and H. Liu, “Multi-modal enhancement transformer network for skeleton-based human interaction recognition,”Biomimetics, vol. 9, no. 3, 2024
2024
-
[48]
Ntu rgb+d: A large scale dataset for 3d human activity analysis,
A. Shahroudy, J. Liu, T.-T. Ng, and G. Wang, “Ntu rgb+d: A large scale dataset for 3d human activity analysis,” in2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 1010– 1019
2016
-
[49]
Ntu rgb+d 120: A large-scale benchmark for 3d human activity understanding,
J. Liu, A. Shahroudy, M. Perez, G. Wang, L.-Y . Duan, and A. C. Kot, “Ntu rgb+d 120: A large-scale benchmark for 3d human activity understanding,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 42, no. 10, pp. 2684–2701, 2020
2020
-
[50]
Cross-view action modeling, learning, and recognition,
J. Wang, X. Nie, Y . Xia, Y . Wu, and S.-C. Zhu, “Cross-view action modeling, learning, and recognition,” inConference on Computer Vision and Pattern Recognition (CVPR), 2014, p. 2649–2656
2014
-
[51]
Pku-mmd: A large scale benchmark for skeleton-based human action understanding,
C. Liu, Y . Hu, Y . Li, S. Song, and J. Liu, “Pku-mmd: A large scale benchmark for skeleton-based human action understanding,” inPro- ceedings of the Workshop on Visual Analysis in Smart and Connected Communities, ser. VSCC ’17. New York, NY , USA: Association for Computing Machinery, 2017, p. 1–8
2017
-
[52]
Toyota smarthome: Real-world activities of daily living,
S. Das, R. Dai, M. Koperski, L. Minciullo, L. Garattoni, F. Bremond, and G. Francesca, “Toyota smarthome: Real-world activities of daily living,” inProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), October 2019
2019
-
[53]
Assemblyhands: Towards egocentric activity understanding via 3d hand pose estimation,
T. Ohkawa, K. He, F. Sener, T. Hodan, L. Tran, and C. Keskin, “Assemblyhands: Towards egocentric activity understanding via 3d hand pose estimation,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2023, pp. 12 999–13 008
2023
-
[54]
Hi4d: 4d instance segmentation of close human interaction,
Y . Yin, C. Guo, M. Kaufmann, J. J. Zarate, J. Song, and O. Hilliges, “Hi4d: 4d instance segmentation of close human interaction,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2023, pp. 17 016–17 027
2023
-
[55]
Inter-x: Towards versatile human-human interaction analysis,
L. Xu, X. Lv, Y . Yan, X. Jin, S. Wu, C. Xu, Y . Liu, Y . Zhou, F. Rao, X. Sheng, Y . Liu, W. Zeng, and X. Yang, “Inter-x: Towards versatile human-human interaction analysis,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024
2024
-
[56]
Intergen: Diffusion- based multi-human motion generation under complex interactions,
H. Liang, W. Zhang, W. Li, J. Yu, and L. Xu, “Intergen: Diffusion- based multi-human motion generation under complex interactions,” International Journal of Computer Vision, Mar 2024
2024
-
[57]
Co- occurrence feature learning for skeleton based action recognition using regularized deep lstm networks,
W. Zhu, C. Lan, J. Xing, W. Zeng, Y . Li, L. Shen, and X. Xie, “Co- occurrence feature learning for skeleton based action recognition using regularized deep lstm networks,” inProceedings of the Thirtieth AAAI Conference on Artificial Intelligence, ser. AAAI’16. AAAI Press, 2016, p. 3697–3703
2016
-
[58]
Spatio-temporal lstm with trust gates for 3d human action recognition,
J. Liu, A. Shahroudy, D. Xu, and G. Wang, “Spatio-temporal lstm with trust gates for 3d human action recognition,” inProceedings of the 17th European Conference on Computer Vision (ECCV), 2016, pp. 816–833
2016
-
[59]
Global context- aware attention lstm networks for 3d action recognition,
J. Liu, G. Wang, P. Hu, L.-Y . Duan, and A. C. Kot, “Global context- aware attention lstm networks for 3d action recognition,” in2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 3671–3680
2017
-
[60]
View adaptive recurrent neural networks for high performance human action recognition from skeleton data,
P. Zhang, C. Lan, J. Xing, W. Zeng, J. Xue, and N. Zheng, “View adaptive recurrent neural networks for high performance human action recognition from skeleton data,” in2017 IEEE International Confer- ence on Computer Vision (ICCV), 2017, pp. 2136–2145
2017
-
[61]
Skeleton- based human action recognition with global context-aware attention lstm networks,
J. Liu, G. Wang, L.-Y . Duan, K. Abdiyeva, and A. C. Kot, “Skeleton- based human action recognition with global context-aware attention lstm networks,”IEEE Transactions on Image Processing, vol. 27, no. 4, pp. 1586–1599, 2018
2018
-
[62]
View adaptive neural networks for high performance skeleton-based human action recognition,
P. Zhang, C. Lan, J. Xing, W. Zeng, J. Xue, and N. Zheng, “View adaptive neural networks for high performance skeleton-based human action recognition,”IEEE Transactions on Pattern Analysis and Ma- chine Intelligence, vol. 41, no. 8, pp. 1963–1978, 2019
1963
-
[63]
Spatial temporal graph convolutional networks for skeleton-based action recognition,
S. Yan, Y . Xiong, and D. Lin, “Spatial temporal graph convolutional networks for skeleton-based action recognition,” inProceedings of the Thirty-Second AAAI Conference on Artificial Intelligence, ser. AAAI’18, vol. 32, no. 1, 2018, pp. 7444–7452
2018
-
[64]
Actional- structural graph convolutional networks for skeleton-based action recognition,
M. Li, S. Chen, X. Chen, Y . Zhang, Y . Wang, and Q. Tian, “Actional- structural graph convolutional networks for skeleton-based action recognition,” in2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019, pp. 3590–3598
2019
-
[65]
Two-stream adaptive graph convolutional networks for skeleton-based action recognition,
L. Shi, Y . Zhang, J. Cheng, and H. Lu, “Two-stream adaptive graph convolutional networks for skeleton-based action recognition,” in2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019, pp. 12 018–12 027
2019
-
[66]
Disentangling and unifying graph convolutions for skeleton-based action recognition,
Z. Liu, H. Zhang, Z. Chen, Z. Wang, and W. Ouyang, “Disentangling and unifying graph convolutions for skeleton-based action recognition,” in2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020, pp. 140–149
2020
-
[67]
Channel- wise topology refinement graph convolution for skeleton-based action recognition,
Y . Chen, Z. Zhang, C. Yuan, B. Li, Y . Deng, and W. Hu, “Channel- wise topology refinement graph convolution for skeleton-based action recognition,” inIEEE International Conference on Computer Vision (ICCV), 2021, pp. 13 359–13 368
2021
-
[68]
Infogcn: Representation learning for human skeleton-based action recognition,
H.-G. Chi, M. H. Ha, S. Chi, S. W. Lee, Q. Huang, and K. Ramani, “Infogcn: Representation learning for human skeleton-based action recognition,” inIEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, pp. 20 154–20 164
2022
-
[69]
Hierarchically decomposed graph convolutional networks for skeleton-based action recognition,
J. Lee, M. Lee, D. Lee, and S. Lee, “Hierarchically decomposed graph convolutional networks for skeleton-based action recognition,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), October 2023, pp. 10 444–10 453. JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 13
2023
-
[70]
Hypergraph neural network for skeleton-based action recognition,
X. Hao, J. Li, Y . Guo, T. Jiang, and M. Yu, “Hypergraph neural network for skeleton-based action recognition,”IEEE Transactions on Image Processing, vol. 30, pp. 2263–2275, 2021
2021
-
[71]
Pyskl: Towards good practices for skeleton action recognition,
H. Duan, J. Wang, K. Chen, and D. Lin, “Pyskl: Towards good practices for skeleton action recognition,” inProceedings of the 30th ACM International Conference on Multimedia (ACMMM), 2022, pp. 7351– 7354
2022
-
[72]
Tsgcnext: Dynamic-static multi- graph convolution for efficient skeleton-based action recognition,
D. Liu, X. Li, Z. Cai, and P. Chen, “Tsgcnext: Dynamic-static multi- graph convolution for efficient skeleton-based action recognition,” Expert Systems with Applications, vol. 276, p. 127081, 2025
2025
-
[73]
Degcn: Deformable graph convolutional networks for skeleton-based action recognition,
W. Myung, N. Su, J.-H. Xue, and G. Wang, “Degcn: Deformable graph convolutional networks for skeleton-based action recognition,”IEEE Transactions on Image Processing, vol. 33, pp. 2477–2490, 2024
2024
-
[74]
Blockgcn: Redefine topology awareness for skeleton-based action recognition,
Y . Zhou, X. Yan, Z.-Q. Cheng, Y . Yan, Q. Dai, and X.-S. Hua, “Blockgcn: Redefine topology awareness for skeleton-based action recognition,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2024, pp. 2049–2058
2024
-
[75]
Decoupled spatial-temporal attention network for skeleton-based action-gesture recognition,
L. Shi, Y . Zhang, J. Cheng, and H. Lu, “Decoupled spatial-temporal attention network for skeleton-based action-gesture recognition,” in 15th Asian Conference on Computer Vision (ACCV), 2020, p. 38–53
2020
-
[76]
Spatio-temporal segments attention for skeleton-based action recognition,
H. Qiu, B. Hou, B. Ren, and X. Zhang, “Spatio-temporal segments attention for skeleton-based action recognition,”Neurocomputing, vol. 518, pp. 30–38, 2023
2023
-
[77]
Hypergraph transformer for skeleton-based action recognition,
Y . Zhou, Z.-Q. Cheng, C. Li, Y . Geng, X. Xie, and M. Keu- per, “Hypergraph transformer for skeleton-based action recognition,” arXiv:2211.09590, 2022
Pith/arXiv arXiv 2022
-
[78]
A cuboid cnn model with an attention mechanism for skeleton-based action recognition,
K. Zhu, R. Wang, Q. Zhao, J. Cheng, and D. Tao, “A cuboid cnn model with an attention mechanism for skeleton-based action recognition,” IEEE Transactions on Multimedia, vol. 22, no. 11, pp. 2977–2989, 2020
2020
-
[79]
N. H. B. Long, “Step catformer: Spatial-temporal effective body- part cross attention transformer for skeleton-based action recognition,” arXiv:2312.03288, 2023
Pith/arXiv arXiv 2023
-
[80]
Masked motion predictors are strong 3d action representation learners,
Y . Mao, J. Deng, W. Zhou, Y . Fang, W. Ouyang, and H. Li, “Masked motion predictors are strong 3d action representation learners,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), October 2023, pp. 10 181–10 191
2023
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.