REVIEW 3 cited by
A Decoupled Spatio-Temporal Framework for Skeleton-based Action Segmentation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Effectively modeling discriminative spatio-temporal information is essential for segmenting activities in long action sequences. However, we observe that existing methods are limited in weak spatio-temporal modeling capability due to two forms of decoupled modeling: (i) cascaded interaction couples spatial and temporal modeling, which over-smooths motion modeling over the long sequence, and (ii) joint-shared temporal modeling adopts shared weights to model each joint, ignoring the distinct motion patterns of different joints. We propose a Decoupled Spatio-Temporal Framework (DeST) to address the above issues. Firstly, we decouple the cascaded spatio-temporal interaction to avoid stacking multiple spatio-temporal blocks, while achieving sufficient spatio-temporal interaction. Specifically, DeST performs once unified spatial modeling and divides the spatial features into different groups of subfeatures, which then adaptively interact with temporal features from different layers. Since the different sub-features contain distinct spatial semantics, the model could learn the optimal interaction pattern at each layer. Meanwhile, inspired by the fact that different joints move at different speeds, we propose joint-decoupled temporal modeling, which employs independent trainable weights to capture distinctive temporal features of each joint. On four large-scale benchmarks of different scenes, DeST significantly outperforms current state-of-the-art methods with less computational complexity.
Forward citations
Cited by 3 Pith papers
-
Skeleton Motion Words for Unsupervised Skeleton-Based Temporal Action Segmentation
An autoencoder with joint-disentangled embeddings and vector-quantized temporal patches segments skeleton sequences into actions without labels, beating prior unsupervised methods on HuGaDB, LARa, and two of three BAB...
-
DuoCLR: Dual-Surrogate Contrastive Learning for Skeleton-based Human Action Segmentation
DuoCLR pretrains on trimmed skeleton sequences using Shuffle-and-Warp multi-action permutations and two surrogate tasks, significantly improving action segmentation on untrimmed videos.
-
Foundation Model for Skeleton-Based Human Action Understanding
The abstract claims a new skeleton-based foundation model (USDRL) with state-of-the-art results on 9 action-understanding tasks, but the supplied full text is an unrelated paper on brain-computer interface cybersecurity.
Discussion (0). Continue with ORCID to comment.