Pith. sign in

REVIEW 3 cited by

A Decoupled Spatio-Temporal Framework for Skeleton-based Action Segmentation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.05830 v1 pith:MLFQ7RQ3 submitted 2023-12-10 cs.CV

classification cs.CV
keywords modelingdifferentspatio-temporaltemporalinteractionspatialdecoupleddest
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Effectively modeling discriminative spatio-temporal information is essential for segmenting activities in long action sequences. However, we observe that existing methods are limited in weak spatio-temporal modeling capability due to two forms of decoupled modeling: (i) cascaded interaction couples spatial and temporal modeling, which over-smooths motion modeling over the long sequence, and (ii) joint-shared temporal modeling adopts shared weights to model each joint, ignoring the distinct motion patterns of different joints. We propose a Decoupled Spatio-Temporal Framework (DeST) to address the above issues. Firstly, we decouple the cascaded spatio-temporal interaction to avoid stacking multiple spatio-temporal blocks, while achieving sufficient spatio-temporal interaction. Specifically, DeST performs once unified spatial modeling and divides the spatial features into different groups of subfeatures, which then adaptively interact with temporal features from different layers. Since the different sub-features contain distinct spatial semantics, the model could learn the optimal interaction pattern at each layer. Meanwhile, inspired by the fact that different joints move at different speeds, we propose joint-decoupled temporal modeling, which employs independent trainable weights to capture distinctive temporal features of each joint. On four large-scale benchmarks of different scenes, DeST significantly outperforms current state-of-the-art methods with less computational complexity.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Skeleton Motion Words for Unsupervised Skeleton-Based Temporal Action Segmentation

    cs.CV 2025-08 conditional novelty 6.0 of 10

    An autoencoder with joint-disentangled embeddings and vector-quantized temporal patches segments skeleton sequences into actions without labels, beating prior unsupervised methods on HuGaDB, LARa, and two of three BAB...

  2. DuoCLR: Dual-Surrogate Contrastive Learning for Skeleton-based Human Action Segmentation

    cs.CV 2025-09 conditional novelty 5.0 of 10

    DuoCLR pretrains on trimmed skeleton sequences using Shuffle-and-Warp multi-action permutations and two surrogate tasks, significantly improving action segmentation on untrimmed videos.

  3. Foundation Model for Skeleton-Based Human Action Understanding

    cs.CV 2025-08 unverdicted novelty 4.0 of 10

    The abstract claims a new skeleton-based foundation model (USDRL) with state-of-the-art results on 9 action-understanding tasks, but the supplied full text is an unrelated paper on brain-computer interface cybersecurity.

Pith tools