SkeletonLLM translates skeleton kinematics into compact image sequences via an end-to-end differentiable renderer DrAction and uses cooperative training to enable MLLMs to perform open-vocabulary action recognition, motion captioning, and QA across heterogeneous skeleton formats.
org/CorpusID:273662196
2 Pith papers cite this work. Polarity classification is still indexing.
fields
cs.CV 2years
2026 2verdicts
UNVERDICTED 2representative citing papers
FDSM adds spectral residual, timestep-adaptive spectral loss, and curriculum semantic abstraction to diffusion models for zero-shot skeleton-text action recognition and claims SOTA on NTU, PKU-MMD, and Kinetics-skeleton.
citing papers explorer
-
Universal Skeleton Understanding via Differentiable Rendering and MLLMs
SkeletonLLM translates skeleton kinematics into compact image sequences via an end-to-end differentiable renderer DrAction and uses cooperative training to enable MLLMs to perform open-vocabulary action recognition, motion captioning, and QA across heterogeneous skeleton formats.
-
Frequency-Enhanced Diffusion Models: Curriculum-Guided Semantic Alignment for Zero-Shot Skeleton Action Recognition
FDSM adds spectral residual, timestep-adaptive spectral loss, and curriculum semantic abstraction to diffusion models for zero-shot skeleton-text action recognition and claims SOTA on NTU, PKU-MMD, and Kinetics-skeleton.