REVIEW 5 cited by
Learning to Generate Diverse Dance Motions with Transformer
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
With the ongoing pandemic, virtual concerts and live events using digitized performances of musicians are getting traction on massive multiplayer online worlds. However, well choreographed dance movements are extremely complex to animate and would involve an expensive and tedious production process. In addition to the use of complex motion capture systems, it typically requires a collaborative effort between animators, dancers, and choreographers. We introduce a complete system for dance motion synthesis, which can generate complex and highly diverse dance sequences given an input music sequence. As motion capture data is limited for the range of dance motions and styles, we introduce a massive dance motion data set that is created from YouTube videos. We also present a novel two-stream motion transformer generative model, which can generate motion sequences with high flexibility. We also introduce new evaluation metrics for the quality of synthesized dance motions, and demonstrate that our system can outperform state-of-the-art methods. Our system provides high-quality animations suitable for large crowds for virtual concerts and can also be used as reference for professional animation pipelines. Most importantly, we show that vast online videos can be effective in training dance motion models.
Forward citations
Cited by 5 Pith papers
-
Text Dictates, Music Decorates: Energy-based Attention for Editable Dance Motion Generation
STREAM decouples text (via AdaLN) from music (via energy-based BEAM attention) to generate editable, musically aligned dance motions with a new annotated dataset and editability metric.
-
Spatial-Temporal Graph Mamba for Music-Guided Dance Video Synthesis
STG-Mamba generates dance videos from music using a spatial-temporal graph Mamba block for skeleton generation and forward-backward self-supervised losses for video synthesis, reporting SOTA on benchmarks.
-
Stochastic Human Motion Prediction with Memory of Action Transition and Action Characteristic
Adding a soft-transition action bank, an action characteristic bank, and adaptive attention fusion to the WAT baseline improves action-conditioned human motion prediction on four benchmarks.
-
DuetGen: Music Driven Two-Person Dance Generation via Hierarchical Masked Modeling
DuetGen is a two-stage masked-modeling system that converts music into synchronized, interactive two-person dance motion, and claims state-of-the-art results on the DD100 duet dataset.
-
MotionRAG-Diff: A Retrieval-Augmented Diffusion Framework for Long-Term Music-to-Dance Generation
MotionRAG-Diff combines contrastive retrieval from a motion graph with a diffusion model to generate long music-synchronized dance, reporting strong beat alignment but mixed quality scores.
Discussion (0). Continue with ORCID to comment.