Pith. sign in

REVIEW 2 cited by

Learning Music-Dance Representations through Explicit-Implicit Rhythm Synchronization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2207.03190 v2 pith:FKCOUXBY submitted 2022-07-07 cs.SD cs.CVcs.MMeess.AS

classification cs.SDcs.CVcs.MMeess.AS
keywords musicmusic-dancerhythmsrepresentationvisualdancelearningrhythm
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Although audio-visual representation has been proved to be applicable in many downstream tasks, the representation of dancing videos, which is more specific and always accompanied by music with complex auditory contents, remains challenging and uninvestigated. Considering the intrinsic alignment between the cadent movement of dancer and music rhythm, we introduce MuDaR, a novel Music-Dance Representation learning framework to perform the synchronization of music and dance rhythms both in explicit and implicit ways. Specifically, we derive the dance rhythms based on visual appearance and motion cues inspired by the music rhythm analysis. Then the visual rhythms are temporally aligned with the music counterparts, which are extracted by the amplitude of sound intensity. Meanwhile, we exploit the implicit coherence of rhythms implied in audio and visual streams by contrastive learning. The model learns the joint embedding by predicting the temporal consistency between audio-visual pairs. The music-dance representation, together with the capability of detecting audio and visual rhythms, can further be applied to three downstream tasks: (a) dance classification, (b) music-dance retrieval, and (c) music-dance retargeting. Extensive experiments demonstrate that our proposed framework outperforms other self-supervised methods by a large margin.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Let Your Video Listen to Your Music!

    cs.CV 2025-06 reject novelty 6.0 of 10

    MVAA aligns a video's motion peaks to music beats via keyframe re-timing and diffusion-based inpainting, aiming to preserve the original content while improving rhythmic synchronization.

  2. MV-Crafter: An Intelligent System for Music-guided Video Generation

    cs.HC 2025-04 conditional novelty 6.0 of 10

    MV-Crafter generates beat-synchronized music videos from music and a text theme by combining LLM-based scripting, diffusion video generation, and a dynamic beat-matching warping algorithm.

Pith tools