Pith. sign in

REVIEW 1 cited by

MotionMix: Weakly-Supervised Diffusion for Controllable Motion Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.11115 v3 pith:GTKOUAIL submitted 2024-01-20 cs.CV

classification cs.CV
keywords motiondiffusionmotionsgenerationmodelmotionmixannotatedcontrollable
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Controllable generation of 3D human motions becomes an important topic as the world embraces digital transformation. Existing works, though making promising progress with the advent of diffusion models, heavily rely on meticulously captured and annotated (e.g., text) high-quality motion corpus, a resource-intensive endeavor in the real world. This motivates our proposed MotionMix, a simple yet effective weakly-supervised diffusion model that leverages both noisy and unannotated motion sequences. Specifically, we separate the denoising objectives of a diffusion model into two stages: obtaining conditional rough motion approximations in the initial $T-T^*$ steps by learning the noisy annotated motions, followed by the unconditional refinement of these preliminary motions during the last $T^*$ steps using unannotated motions. Notably, though learning from two sources of imperfect data, our model does not compromise motion generation quality compared to fully supervised approaches that access gold data. Extensive experiments on several benchmarks demonstrate that our MotionMix, as a versatile framework, consistently achieves state-of-the-art performances on text-to-motion, action-to-motion, and music-to-dance tasks. Project page: https://nhathoang2002.github.io/MotionMix-page/

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. PackDiT: Joint Human Motion and Text Generation via Mutual Prompting

    cs.CV 2025-01 conditional novelty 5.0 of 10

    A two-transformer diffusion framework with mutual cross-attention claims the first diffusion-based joint human motion and text generation, with a reported HumanML3D text-to-motion FID of 0.106.

Pith tools