Pith. sign in

REVIEW 4 cited by

HumanDreamer: Generating Controllable Human-Motion Videos via Decoupled Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.24026 v2 pith:ULFAT2M5 submitted 2025-03-31 cs.CV

classification cs.CV
keywords human-motiongenerationposesvideosposevideocontroldataset
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Human-motion video generation has been a challenging task, primarily due to the difficulty inherent in learning human body movements. While some approaches have attempted to drive human-centric video generation explicitly through pose control, these methods typically rely on poses derived from existing videos, thereby lacking flexibility. To address this, we propose HumanDreamer, a decoupled human video generation framework that first generates diverse poses from text prompts and then leverages these poses to generate human-motion videos. Specifically, we propose MotionVid, the largest dataset for human-motion pose generation. Based on the dataset, we present MotionDiT, which is trained to generate structured human-motion poses from text prompts. Besides, a novel LAMA loss is introduced, which together contribute to a significant improvement in FID by 62.4%, along with respective enhancements in R-precision for top1, top2, and top3 by 41.8%, 26.3%, and 18.3%, thereby advancing both the Text-to-Pose control accuracy and FID metrics. Our experiments across various Pose-to-Video baselines demonstrate that the poses generated by our method can produce diverse and high-quality human-motion videos. Furthermore, our model can facilitate other downstream tasks, such as pose sequence prediction and 2D-3D motion lifting.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. GigaWorld-1: A Roadmap to Build World Models for Robot Policy Evaluation

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Long-horizon action-faithful consistency, not short-term visual realism, dominates world-model reliability for robot policy evaluation; GigaWorld-1 implements that roadmap and gains 14.9% on evaluator-alignment metrics.

  2. WonderFree: Enhancing Novel View Quality and Cross-View Consistency for 3D Scene Exploration

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A pipeline that restores corrupted novel-view videos with a video diffusion model and jointly denoises multiple viewpoints to improve 3D scene exploration from a single image.

  3. EmbodieDreamer: Advancing Real2Sim2Real Transfer for Policy Training via Embodied World Modeling

    cs.RO 2025-07 conditional novelty 5.0 of 10

    A Real2Sim2Real framework that aligns simulator dynamics via differentiable parameter fitting and renders photorealistic policy-training videos with a diffusion model, improving real-world manipulation success.

  4. Toward Rich Video Human-Motion2D Generation

    cs.CV 2025-06 reject novelty 4.0 of 10

    A new 150K-video 2D skeleton dataset with text captions and a diffusion model for single- and double-character motion generation, though the claimed FID-rewarded RL training is misrepresented.

Pith tools