Pith. sign in

REVIEW 5 cited by

MagicPose4D: Crafting Articulated Models with Appearance and Motion Control

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.14017 v4 pith:J3EBSQEC submitted 2024-05-22 cs.CV

classification cs.CV
keywords motionmagicpose4dskeletonaccuratecontentcontrolgenerationmethods
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

With the success of 2D and 3D visual generative models, there is growing interest in generating 4D content. Existing methods primarily rely on text prompts to produce 4D content, but they often fall short of accurately defining complex or rare motions. To address this limitation, we propose MagicPose4D, a novel framework for refined control over both appearance and motion in 4D generation. Unlike current 4D generation methods, MagicPose4D accepts monocular videos or mesh sequences as motion prompts, enabling precise and customizable motion control. MagicPose4D comprises two key modules: (i) Dual-Phase 4D Reconstruction Module, which operates in two phases. The first phase focuses on capturing the model's shape using accurate 2D supervision and less accurate but geometrically informative 3D pseudo-supervision without imposing skeleton constraints. The second phase extracts the 3D motion (skeleton poses) using more accurate pseudo-3D supervision, obtained in the first phase and introduces kinematic chain-based skeleton constraints to ensure physical plausibility. Additionally, we propose a Global-local Chamfer loss that aligns the overall distribution of predicted mesh vertices with the supervision while maintaining part-level alignment without extra annotations. (ii) Cross-category Motion Transfer Module, which leverages the extracted motion from the 4D reconstruction module and uses a kinematic-chain-based skeleton to achieve cross-category motion transfer. It ensures smooth transitions between frames through dynamic rigidity, facilitating robust generalization without additional training. Through extensive experiments, we demonstrate that MagicPose4D significantly improves the accuracy and consistency of 4D content generation, outperforming existing methods in various benchmarks.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. UniMC: Taming Diffusion Transformer for Unified Keypoint-Guided Multi-Class Image Generation

    cs.CV 2025-07 conditional novelty 7.0 of 10

    UniMC uses tokenized instance conditions (class, box, keypoints) and a timestep-aware modulator in a DiT backbone to control multi-class human and animal image generation, trained and evaluated on the new HAIG-2.9M dataset.

  2. MorphGS: Morphology-Adaptive Articulated 3D Motion Transfer from Videos

    cs.CV 2026-01 conditional novelty 6.0 of 10

    MorphGS retargets motion from a monocular video onto a rigged 3D character by optimizing target morphology and pose with image-space losses, without 3D source reconstruction or parametric templates.

  3. Behave Your Motion: Habit-preserved Cross-category Animal Motion Transfer

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A habit-preserving VQ-VAE with category-specific habit encoders and LLM-based habit retrieval transfers animal motions across species, including zero-shot to unseen species, validated on a new skeletal quadruped dataset.

  4. PhysRig: Differentiable Physics-Based Skinning and Rigging Framework for Realistic Articulated Object Modeling

    cs.CV 2025-06 reject novelty 5.0 of 10

    PhysRig animates articulated 3D objects by simulating them as deformable soft bodies driven by an embedded skeleton, and learns the material and motion parameters with a differentiable physics simulator.

  5. Drive Any Mesh: 4D Latent Diffusion for Mesh Deformation from Video

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A video-conditioned latent diffusion model generates mesh vertex trajectories that deform an input 3D asset into render-ready 4D animations.

Pith tools