Pith. sign in

REVIEW 17 cited by

MIMO: Controllable Character Video Synthesis with Spatial Decomposed Modeling

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.16160 v2 pith:WBQ2B3MU submitted 2024-09-24 cs.CV

classification cs.CV
keywords scenevideocharacterspatialsynthesischaracterscodemodeling
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Character video synthesis aims to produce realistic videos of animatable characters within lifelike scenes. As a fundamental problem in the computer vision and graphics community, 3D works typically require multi-view captures for per-case training, which severely limits their applicability of modeling arbitrary characters in a short time. Recent 2D methods break this limitation via pre-trained diffusion models, but they struggle for pose generality and scene interaction. To this end, we propose MIMO, a novel framework which can not only synthesize character videos with controllable attributes (i.e., character, motion and scene) provided by simple user inputs, but also simultaneously achieve advanced scalability to arbitrary characters, generality to novel 3D motions, and applicability to interactive real-world scenes in a unified framework. The core idea is to encode the 2D video to compact spatial codes, considering the inherent 3D nature of video occurrence. Concretely, we lift the 2D frame pixels into 3D using monocular depth estimators, and decompose the video clip to three spatial components (i.e., main human, underlying scene, and floating occlusion) in hierarchical layers based on the 3D depth. These components are further encoded to canonical identity code, structured motion code and full scene code, which are utilized as control signals of synthesis process. The design of spatial decomposed modeling enables flexible user control, complex motion expression, as well as 3D-aware synthesis for scene interactions. Experimental results demonstrate effectiveness and robustness of the proposed method.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 17 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Perception-as-Control: Fine-grained Controllable Image Animation with 3D-aware Motion Representation

    cs.CV 2025-01 conditional novelty 7.0 of 10

    A diffusion-based image animation method that uses rendered colored spheres and a checkerboard world envelope as 3D-aware motion control signals, enabling joint camera and object motion control with fine granularity.

  2. MultiAnimate: A Unified Framework for Controllable Multi-Character Animation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A diffusion-based framework that animates multiple characters in one scene from separate reference images and pose sequences while preserving each character's identity.

  3. Precise Action-to-Video Generation Through Visual Action Prompts

    cs.CV 2025-08 conditional novelty 6.0 of 10

    Skeleton-based visual action prompts give precise, cross-domain action control for video generation of human and robot interactions.

  4. PERSONA: Personalized Whole-Body 3D Avatar with Pose-Driven Deformations from a Single Image

    cs.CV 2025-08 conditional novelty 6.0 of 10

    PERSONA creates a personalized 3D avatar from one image by using diffusion-generated pose-rich videos to train a 3D Gaussian avatar with balanced sampling and geometry-weighted optimization.

  5. Vid-CamEdit: Video Camera Trajectory Editing with Generative Rendering from Estimated Geometry

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Vid-CamEdit re-synthesizes monocular videos along user-defined camera paths by conditioning a video diffusion model on 2D flows derived from estimated 3D geometry, without training on multi-view video data.

  6. DreamActor-H1: High-Fidelity Human-Product Demonstration Video Generation via Motion-designed Diffusion Transformers

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A diffusion transformer model generates human-product demonstration videos from paired human and product images while preserving both identities through masked cross-attention and motion template guidance.

  7. HunyuanVideo-HOMA: Generic Human-Object Interaction in Multimodal Driven Human Animation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    HunyuanVideo-HOMA generates human-object interaction videos from weak, sparse inputs: one arm pose, an object center dot, a human photo, and an object photo.

  8. DreamDance: Animating Character Art via Inpainting Stable Gaussian Worlds

    cs.CV 2025-05 conditional novelty 6.0 of 10

    DreamDance animates a single character artwork by reconstructing its background as a 3D Gaussian scene and then inpainting the animated character into the rendered video.

  9. AniCrafter: Customizing Realistic Human-Centric Animation via Avatar-Background Conditioning in Video Diffusion Models

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A diffusion model animates a character into arbitrary dynamic backgrounds by conditioning on a rendered 3D-avatar video, reframing open-domain animation as a restoration problem.

  10. AnyCharV: Bootstrap Controllable Character Video Generation with Fine-to-Coarse Guidance

    cs.CV 2025-02 conditional novelty 6.0 of 10

    AnyCharV is a two-stage diffusion method that places a reference character into a target video scene using pose and mask guidance, with a self-boosting stage that improves identity preservation.

  11. ManiVideo: Generating Hand-Object Manipulation Video with Dexterous and Generalizable Grasping

    cs.CV 2024-12 conditional novelty 6.0 of 10

    ManiVideo generates bimanual hand-object manipulation videos conditioned on 3D motion sequences, using a multi-layer occlusion representation and Objaverse-based training to improve 3D consistency and object generalization.

  12. AniGS: Animatable Gaussian Avatar from a Single Image with Inconsistent Gaussian Reconstruction

    cs.CV 2024-12 conditional novelty 6.0 of 10

    AniGS produces an animatable 3D avatar from a single image by synthesizing multi-view canonical images and normals with a video diffusion model and reconstructing them via 4D Gaussian Splatting.

  13. DreamDance: Animating Human Images by Enriching 3D Geometry Cues from 2D Poses

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A two-stage diffusion pipeline that enriches 2D pose guidance with generated depth and normal maps to achieve state-of-the-art human image animation.

  14. PersonaCraft: Personalized and Controllable Full-Body Multi-Human Scene Generation Using Occlusion-Aware 3D-Conditioned Diffusion

    cs.CV 2024-11 conditional novelty 6.0 of 10

    PersonaCraft adds SMPLx depth and normal conditioning, occlusion boundary enhancement, and occlusion-aware classifier-free guidance to diffusion models, enabling controllable multi-person images that preserve both fac...

  15. Drive Any Mesh: 4D Latent Diffusion for Mesh Deformation from Video

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A video-conditioned latent diffusion model generates mesh vertex trajectories that deform an input 3D asset into render-ready 4D animations.

  16. OmniV2V: Versatile Video Generation and Editing via Dynamic Content Manipulation

    cs.CV 2025-06 conditional novelty 5.0 of 10

    OmniV2V is one diffusion-transformer model that performs eight video generation and editing tasks by combining mask, pose, image, and text-instruction conditions.

  17. Animate Anyone 2: High-Fidelity Character Image Animation with Environment Affordance

    cs.CV 2025-02 conditional novelty 5.0 of 10

    A diffusion-based character animation method that conditions on environment, object, and depth signals to produce videos where characters interact naturally with their surroundings.

Pith tools