Pith. sign in

REVIEW 11 cited by

DwNet: Dense warp-based network for pose-guided human video generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1910.09139 v1 pith:HAGA3BSR submitted 2019-10-21 cs.CV cs.LG

classification cs.CVcs.LG
keywords videogenerationhumandensedwnetfashiongeneratedimage
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Generation of realistic high-resolution videos of human subjects is a challenging and important task in computer vision. In this paper, we focus on human motion transfer - generation of a video depicting a particular subject, observed in a single image, performing a series of motions exemplified by an auxiliary (driving) video. Our GAN-based architecture, DwNet, leverages dense intermediate pose-guided representation and refinement process to warp the required subject appearance, in the form of the texture, from a source image into a desired pose. Temporal consistency is maintained by further conditioning the decoding process within a GAN on the previously generated frame. In this way a video is generated in an iterative and recurrent fashion. We illustrate the efficacy of our approach by showing state-of-the-art quantitative and qualitative performance on two benchmark datasets: TaiChi and Fashion Modeling. The latter is collected by us and will be made publicly available to the community.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Forwardrobe: Garment-Aware Gaussian Avatars from a Single Image

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Forwardrobe reconstructs from a single image an animatable Gaussian avatar whose clothes are a separable, editable, and transferable 3D garment asset.

  2. PoseGuard: Pose-Guided Generation with Safety Guardrails

    cs.CR 2025-08 unverdicted novelty 6.0 of 10

    PoseGuard degrades output quality of pose-guided video generators for unsafe poses while preserving fidelity for benign poses, using LoRA-based safety alignment.

  3. Can Pose Transfer Models Generate Realistic Human Motion?

    cs.CV 2025-01 conditional novelty 6.0 of 10

    State-of-the-art pose transfer models produce videos that human viewers can correctly identify only 42.92% of the time when actions and identities are out of distribution.

  4. Motion-X++: A Large-Scale Multimodal 3D Whole-body Human Motion Dataset

    cs.CV 2025-01 conditional novelty 6.0 of 10

    Motion-X++ provides 19.5M 3D whole-body pose annotations across 120.5K sequences with text, audio, video, and motion modalities.

  5. From Rigging to Waving: 3D-Guided Diffusion for Natural Animation of Hand-Drawn Characters

    cs.GR 2025-09 conditional novelty 5.0 of 10

    A hybrid skeletal-diffusion pipeline animates hand-drawn characters from a single image with 3D guidance, inpainting coarse renders and injecting secondary dynamics via latent blending and hair-body separation.

  6. Animate-X++: Universal Character Image Animation with Dynamic Backgrounds

    cs.CV 2025-08 conditional novelty 5.0 of 10

    Animate-X++ turns cartoon images into pose-driven animations with text-controlled moving backgrounds, claiming state-of-the-art results on a new synthetic anthropomorphic benchmark.

  7. MVHumanNet++: A Large-scale Dataset of Multi-view Daily Dressing Human Captures with Richer Annotations for 3D Human Digitization

    cs.CV 2025-05 conditional novelty 5.0 of 10

    MVHumanNet++ is a very large multi-view human dataset whose scale and annotations improve downstream 3D human reconstruction and generation models.

  8. Animate Anyone 2: High-Fidelity Character Image Animation with Environment Affordance

    cs.CV 2025-02 conditional novelty 5.0 of 10

    A diffusion-based character animation method that conditions on environment, object, and depth signals to produce videos where characters interact naturally with their surroundings.

  9. RAIN: Real-time Animation of Infinite Video Stream

    cs.CV 2024-12 conditional novelty 5.0 of 10

    RAIN makes real-time character animation and video style transfer practical on consumer GPUs by grouping streamed frames into shared noise levels and running cross-noise-level temporal attention.

  10. Consistent Human Image and Video Generation with Spatially Conditioned Diffusion

    cs.CV 2024-12 conditional novelty 5.0 of 10

    Spatially conditioning a diffusion model by concatenating a reference human image with the noisy target, and adding causal self-attention, improves appearance consistency in human image and video animation.

  11. Identity-Preserving Pose-Guided Character Animation via Facial Landmarks Transformation

    cs.CV 2024-12 conditional novelty 4.0 of 10

    FLT improves identity preservation in landmark-conditioned character animation by combining reference face shape with driving expressions in 3D and re-rendering landmarks.

Pith tools