Pith. sign in

REVIEW 9 cited by

DreamPose: Fashion Image-to-Video Synthesis via Stable Diffusion

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.06025 v4 pith:SNSR26KB submitted 2023-04-12 cs.CV

classification cs.CV
keywords fashionmethodvideodiffusiondreamposehumanmodelposes
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We present DreamPose, a diffusion-based method for generating animated fashion videos from still images. Given an image and a sequence of human body poses, our method synthesizes a video containing both human and fabric motion. To achieve this, we transform a pretrained text-to-image model (Stable Diffusion) into a pose-and-image guided video synthesis model, using a novel fine-tuning strategy, a set of architectural changes to support the added conditioning signals, and techniques to encourage temporal consistency. We fine-tune on a collection of fashion videos from the UBC Fashion dataset. We evaluate our method on a variety of clothing styles and poses, and demonstrate that our method produces state-of-the-art results on fashion video animation.Video results are available on our project page.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MultiAnimate: A Unified Framework for Controllable Multi-Character Animation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A diffusion-based framework that animates multiple characters in one scene from separate reference images and pose sequences while preserving each character's identity.

  2. VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A zero-shot diffusion framework that inserts a reference object into a video with high-fidelity appearance preservation and precise key-point trajectory motion control.

  3. Dynamic Try-On: Taming Video Virtual Try-on with Dynamic Attention Mechanism

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A DiT-based video try-on framework that reuses the backbone as garment encoder and uses limb-aware dynamic attention to improve temporal consistency.

  4. Perturb-and-Revise: Flexible 3D Editing with Generative Trajectories

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Perturb-and-Revise edits 3D scenes by mixing a NeRF's trained parameters with random ones, running multi-view score distillation toward the edit prompt, and refining with identity-preserving gradients.

  5. Wan-Animate-2: Pushing the Application Boundaries of Character Animation

    cs.CV 2026-08 conditional novelty 5.0 of 10

    Wan-Animate-2 animates a reference character from a driving video in one end-to-end diffusion transformer, adds text-driven viewpoint control, and distills a real-time streaming variant.

  6. X-Dyna: Expressive Dynamic Human Image Animation

    cs.CV 2025-01 conditional novelty 5.0 of 10

    A diffusion-based pipeline that animates a single human image with pose, expression, and dynamic background effects from a driving video, outperforming prior methods on dynamic detail metrics.

  7. One-shot Human Motion Transfer via Occlusion-Robust Flow Prediction and Neural Texturing

    cs.CV 2024-12 conditional novelty 5.0 of 10

    Combining multi-scale flow warping with DensePose-based neural texture mapping improves transferred pose accuracy in one-shot human animation, while appearance preservation stays mixed.

  8. VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models

    cs.CV 2024-11 conditional novelty 5.0 of 10

    VBench++ is a benchmark that scores text-to-video and image-to-video models on 16 quality dimensions plus trustworthiness, reporting human-alignment correlations for each.

  9. Generative AI for Character Animation: A Comprehensive Survey of Techniques, Applications, and Future Directions

    cs.CV 2025-04 conditional novelty 3.0 of 10

    A comprehensive survey that unifies generative AI techniques for character animation across facial, gesture, motion, and 3D asset generation, with a shared taxonomy and resource list.

Pith tools