REVIEW 9 cited by
DreamPose: Fashion Image-to-Video Synthesis via Stable Diffusion
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We present DreamPose, a diffusion-based method for generating animated fashion videos from still images. Given an image and a sequence of human body poses, our method synthesizes a video containing both human and fabric motion. To achieve this, we transform a pretrained text-to-image model (Stable Diffusion) into a pose-and-image guided video synthesis model, using a novel fine-tuning strategy, a set of architectural changes to support the added conditioning signals, and techniques to encourage temporal consistency. We fine-tune on a collection of fashion videos from the UBC Fashion dataset. We evaluate our method on a variety of clothing styles and poses, and demonstrate that our method produces state-of-the-art results on fashion video animation.Video results are available on our project page.
Forward citations
Cited by 9 Pith papers
-
MultiAnimate: A Unified Framework for Controllable Multi-Character Animation
A diffusion-based framework that animates multiple characters in one scene from separate reference images and pose sequences while preserving each character's identity.
-
VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control
A zero-shot diffusion framework that inserts a reference object into a video with high-fidelity appearance preservation and precise key-point trajectory motion control.
-
Dynamic Try-On: Taming Video Virtual Try-on with Dynamic Attention Mechanism
A DiT-based video try-on framework that reuses the backbone as garment encoder and uses limb-aware dynamic attention to improve temporal consistency.
-
Perturb-and-Revise: Flexible 3D Editing with Generative Trajectories
Perturb-and-Revise edits 3D scenes by mixing a NeRF's trained parameters with random ones, running multi-view score distillation toward the edit prompt, and refining with identity-preserving gradients.
-
Wan-Animate-2: Pushing the Application Boundaries of Character Animation
Wan-Animate-2 animates a reference character from a driving video in one end-to-end diffusion transformer, adds text-driven viewpoint control, and distills a real-time streaming variant.
-
X-Dyna: Expressive Dynamic Human Image Animation
A diffusion-based pipeline that animates a single human image with pose, expression, and dynamic background effects from a driving video, outperforming prior methods on dynamic detail metrics.
-
One-shot Human Motion Transfer via Occlusion-Robust Flow Prediction and Neural Texturing
Combining multi-scale flow warping with DensePose-based neural texture mapping improves transferred pose accuracy in one-shot human animation, while appearance preservation stays mixed.
-
VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models
VBench++ is a benchmark that scores text-to-video and image-to-video models on 16 quality dimensions plus trustworthiness, reporting human-alignment correlations for each.
-
Generative AI for Character Animation: A Comprehensive Survey of Techniques, Applications, and Future Directions
A comprehensive survey that unifies generative AI techniques for character animation across facial, gesture, motion, and 3D asset generation, with a shared taxonomy and resource list.
Discussion (0). Continue with ORCID to comment.