Pith. sign in

REVIEW 2 cited by

Fashion-VDM: Video Diffusion Model for Virtual Try-On

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.00225 v2 pith:WZ6A7V27 submitted 2024-10-31 cs.CV

classification cs.CV
keywords videotry-onvirtualfashion-vdmgarmentpersondiffusiongiven
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present Fashion-VDM, a video diffusion model (VDM) for generating virtual try-on videos. Given an input garment image and person video, our method aims to generate a high-quality try-on video of the person wearing the given garment, while preserving the person's identity and motion. Image-based virtual try-on has shown impressive results; however, existing video virtual try-on (VVT) methods are still lacking garment details and temporal consistency. To address these issues, we propose a diffusion-based architecture for video virtual try-on, split classifier-free guidance for increased control over the conditioning inputs, and a progressive temporal training strategy for single-pass 64-frame, 512px video generation. We also demonstrate the effectiveness of joint image-video training for video try-on, especially when video data is limited. Our qualitative and quantitative experiments show that our approach sets the new state-of-the-art for video virtual try-on. For additional results, visit our project page: https://johannakarras.github.io/Fashion-VDM.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Low-Barrier Dataset Collection with Real Human Body for Interactive Per-Garment Virtual Try-On

    cs.GR 2025-06 conditional novelty 6.0 of 10

    A per-garment virtual try-on pipeline that trains a GAN from a two-minute real-human video capture and uses a hybrid pose-plus-DensePose input to synthesize the garment with accurate alignment.

  2. Real-Time Per-Garment Virtual Try-On with Temporal Consistency for Loose-Fitting Garments

    cs.GR 2025-06 conditional novelty 5.0 of 10

    A per-garment virtual try-on method for loose-fitting garments uses a garment-invariant pose representation and a recurrent ConvLSTM synthesis network to achieve temporally smoother try-on video at about 10 fps.

Pith tools