A new large-scale triplet dataset and diffusion transformer model using coarse human masks deliver improved video virtual try-on quality and generalization in challenging real-world conditions.
arXiv preprint arXiv:2407.15886 , year=
11 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
fields
cs.CV 11roles
background 1polarities
background 1representative citing papers
FIT is a large-scale dataset of 1.13M try-on triplets with exact size data plus a synthetic generation pipeline that enables training of virtual try-on models capable of depicting realistic garment fit including ill-fit cases.
Durian introduces a dual-reference diffusion model trained via self-reconstruction on video frames to enable cross-identity attribute transfer in portrait animations, supporting multi-attribute composition and interpolation.
LPH-VTON uses a single denoising process with staged handover from structure-biased to texture-biased diffusion models to improve both geometric alignment and textural fidelity in virtual try-on.
Vanast produces coherent garment-transferred human animation videos from a single human image, garment images, and pose guidance video using synthetic triplet supervision and a Dual Module video diffusion transformer architecture.
RefTon is a flux-based virtual try-on method that uses unpaired reference images of the target garment on different people to guide texture and detail preservation in a streamlined person-to-person pipeline without body parsing or masks.
FDM-MFVT is a few-step mask-free virtual try-on diffusion model using OANO and IDT modules plus a new 30,000-pair MFVT dataset, claiming better efficiency and quality than baselines.
ModaFlow is a modality-aware flow matching framework for virtual try-on that uses visual embeddings for structural guidance, text embeddings with adaptive CFG, regularization losses, and stochastic mask sampling to achieve lower FID scores than prior methods.
FitVTON introduces a fit-aware virtual try-on model using text prompts for size control, auxiliary garment/body mask prediction, and texture rectification to achieve better sizing accuracy on diverse bodies than prior diffusion methods.
Introduces dual pose-image representation, cross-modal alignment, and iterative construction to improve prompt alignment and diversity in multi-person text-to-image generation.
Tstars-Tryon 1.0 is a deployed virtual try-on system claiming high robustness, photorealism, multi-reference flexibility, and near real-time speed for diverse fashion items.
citing papers explorer
-
TripVVT: A Large-Scale Triplet Dataset and a Coarse-Mask Baseline for In-the-Wild Video Virtual Try-On
A new large-scale triplet dataset and diffusion transformer model using coarse human masks deliver improved video virtual try-on quality and generalization in challenging real-world conditions.
-
FIT: A Large-Scale Dataset for Fit-Aware Virtual Try-On
FIT is a large-scale dataset of 1.13M try-on triplets with exact size data plus a synthetic generation pipeline that enables training of virtual try-on models capable of depicting realistic garment fit including ill-fit cases.
-
Durian: Dual Reference Image-Guided Portrait Animation with Attribute Transfer
Durian introduces a dual-reference diffusion model trained via self-reconstruction on video frames to enable cross-identity attribute transfer in portrait animations, supporting multi-attribute composition and interpolation.
-
LPH-VTON: Resolving the Structure-Texture Dilemma of Virtual Try-On via Latent Process Handover
LPH-VTON uses a single denoising process with staged handover from structure-biased to texture-biased diffusion models to improve both geometric alignment and textural fidelity in virtual try-on.
-
Vanast: Virtual Try-On with Human Image Animation via Synthetic Triplet Supervision
Vanast produces coherent garment-transferred human animation videos from a single human image, garment images, and pose guidance video using synthetic triplet supervision and a Dual Module video diffusion transformer architecture.
-
RefTon: Reference person shot assist virtual Try-on
RefTon is a flux-based virtual try-on method that uses unpaired reference images of the target garment on different people to guide texture and detail preservation in a streamlined person-to-person pipeline without body parsing or masks.
-
FDM-MFVT: Few-step Sampling Diffusion Model for Mask-Free Virtual Try-On
FDM-MFVT is a few-step mask-free virtual try-on diffusion model using OANO and IDT modules plus a new 30,000-pair MFVT dataset, claiming better efficiency and quality than baselines.
-
ModaFlow: Modality-Aware Flow Matching for High-Fidelity Virtual Try-On
ModaFlow is a modality-aware flow matching framework for virtual try-on that uses visual embeddings for structural guidance, text embeddings with adaptive CFG, regularization losses, and stochastic mask sampling to achieve lower FID scores than prior methods.
-
FitVTON: Fit-aware Virtual Try-On via Body-Garment Size Control
FitVTON introduces a fit-aware virtual try-on model using text prompts for size control, auxiliary garment/body mask prediction, and texture rectification to achieve better sizing accuracy on diverse bodies than prior diffusion methods.
-
Composing People Together: Iterative Pose-Image Generation for Multi-Person Interaction Scenes
Introduces dual pose-image representation, cross-modal alignment, and iterative construction to improve prompt alignment and diversity in multi-person text-to-image generation.
-
Tstars-Tryon 1.0: Robust and Realistic Virtual Try-On for Diverse Fashion Items
Tstars-Tryon 1.0 is a deployed virtual try-on system claiming high robustness, photorealism, multi-reference flexibility, and near real-time speed for diverse fashion items.