REVIEW 12 cited by
OOTDiffusion: Outfitting Fusion based Latent Diffusion for Controllable Virtual Try-on
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We present OOTDiffusion, a novel network architecture for realistic and controllable image-based virtual try-on (VTON). We leverage the power of pretrained latent diffusion models, designing an outfitting UNet to learn the garment detail features. Without a redundant warping process, the garment features are precisely aligned with the target human body via the proposed outfitting fusion in the self-attention layers of the denoising UNet. In order to further enhance the controllability, we introduce outfitting dropout to the training process, which enables us to adjust the strength of the garment features through classifier-free guidance. Our comprehensive experiments on the VITON-HD and Dress Code datasets demonstrate that OOTDiffusion efficiently generates high-quality try-on results for arbitrary human and garment images, which outperforms other VTON methods in both realism and controllability, indicating an impressive breakthrough in virtual try-on. Our source code is available at https://github.com/levihsu/OOTDiffusion.
Forward citations
Cited by 12 Pith papers
-
OmniVTON: Training-Free Universal Virtual Try-On
OmniVTON uses pretrained diffusion models with no training to transfer garments between people across shop and street scenes, and extends to multi-human try-on.
-
VTBench: Comprehensive Benchmark Suite Towards Real-World Virtual Try-on Models
VTBench is a multi-dimensional benchmark with novel unpaired metrics and human preference data for evaluating image-based virtual try-on models, though the human-alignment evidence is incomplete.
-
Incorporating Visual Correspondence into Diffusion Model for Virtual Try-On
SPM-Diff injects flow-warped garment point features into a diffusion model's self-attention to improve detail preservation in virtual try-on.
-
WearWow: Native 2K Multi-Garment Virtual Try-On via Adaptive Token Packing and Preference Alignment
WearWow generates native 2K multi-garment virtual try-on images without masks, using token packing plus dual preference rewards to preserve fabric texture.
-
Dress&Dance: Dress up and Dance as You Like It - Technical Preview
A video diffusion framework that unifies text, image, and video conditioning through attention to produce high-resolution virtual try-on videos with reference-driven motion.
-
FW-VTON: Flattening-and-Warping for Person-to-Person Virtual Try-on
FW-VTON reports state-of-the-art person-to-person virtual try-on results using a flattening, warping, and integration pipeline plus a new P2P-VTON dataset.
-
Video Virtual Try-on with Conditional Diffusion Transformer Inpainter
ViTI reformulates video virtual try-on as conditional video inpainting with a full 3D attention diffusion transformer, and reports the best VFID score on VVT (2.121).
-
Low-Barrier Dataset Collection with Real Human Body for Interactive Per-Garment Virtual Try-On
A per-garment virtual try-on pipeline that trains a GAN from a two-minute real-human video capture and uses a hybrid pose-plus-DensePose input to synthesize the garment with accurate alignment.
-
Pursuing Temporal-Consistent Video Virtual Try-On via Dynamic Pose Interaction
DPIDM, a diffusion model with pose-aware spatial and temporal attention plus a temporal attention loss, reports state-of-the-art video virtual try-on and cuts VFID on VVT from 1.280 to 0.506.
-
Real-Time Per-Garment Virtual Try-On with Temporal Consistency for Loose-Fitting Garments
A per-garment virtual try-on method for loose-fitting garments uses a garment-invariant pose representation and a recurrent ConvLSTM synthesis network to achieve temporally smoother try-on video at about 10 fps.
-
TalkFashion: Intelligent Virtual Try-On Assistant Based on Multimodal Large Language Model
TalkFashion, a text-driven virtual try-on assistant, reports better semantic consistency and visual quality than four baselines on VITON-HD by combining an LLM router, catalog matching, and automatic mask generation.
-
DiffFit: Disentangled Garment Warping and Texture Refinement for Virtual Try-On
DiffFit synthesizes virtual try-on images by separately warping the garment geometry and then refining texture with a conditional diffusion model, reporting SOTA metrics on VITON-HD and DressCode but with inconsistent...
Discussion (0). Sign in to comment.