REVIEW 35 cited by
OOTDiffusion: Outfitting Fusion based Latent Diffusion for Controllable Virtual Try-on
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We present OOTDiffusion, a novel network architecture for realistic and controllable image-based virtual try-on (VTON). We leverage the power of pretrained latent diffusion models, designing an outfitting UNet to learn the garment detail features. Without a redundant warping process, the garment features are precisely aligned with the target human body via the proposed outfitting fusion in the self-attention layers of the denoising UNet. In order to further enhance the controllability, we introduce outfitting dropout to the training process, which enables us to adjust the strength of the garment features through classifier-free guidance. Our comprehensive experiments on the VITON-HD and Dress Code datasets demonstrate that OOTDiffusion efficiently generates high-quality try-on results for arbitrary human and garment images, which outperforms other VTON methods in both realism and controllability, indicating an impressive breakthrough in virtual try-on. Our source code is available at https://github.com/levihsu/OOTDiffusion.
Forward citations
Cited by 35 Pith papers
-
OmniVTON: Training-Free Universal Virtual Try-On
OmniVTON uses pretrained diffusion models with no training to transfer garments between people across shop and street scenes, and extends to multi-human try-on.
-
VTBench: Comprehensive Benchmark Suite Towards Real-World Virtual Try-on Models
VTBench is a multi-dimensional benchmark with novel unpaired metrics and human preference data for evaluating image-based virtual try-on models, though the human-alignment evidence is incomplete.
-
Incorporating Visual Correspondence into Diffusion Model for Virtual Try-On
SPM-Diff injects flow-warped garment point features into a diffusion model's self-attention to improve detail preservation in virtual try-on.
-
WearWow: Native 2K Multi-Garment Virtual Try-On via Adaptive Token Packing and Preference Alignment
WearWow generates native 2K multi-garment virtual try-on images without masks, using token packing plus dual preference rewards to preserve fabric texture.
-
Dress&Dance: Dress up and Dance as You Like It - Technical Preview
A video diffusion framework that unifies text, image, and video conditioning through attention to produce high-resolution virtual try-on videos with reference-driven motion.
-
FastFit: Accelerating Multi-Reference Virtual Try-On via Cacheable Diffusion Models
FastFit uses a cacheable diffusion UNet to compute multi-reference garment features once per generation, enabling about 3.5x faster multi-item virtual try-on with comparable or better fidelity.
-
FW-VTON: Flattening-and-Warping for Person-to-Person Virtual Try-on
FW-VTON reports state-of-the-art person-to-person virtual try-on results using a flattening, warping, and integration pipeline plus a new P2P-VTON dataset.
-
Video Virtual Try-on with Conditional Diffusion Transformer Inpainter
ViTI reformulates video virtual try-on as conditional video inpainting with a full 3D attention diffusion transformer, and reports the best VFID score on VVT (2.121).
-
Low-Barrier Dataset Collection with Real Human Body for Interactive Per-Garment Virtual Try-On
A per-garment virtual try-on pipeline that trains a GAN from a two-minute real-human video capture and uses a hybrid pose-plus-DensePose input to synthesize the garment with accurate alignment.
-
Pursuing Temporal-Consistent Video Virtual Try-On via Dynamic Pose Interaction
DPIDM, a diffusion model with pose-aware spatial and temporal attention plus a temporal attention loss, reports state-of-the-art video virtual try-on and cuts VFID on VVT from 1.280 to 0.506.
-
3DV-TON: Textured 3D-Guided Consistent Video Try-on via Diffusion Models
A diffusion video try-on framework that uses animated textured 3D meshes as frame-level guidance, plus a new high-resolution benchmark, achieves stronger temporal consistency and garment fidelity than two released baselines.
-
Any2AnyTryon: Leveraging Adaptive Position Embeddings for Versatile Virtual Clothing Tasks
A single DiT-based model with adaptive position embeddings performs virtual try-on, garment reconstruction, model-free try-on, and layered try-on from text and variable-size image inputs.
-
DreamFit: Garment-Centric Human Generation via a Lightweight Anything-Dressing Encoder
DreamFit generates human images from a garment reference and text by encoding the reference through LoRA-activated layers of a frozen Stable Diffusion UNet and injecting features with adaptive attention.
-
DiffusionTrend: A Minimalist Approach to Virtual Fashion Try-On
A training-free virtual try-on pipeline that blends DDIM-inverted garment latents into masked model latents, guided by a lightweight CNN apparel mask.
-
FashionComposer: Compositional Fashion Image Generation
A single diffusion framework composes multiple garment and face references into one fashion image using an asset library and subject-binding attention.
-
IGR: Improving Diffusion Model for Garment Restoration from Person Image
IGR restores a clean garment image from a person photo using Stable Diffusion, dual extractors, attention fusion blocks, and a VITON-to-GarmRe fine-tuning strategy, beating TryOffDiff on the reported benchmarks.
-
Dynamic Try-On: Taming Video Virtual Try-on with Dynamic Attention Mechanism
A DiT-based video try-on framework that reuses the backbone as garment encoder and uses limb-aware dynamic attention to improve temporal consistency.
-
Learning Flow Fields in Attention for Controllable Person Image Generation
Leffa converts attention maps into flow fields that warp the reference image, then uses the mismatch with the target as a loss, improving fine-detail fidelity in virtual try-on and pose transfer.
-
UniReal: Universal Image Generation and Editing via Learning Real-world Dynamics
Treating image editing and generation as discontinuous video generation, UniReal trains one 5B diffusion transformer on video frame pairs and labeled datasets to handle diverse image tasks in a unified framework.
-
DreamDance: Animating Human Images by Enriching 3D Geometry Cues from 2D Poses
A two-stage diffusion pipeline that enriches 2D pose guidance with generated depth and normal maps to achieve state-of-the-art human image animation.
-
TKG-DM: Training-free Chroma Key Content Generation Diffusion Model
Adjusting the mean of specific channels in the initial noise of Stable Diffusion produces foreground objects on a uniform, user-selected chroma key background without any fine-tuning.
-
Try-On-Adapter: A Simple and Flexible Try-On Paradigm
A diffusion-based adapter performs virtual try-on as outpainting from a reference face and garment, reporting FID 5.56 and 7.23 on VITON-HD.
-
FitDiT: Advancing the Authentic Garment Details for High-fidelity Virtual Try-on
FitDiT applies a customized Diffusion Transformer to image-based virtual try-on, adding a garment feature evolution stage, a frequency-domain loss, and a relaxed mask strategy to improve texture and size fidelity.
-
Virtual Fitting Room: Generating Arbitrarily Long Videos of Virtual Try-On from a Single Image -- Technical Preview
VFR generates minute-long virtual try-on videos by auto-regressively chaining short diffusion-generated segments that are kept consistent with a 360-degree anchor video of the user.
-
JCo-MVTON: Jointly Controllable Multi-Modal Diffusion Transformer for Mask-Free Virtual Try-on
A mask-free diffusion transformer for virtual try-on, trained with a self-generated and manually curated triplet dataset, achieves state-of-the-art scores on DressCode and competitive results on VITON-HD.
-
Real-Time Per-Garment Virtual Try-On with Temporal Consistency for Loose-Fitting Garments
A per-garment virtual try-on method for loose-fitting garments uses a garment-invariant pose representation and a recurrent ConvLSTM synthesis network to achieve temporally smoother try-on video at about 10 fps.
-
Insert Anything: Image Insertion via In-Context Editing in DiT
Insert Anything is a single model fine-tuned on 159,908 prompt-image pairs that performs mask- or text-guided insertion of people, objects, and garments from reference images into target scenes.
-
CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation
CatV2TON unifies image and video virtual try-on in one diffusion transformer, using temporal garment-person concatenation and clip-based inference with AdaCN for long, consistent try-on videos.
-
RealVVT: Towards Photorealistic Video Virtual Try-on via Spatio-Temporal Consistency
A diffusion-based framework improves spatial and temporal consistency of clothes in virtual try-on videos, reporting the best FID/KID and several video metrics on four public datasets.
-
1-2-1: Renaissance of Single-Network Paradigm for Virtual Try-On
A single-network virtual try-on model with modality-specific normalization and shared attention matches or beats dual-network reference-based models on image and video try-on benchmarks.
-
Consistent Human Image and Video Generation with Spatially Conditioned Diffusion
Spatially conditioning a diffusion model by concatenating a reference human image with the noisy target, and adding causal self-attention, improves appearance consistency in human image and video animation.
-
AnyDressing: Customizable Multi-Garment Virtual Dressing via Latent Diffusion Models
AnyDressing combines a parallel garment encoder with localized attention to generate a person wearing multiple specified garments from a text prompt.
-
TalkFashion: Intelligent Virtual Try-On Assistant Based on Multimodal Large Language Model
TalkFashion, a text-driven virtual try-on assistant, reports better semantic consistency and visual quality than four baselines on VITON-HD by combining an LLM router, catalog matching, and automatic mask generation.
-
DiffFit: Disentangled Garment Warping and Texture Refinement for Virtual Try-On
DiffFit synthesizes virtual try-on images by separately warping the garment geometry and then refining texture with a conditional diffusion model, reporting SOTA metrics on VITON-HD and DressCode but with inconsistent...
-
MFP-VTON: Enhancing Mask-Free Person-to-Person Virtual Try-On via Diffusion Transformer
A mask-free person-to-person virtual try-on model built on FLUX-Fill-dev, trained with pseudo data generated by IDM and a Focus Attention loss.
Discussion (0). Continue with ORCID to comment.