Pith. sign in

REVIEW 35 cited by

OOTDiffusion: Outfitting Fusion based Latent Diffusion for Controllable Virtual Try-on

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.01779 v2 pith:DR647JWT submitted 2024-03-04 cs.CV

classification cs.CV
keywords garmentootdiffusionoutfittingtry-onfeaturesvirtualcodecontrollability
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We present OOTDiffusion, a novel network architecture for realistic and controllable image-based virtual try-on (VTON). We leverage the power of pretrained latent diffusion models, designing an outfitting UNet to learn the garment detail features. Without a redundant warping process, the garment features are precisely aligned with the target human body via the proposed outfitting fusion in the self-attention layers of the denoising UNet. In order to further enhance the controllability, we introduce outfitting dropout to the training process, which enables us to adjust the strength of the garment features through classifier-free guidance. Our comprehensive experiments on the VITON-HD and Dress Code datasets demonstrate that OOTDiffusion efficiently generates high-quality try-on results for arbitrary human and garment images, which outperforms other VTON methods in both realism and controllability, indicating an impressive breakthrough in virtual try-on. Our source code is available at https://github.com/levihsu/OOTDiffusion.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 35 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. OmniVTON: Training-Free Universal Virtual Try-On

    cs.CV 2025-07 conditional novelty 7.0 of 10

    OmniVTON uses pretrained diffusion models with no training to transfer garments between people across shop and street scenes, and extends to multi-human try-on.

  2. VTBench: Comprehensive Benchmark Suite Towards Real-World Virtual Try-on Models

    cs.CV 2025-05 conditional novelty 7.0 of 10

    VTBench is a multi-dimensional benchmark with novel unpaired metrics and human preference data for evaluating image-based virtual try-on models, though the human-alignment evidence is incomplete.

  3. Incorporating Visual Correspondence into Diffusion Model for Virtual Try-On

    cs.CV 2025-05 conditional novelty 7.0 of 10

    SPM-Diff injects flow-warped garment point features into a diffusion model's self-attention to improve detail preservation in virtual try-on.

  4. WearWow: Native 2K Multi-Garment Virtual Try-On via Adaptive Token Packing and Preference Alignment

    cs.CV 2026-07 conditional novelty 6.0 of 10

    WearWow generates native 2K multi-garment virtual try-on images without masks, using token packing plus dual preference rewards to preserve fabric texture.

  5. Dress&Dance: Dress up and Dance as You Like It - Technical Preview

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A video diffusion framework that unifies text, image, and video conditioning through attention to produce high-resolution virtual try-on videos with reference-driven motion.

  6. FastFit: Accelerating Multi-Reference Virtual Try-On via Cacheable Diffusion Models

    cs.CV 2025-08 conditional novelty 6.0 of 10

    FastFit uses a cacheable diffusion UNet to compute multi-reference garment features once per generation, enabling about 3.5x faster multi-item virtual try-on with comparable or better fidelity.

  7. FW-VTON: Flattening-and-Warping for Person-to-Person Virtual Try-on

    cs.CV 2025-07 conditional novelty 6.0 of 10

    FW-VTON reports state-of-the-art person-to-person virtual try-on results using a flattening, warping, and integration pipeline plus a new P2P-VTON dataset.

  8. Video Virtual Try-on with Conditional Diffusion Transformer Inpainter

    cs.CV 2025-06 conditional novelty 6.0 of 10

    ViTI reformulates video virtual try-on as conditional video inpainting with a full 3D attention diffusion transformer, and reports the best VFID score on VVT (2.121).

  9. Low-Barrier Dataset Collection with Real Human Body for Interactive Per-Garment Virtual Try-On

    cs.GR 2025-06 conditional novelty 6.0 of 10

    A per-garment virtual try-on pipeline that trains a GAN from a two-minute real-human video capture and uses a hybrid pose-plus-DensePose input to synthesize the garment with accurate alignment.

  10. Pursuing Temporal-Consistent Video Virtual Try-On via Dynamic Pose Interaction

    cs.CV 2025-05 conditional novelty 6.0 of 10

    DPIDM, a diffusion model with pose-aware spatial and temporal attention plus a temporal attention loss, reports state-of-the-art video virtual try-on and cuts VFID on VVT from 1.280 to 0.506.

  11. 3DV-TON: Textured 3D-Guided Consistent Video Try-on via Diffusion Models

    cs.CV 2025-04 conditional novelty 6.0 of 10

    A diffusion video try-on framework that uses animated textured 3D meshes as frame-level guidance, plus a new high-resolution benchmark, achieves stronger temporal consistency and garment fidelity than two released baselines.

  12. Any2AnyTryon: Leveraging Adaptive Position Embeddings for Versatile Virtual Clothing Tasks

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A single DiT-based model with adaptive position embeddings performs virtual try-on, garment reconstruction, model-free try-on, and layered try-on from text and variable-size image inputs.

  13. DreamFit: Garment-Centric Human Generation via a Lightweight Anything-Dressing Encoder

    cs.CV 2024-12 conditional novelty 6.0 of 10

    DreamFit generates human images from a garment reference and text by encoding the reference through LoRA-activated layers of a frozen Stable Diffusion UNet and injecting features with adaptive attention.

  14. DiffusionTrend: A Minimalist Approach to Virtual Fashion Try-On

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A training-free virtual try-on pipeline that blends DDIM-inverted garment latents into masked model latents, guided by a lightweight CNN apparel mask.

  15. FashionComposer: Compositional Fashion Image Generation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A single diffusion framework composes multiple garment and face references into one fashion image using an asset library and subject-binding attention.

  16. IGR: Improving Diffusion Model for Garment Restoration from Person Image

    cs.CV 2024-12 conditional novelty 6.0 of 10

    IGR restores a clean garment image from a person photo using Stable Diffusion, dual extractors, attention fusion blocks, and a VITON-to-GarmRe fine-tuning strategy, beating TryOffDiff on the reported benchmarks.

  17. Dynamic Try-On: Taming Video Virtual Try-on with Dynamic Attention Mechanism

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A DiT-based video try-on framework that reuses the backbone as garment encoder and uses limb-aware dynamic attention to improve temporal consistency.

  18. Learning Flow Fields in Attention for Controllable Person Image Generation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Leffa converts attention maps into flow fields that warp the reference image, then uses the mismatch with the target as a loss, improving fine-detail fidelity in virtual try-on and pose transfer.

  19. UniReal: Universal Image Generation and Editing via Learning Real-world Dynamics

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Treating image editing and generation as discontinuous video generation, UniReal trains one 5B diffusion transformer on video frame pairs and labeled datasets to handle diverse image tasks in a unified framework.

  20. DreamDance: Animating Human Images by Enriching 3D Geometry Cues from 2D Poses

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A two-stage diffusion pipeline that enriches 2D pose guidance with generated depth and normal maps to achieve state-of-the-art human image animation.

  21. TKG-DM: Training-free Chroma Key Content Generation Diffusion Model

    cs.CV 2024-11 conditional novelty 6.0 of 10

    Adjusting the mean of specific channels in the initial noise of Stable Diffusion produces foreground objects on a uniform, user-selected chroma key background without any fine-tuning.

  22. Try-On-Adapter: A Simple and Flexible Try-On Paradigm

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A diffusion-based adapter performs virtual try-on as outpainting from a reference face and garment, reporting FID 5.56 and 7.23 on VITON-HD.

  23. FitDiT: Advancing the Authentic Garment Details for High-fidelity Virtual Try-on

    cs.CV 2024-11 conditional novelty 6.0 of 10

    FitDiT applies a customized Diffusion Transformer to image-based virtual try-on, adding a garment feature evolution stage, a frequency-domain loss, and a relaxed mask strategy to improve texture and size fidelity.

  24. Virtual Fitting Room: Generating Arbitrarily Long Videos of Virtual Try-On from a Single Image -- Technical Preview

    cs.CV 2025-09 conditional novelty 5.0 of 10

    VFR generates minute-long virtual try-on videos by auto-regressively chaining short diffusion-generated segments that are kept consistent with a 360-degree anchor video of the user.

  25. JCo-MVTON: Jointly Controllable Multi-Modal Diffusion Transformer for Mask-Free Virtual Try-on

    cs.CV 2025-08 conditional novelty 5.0 of 10

    A mask-free diffusion transformer for virtual try-on, trained with a self-generated and manually curated triplet dataset, achieves state-of-the-art scores on DressCode and competitive results on VITON-HD.

  26. Real-Time Per-Garment Virtual Try-On with Temporal Consistency for Loose-Fitting Garments

    cs.GR 2025-06 conditional novelty 5.0 of 10

    A per-garment virtual try-on method for loose-fitting garments uses a garment-invariant pose representation and a recurrent ConvLSTM synthesis network to achieve temporally smoother try-on video at about 10 fps.

  27. Insert Anything: Image Insertion via In-Context Editing in DiT

    cs.CV 2025-04 conditional novelty 5.0 of 10

    Insert Anything is a single model fine-tuned on 159,908 prompt-image pairs that performs mask- or text-guided insertion of people, objects, and garments from reference images into target scenes.

  28. CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation

    cs.CV 2025-01 conditional novelty 5.0 of 10

    CatV2TON unifies image and video virtual try-on in one diffusion transformer, using temporal garment-person concatenation and clip-based inference with AdaCN for long, consistent try-on videos.

  29. RealVVT: Towards Photorealistic Video Virtual Try-on via Spatio-Temporal Consistency

    cs.CV 2025-01 conditional novelty 5.0 of 10

    A diffusion-based framework improves spatial and temporal consistency of clothes in virtual try-on videos, reporting the best FID/KID and several video metrics on four public datasets.

  30. 1-2-1: Renaissance of Single-Network Paradigm for Virtual Try-On

    cs.CV 2025-01 conditional novelty 5.0 of 10

    A single-network virtual try-on model with modality-specific normalization and shared attention matches or beats dual-network reference-based models on image and video try-on benchmarks.

  31. Consistent Human Image and Video Generation with Spatially Conditioned Diffusion

    cs.CV 2024-12 conditional novelty 5.0 of 10

    Spatially conditioning a diffusion model by concatenating a reference human image with the noisy target, and adding causal self-attention, improves appearance consistency in human image and video animation.

  32. AnyDressing: Customizable Multi-Garment Virtual Dressing via Latent Diffusion Models

    cs.CV 2024-12 conditional novelty 5.0 of 10

    AnyDressing combines a parallel garment encoder with localized attention to generate a person wearing multiple specified garments from a text prompt.

  33. TalkFashion: Intelligent Virtual Try-On Assistant Based on Multimodal Large Language Model

    cs.CV 2025-07 conditional novelty 4.0 of 10

    TalkFashion, a text-driven virtual try-on assistant, reports better semantic consistency and visual quality than four baselines on VITON-HD by combining an LLM router, catalog matching, and automatic mask generation.

  34. DiffFit: Disentangled Garment Warping and Texture Refinement for Virtual Try-On

    cs.CV 2025-06 reject novelty 4.0 of 10

    DiffFit synthesizes virtual try-on images by separately warping the garment geometry and then refining texture with a conditional diffusion model, reporting SOTA metrics on VITON-HD and DressCode but with inconsistent...

  35. MFP-VTON: Enhancing Mask-Free Person-to-Person Virtual Try-On via Diffusion Transformer

    cs.CV 2025-02 reject novelty 4.0 of 10

    A mask-free person-to-person virtual try-on model built on FLUX-Fill-dev, trained with pseudo data generated by IDM and a Focus Attention loss.

Pith tools