Pith. sign in

REVIEW 9 cited by

OOTDiffusion: Outfitting Fusion based Latent Diffusion for Controllable Virtual Try-on

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.01779 v2 pith:DR647JWT submitted 2024-03-04 cs.CV

classification cs.CV
keywords garmentootdiffusionoutfittingtry-onfeaturesvirtualcodecontrollability
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present OOTDiffusion, a novel network architecture for realistic and controllable image-based virtual try-on (VTON). We leverage the power of pretrained latent diffusion models, designing an outfitting UNet to learn the garment detail features. Without a redundant warping process, the garment features are precisely aligned with the target human body via the proposed outfitting fusion in the self-attention layers of the denoising UNet. In order to further enhance the controllability, we introduce outfitting dropout to the training process, which enables us to adjust the strength of the garment features through classifier-free guidance. Our comprehensive experiments on the VITON-HD and Dress Code datasets demonstrate that OOTDiffusion efficiently generates high-quality try-on results for arbitrary human and garment images, which outperforms other VTON methods in both realism and controllability, indicating an impressive breakthrough in virtual try-on. Our source code is available at https://github.com/levihsu/OOTDiffusion.

Discussion (0). Sign in to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. OmniVTON: Training-Free Universal Virtual Try-On

    cs.CV 2025-07 conditional novelty 7.0 of 10

    OmniVTON uses pretrained diffusion models with no training to transfer garments between people across shop and street scenes, and extends to multi-human try-on.

  2. WearWow: Native 2K Multi-Garment Virtual Try-On via Adaptive Token Packing and Preference Alignment

    cs.CV 2026-07 conditional novelty 6.0 of 10

    WearWow generates native 2K multi-garment virtual try-on images without masks, using token packing plus dual preference rewards to preserve fabric texture.

  3. Dress&Dance: Dress up and Dance as You Like It - Technical Preview

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A video diffusion framework that unifies text, image, and video conditioning through attention to produce high-resolution virtual try-on videos with reference-driven motion.

  4. FW-VTON: Flattening-and-Warping for Person-to-Person Virtual Try-on

    cs.CV 2025-07 conditional novelty 6.0 of 10

    FW-VTON reports state-of-the-art person-to-person virtual try-on results using a flattening, warping, and integration pipeline plus a new P2P-VTON dataset.

  5. Video Virtual Try-on with Conditional Diffusion Transformer Inpainter

    cs.CV 2025-06 conditional novelty 6.0 of 10

    ViTI reformulates video virtual try-on as conditional video inpainting with a full 3D attention diffusion transformer, and reports the best VFID score on VVT (2.121).

  6. Low-Barrier Dataset Collection with Real Human Body for Interactive Per-Garment Virtual Try-On

    cs.GR 2025-06 conditional novelty 6.0 of 10

    A per-garment virtual try-on pipeline that trains a GAN from a two-minute real-human video capture and uses a hybrid pose-plus-DensePose input to synthesize the garment with accurate alignment.

  7. Real-Time Per-Garment Virtual Try-On with Temporal Consistency for Loose-Fitting Garments

    cs.GR 2025-06 conditional novelty 5.0 of 10

    A per-garment virtual try-on method for loose-fitting garments uses a garment-invariant pose representation and a recurrent ConvLSTM synthesis network to achieve temporally smoother try-on video at about 10 fps.

  8. TalkFashion: Intelligent Virtual Try-On Assistant Based on Multimodal Large Language Model

    cs.CV 2025-07 conditional novelty 4.0 of 10

    TalkFashion, a text-driven virtual try-on assistant, reports better semantic consistency and visual quality than four baselines on VITON-HD by combining an LLM router, catalog matching, and automatic mask generation.

  9. DiffFit: Disentangled Garment Warping and Texture Refinement for Virtual Try-On

    cs.CV 2025-06 reject novelty 4.0 of 10

    DiffFit synthesizes virtual try-on images by separately warping the garment geometry and then refining texture with a conditional diffusion model, reporting SOTA metrics on VITON-HD and DressCode but with inconsistent...

Pith tools