Pith. sign in

REVIEW 29 cited by

Improving Diffusion Models for Authentic Virtual Try-on in the Wild

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.05139 v3 pith:JN2KOVTW submitted 2024-03-08 cs.CV

classification cs.CV
keywords garmentimagestry-onvirtualdiffusionmethodauthenticperson
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

This paper considers image-based virtual try-on, which renders an image of a person wearing a curated garment, given a pair of images depicting the person and the garment, respectively. Previous works adapt existing exemplar-based inpainting diffusion models for virtual try-on to improve the naturalness of the generated visuals compared to other methods (e.g., GAN-based), but they fail to preserve the identity of the garments. To overcome this limitation, we propose a novel diffusion model that improves garment fidelity and generates authentic virtual try-on images. Our method, coined IDM-VTON, uses two different modules to encode the semantics of garment image; given the base UNet of the diffusion model, 1) the high-level semantics extracted from a visual encoder are fused to the cross-attention layer, and then 2) the low-level features extracted from parallel UNet are fused to the self-attention layer. In addition, we provide detailed textual prompts for both garment and person images to enhance the authenticity of the generated visuals. Finally, we present a customization method using a pair of person-garment images, which significantly improves fidelity and authenticity. Our experimental results show that our method outperforms previous approaches (both diffusion-based and GAN-based) in preserving garment details and generating authentic virtual try-on images, both qualitatively and quantitatively. Furthermore, the proposed customization method demonstrates its effectiveness in a real-world scenario. More visualizations are available in our project page: https://idm-vton.github.io

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 29 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AssetDropper: Asset Extraction via Diffusion Models with Reward-Driven Optimization

    cs.CV 2025-06 conditional novelty 7.0 of 10

    AssetDropper introduces a task-specific diffusion model, a 212k-pair synthetic dataset, and a generative reward model to extract standardized assets from reference images.

  2. WearWow: Native 2K Multi-Garment Virtual Try-On via Adaptive Token Packing and Preference Alignment

    cs.CV 2026-07 conditional novelty 6.0 of 10

    WearWow generates native 2K multi-garment virtual try-on images without masks, using token packing plus dual preference rewards to preserve fabric texture.

  3. FastFit: Accelerating Multi-Reference Virtual Try-On via Cacheable Diffusion Models

    cs.CV 2025-08 conditional novelty 6.0 of 10

    FastFit uses a cacheable diffusion UNet to compute multi-reference garment features once per generation, enabling about 3.5x faster multi-item virtual try-on with comparable or better fidelity.

  4. FontAdapter: Instant Font Adaptation in Visual Text Generation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A two-stage curriculum with synthetic paired font data enables instant adaptation of unseen fonts in text-to-image generation using one reference glyph, without test-time fine-tuning.

  5. Creating Your Editable 3D Photorealistic Avatar with Tetrahedron-constrained Gaussian Splatting

    cs.GR 2025-04 conditional novelty 6.0 of 10

    TetGS is a hybrid representation that embeds Gaussian kernels inside tetrahedral grids, enabling locally controlled geometric and appearance edits of 3D avatars reconstructed from monocular video.

  6. 3DV-TON: Textured 3D-Guided Consistent Video Try-on via Diffusion Models

    cs.CV 2025-04 conditional novelty 6.0 of 10

    A diffusion video try-on framework that uses animated textured 3D meshes as frame-level guidance, plus a new high-resolution benchmark, achieves stronger temporal consistency and garment fidelity than two released baselines.

  7. Any2AnyTryon: Leveraging Adaptive Position Embeddings for Versatile Virtual Clothing Tasks

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A single DiT-based model with adaptive position embeddings performs virtual try-on, garment reconstruction, model-free try-on, and layered try-on from text and variable-size image inputs.

  8. IPVTON: Image-based 3D Virtual Try-on with Image Prompt Adapter

    cs.CV 2025-01 conditional novelty 6.0 of 10

    IPVTON produces a 3D human model wearing a target garment from one person image and one garment image by combining score distillation with mask-guided image prompts and a pseudo silhouette loss.

  9. DreamFit: Garment-Centric Human Generation via a Lightweight Anything-Dressing Encoder

    cs.CV 2024-12 conditional novelty 6.0 of 10

    DreamFit generates human images from a garment reference and text by encoding the reference through LoRA-activated layers of a frozen Stable Diffusion UNet and injecting features with adaptive attention.

  10. PromptDresser: Improving the Quality and Controllability of Virtual Try-On via Generative Textual Prompt and Prompt-aware Mask

    cs.CV 2024-12 conditional novelty 6.0 of 10

    PromptDresser improves text-editable virtual try-on by combining LMM-generated structured captions with a prompt-aware adaptive mask.

  11. DiffusionTrend: A Minimalist Approach to Virtual Fashion Try-On

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A training-free virtual try-on pipeline that blends DDIM-inverted garment latents into masked model latents, guided by a lightweight CNN apparel mask.

  12. FashionComposer: Compositional Fashion Image Generation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A single diffusion framework composes multiple garment and face references into one fashion image using an asset library and subject-binding attention.

  13. IGR: Improving Diffusion Model for Garment Restoration from Person Image

    cs.CV 2024-12 conditional novelty 6.0 of 10

    IGR restores a clean garment image from a person photo using Stable Diffusion, dual extractors, attention fusion blocks, and a VITON-to-GarmRe fine-tuning strategy, beating TryOffDiff on the reported benchmarks.

  14. Learning Implicit Features with Flow Infused Attention for Realistic Virtual Try-On

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A virtual try-on model replaces explicit garment warping with flow-infused cross-attention guidance in a Stable Diffusion UNet, reporting SOTA scores on VITON-HD and DressCode.

  15. SwiftTry: Fast and Consistent Video Virtual Try-On with Diffusion Models

    cs.CV 2024-12 conditional novelty 6.0 of 10

    SwiftTry makes diffusion-based video virtual try-on faster and more consistent by shifting non-overlapping video chunks during sampling and caching features across denoising steps.

  16. InstantRestore: Single-Step Personalized Face Restoration with Shared-Image Attention

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A single-step diffusion-based face restoration model uses reference-image attention to preserve identity in about 0.5 seconds per image, with no per-identity tuning.

  17. Refine-by-Align: Reference-Guided Artifacts Refinement through Semantic Alignment

    cs.CV 2024-11 conditional novelty 6.0 of 10

    Refine-by-Align uses diffusion cross-attention maps to locate the reference region matching a masked artifact, then re-inpaints the artifact with that reference detail.

  18. TKG-DM: Training-free Chroma Key Content Generation Diffusion Model

    cs.CV 2024-11 conditional novelty 6.0 of 10

    Adjusting the mean of specific channels in the initial noise of Stable Diffusion produces foreground objects on a uniform, user-selected chroma key background without any fine-tuning.

  19. JCo-MVTON: Jointly Controllable Multi-Modal Diffusion Transformer for Mask-Free Virtual Try-on

    cs.CV 2025-08 conditional novelty 5.0 of 10

    A mask-free diffusion transformer for virtual try-on, trained with a self-generated and manually curated triplet dataset, achieves state-of-the-art scores on DressCode and competitive results on VITON-HD.

  20. CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation

    cs.CV 2025-01 conditional novelty 5.0 of 10

    CatV2TON unifies image and video virtual try-on in one diffusion transformer, using temporal garment-person concatenation and clip-based inference with AdaCN for long, consistent try-on videos.

  21. 1-2-1: Renaissance of Single-Network Paradigm for Virtual Try-On

    cs.CV 2025-01 conditional novelty 5.0 of 10

    A single-network virtual try-on model with modality-specific normalization and shared attention matches or beats dual-network reference-based models on image and video try-on benchmarks.

  22. DRDM: A Disentangled Representations Diffusion Model for Synthesizing Realistic Person Images

    cs.CV 2024-12 conditional novelty 5.0 of 10

    The paper introduces a disentangled diffusion model with parsing-map-guided sampling that achieves modest improvements in pose transfer and appearance control on DeepFashion.

  23. Consistent Human Image and Video Generation with Spatially Conditioned Diffusion

    cs.CV 2024-12 conditional novelty 5.0 of 10

    Spatially conditioning a diffusion model by concatenating a reference human image with the noisy target, and adding causal self-attention, improves appearance consistency in human image and video animation.

  24. AnyDressing: Customizable Multi-Garment Virtual Dressing via Latent Diffusion Models

    cs.CV 2024-12 conditional novelty 5.0 of 10

    AnyDressing combines a parallel garment encoder with localized attention to generate a person wearing multiple specified garments from a text prompt.

  25. PEMF-VTO: Point-Enhanced Video Virtual Try-on via Mask-free Paradigm

    cs.CV 2024-12 conditional novelty 5.0 of 10

    A mask-free video virtual try-on model that uses sparse point correspondences between garment and frames, plus frame-to-frame tracking, to improve garment transfer and temporal coherence.

  26. TED-VITON: Transformer-Empowered Diffusion Models for Virtual Try-On

    cs.CV 2024-11 conditional novelty 5.0 of 10

    TED-VITON adapts a transformer-based diffusion model (SD3) for virtual try-on with a garment adapter, a text-preservation loss, and LLM-generated prompts, achieving top scores on VITON-HD and DressCode.

  27. Borrowing from anything: A generalizable framework for reference-guided instance editing

    cs.CV 2025-12 conditional novelty 4.0 of 10

    GENIE uses spatial alignment, residual feature scaling, and progressive attention fusion to transfer a reference's appearance onto a target, achieving state-of-the-art scores on AnyInsertion.

  28. CONVERGE: A Multi-Agent Vision-Radio Architecture for xApps

    cs.NI 2025-08 reject novelty 4.0 of 10

    CONVERGE fuses camera and radio sensing inside O-RAN xApps via a multi-agent architecture, reporting under-one-millisecond sensing delay for real-time blockage-driven RAN control.

  29. MFP-VTON: Enhancing Mask-Free Person-to-Person Virtual Try-On via Diffusion Transformer

    cs.CV 2025-02 reject novelty 4.0 of 10

    A mask-free person-to-person virtual try-on model built on FLUX-Fill-dev, trained with pseudo data generated by IDM and a Focus Attention loss.

Pith tools