REVIEW 29 cited by
Improving Diffusion Models for Authentic Virtual Try-on in the Wild
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
This paper considers image-based virtual try-on, which renders an image of a person wearing a curated garment, given a pair of images depicting the person and the garment, respectively. Previous works adapt existing exemplar-based inpainting diffusion models for virtual try-on to improve the naturalness of the generated visuals compared to other methods (e.g., GAN-based), but they fail to preserve the identity of the garments. To overcome this limitation, we propose a novel diffusion model that improves garment fidelity and generates authentic virtual try-on images. Our method, coined IDM-VTON, uses two different modules to encode the semantics of garment image; given the base UNet of the diffusion model, 1) the high-level semantics extracted from a visual encoder are fused to the cross-attention layer, and then 2) the low-level features extracted from parallel UNet are fused to the self-attention layer. In addition, we provide detailed textual prompts for both garment and person images to enhance the authenticity of the generated visuals. Finally, we present a customization method using a pair of person-garment images, which significantly improves fidelity and authenticity. Our experimental results show that our method outperforms previous approaches (both diffusion-based and GAN-based) in preserving garment details and generating authentic virtual try-on images, both qualitatively and quantitatively. Furthermore, the proposed customization method demonstrates its effectiveness in a real-world scenario. More visualizations are available in our project page: https://idm-vton.github.io
Forward citations
Cited by 29 Pith papers
-
AssetDropper: Asset Extraction via Diffusion Models with Reward-Driven Optimization
AssetDropper introduces a task-specific diffusion model, a 212k-pair synthetic dataset, and a generative reward model to extract standardized assets from reference images.
-
WearWow: Native 2K Multi-Garment Virtual Try-On via Adaptive Token Packing and Preference Alignment
WearWow generates native 2K multi-garment virtual try-on images without masks, using token packing plus dual preference rewards to preserve fabric texture.
-
FastFit: Accelerating Multi-Reference Virtual Try-On via Cacheable Diffusion Models
FastFit uses a cacheable diffusion UNet to compute multi-reference garment features once per generation, enabling about 3.5x faster multi-item virtual try-on with comparable or better fidelity.
-
FontAdapter: Instant Font Adaptation in Visual Text Generation
A two-stage curriculum with synthetic paired font data enables instant adaptation of unseen fonts in text-to-image generation using one reference glyph, without test-time fine-tuning.
-
Creating Your Editable 3D Photorealistic Avatar with Tetrahedron-constrained Gaussian Splatting
TetGS is a hybrid representation that embeds Gaussian kernels inside tetrahedral grids, enabling locally controlled geometric and appearance edits of 3D avatars reconstructed from monocular video.
-
3DV-TON: Textured 3D-Guided Consistent Video Try-on via Diffusion Models
A diffusion video try-on framework that uses animated textured 3D meshes as frame-level guidance, plus a new high-resolution benchmark, achieves stronger temporal consistency and garment fidelity than two released baselines.
-
Any2AnyTryon: Leveraging Adaptive Position Embeddings for Versatile Virtual Clothing Tasks
A single DiT-based model with adaptive position embeddings performs virtual try-on, garment reconstruction, model-free try-on, and layered try-on from text and variable-size image inputs.
-
IPVTON: Image-based 3D Virtual Try-on with Image Prompt Adapter
IPVTON produces a 3D human model wearing a target garment from one person image and one garment image by combining score distillation with mask-guided image prompts and a pseudo silhouette loss.
-
DreamFit: Garment-Centric Human Generation via a Lightweight Anything-Dressing Encoder
DreamFit generates human images from a garment reference and text by encoding the reference through LoRA-activated layers of a frozen Stable Diffusion UNet and injecting features with adaptive attention.
-
PromptDresser: Improving the Quality and Controllability of Virtual Try-On via Generative Textual Prompt and Prompt-aware Mask
PromptDresser improves text-editable virtual try-on by combining LMM-generated structured captions with a prompt-aware adaptive mask.
-
DiffusionTrend: A Minimalist Approach to Virtual Fashion Try-On
A training-free virtual try-on pipeline that blends DDIM-inverted garment latents into masked model latents, guided by a lightweight CNN apparel mask.
-
FashionComposer: Compositional Fashion Image Generation
A single diffusion framework composes multiple garment and face references into one fashion image using an asset library and subject-binding attention.
-
IGR: Improving Diffusion Model for Garment Restoration from Person Image
IGR restores a clean garment image from a person photo using Stable Diffusion, dual extractors, attention fusion blocks, and a VITON-to-GarmRe fine-tuning strategy, beating TryOffDiff on the reported benchmarks.
-
Learning Implicit Features with Flow Infused Attention for Realistic Virtual Try-On
A virtual try-on model replaces explicit garment warping with flow-infused cross-attention guidance in a Stable Diffusion UNet, reporting SOTA scores on VITON-HD and DressCode.
-
SwiftTry: Fast and Consistent Video Virtual Try-On with Diffusion Models
SwiftTry makes diffusion-based video virtual try-on faster and more consistent by shifting non-overlapping video chunks during sampling and caching features across denoising steps.
-
InstantRestore: Single-Step Personalized Face Restoration with Shared-Image Attention
A single-step diffusion-based face restoration model uses reference-image attention to preserve identity in about 0.5 seconds per image, with no per-identity tuning.
-
Refine-by-Align: Reference-Guided Artifacts Refinement through Semantic Alignment
Refine-by-Align uses diffusion cross-attention maps to locate the reference region matching a masked artifact, then re-inpaints the artifact with that reference detail.
-
TKG-DM: Training-free Chroma Key Content Generation Diffusion Model
Adjusting the mean of specific channels in the initial noise of Stable Diffusion produces foreground objects on a uniform, user-selected chroma key background without any fine-tuning.
-
JCo-MVTON: Jointly Controllable Multi-Modal Diffusion Transformer for Mask-Free Virtual Try-on
A mask-free diffusion transformer for virtual try-on, trained with a self-generated and manually curated triplet dataset, achieves state-of-the-art scores on DressCode and competitive results on VITON-HD.
-
CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation
CatV2TON unifies image and video virtual try-on in one diffusion transformer, using temporal garment-person concatenation and clip-based inference with AdaCN for long, consistent try-on videos.
-
1-2-1: Renaissance of Single-Network Paradigm for Virtual Try-On
A single-network virtual try-on model with modality-specific normalization and shared attention matches or beats dual-network reference-based models on image and video try-on benchmarks.
-
DRDM: A Disentangled Representations Diffusion Model for Synthesizing Realistic Person Images
The paper introduces a disentangled diffusion model with parsing-map-guided sampling that achieves modest improvements in pose transfer and appearance control on DeepFashion.
-
Consistent Human Image and Video Generation with Spatially Conditioned Diffusion
Spatially conditioning a diffusion model by concatenating a reference human image with the noisy target, and adding causal self-attention, improves appearance consistency in human image and video animation.
-
AnyDressing: Customizable Multi-Garment Virtual Dressing via Latent Diffusion Models
AnyDressing combines a parallel garment encoder with localized attention to generate a person wearing multiple specified garments from a text prompt.
-
PEMF-VTO: Point-Enhanced Video Virtual Try-on via Mask-free Paradigm
A mask-free video virtual try-on model that uses sparse point correspondences between garment and frames, plus frame-to-frame tracking, to improve garment transfer and temporal coherence.
-
TED-VITON: Transformer-Empowered Diffusion Models for Virtual Try-On
TED-VITON adapts a transformer-based diffusion model (SD3) for virtual try-on with a garment adapter, a text-preservation loss, and LLM-generated prompts, achieving top scores on VITON-HD and DressCode.
-
Borrowing from anything: A generalizable framework for reference-guided instance editing
GENIE uses spatial alignment, residual feature scaling, and progressive attention fusion to transfer a reference's appearance onto a target, achieving state-of-the-art scores on AnyInsertion.
-
CONVERGE: A Multi-Agent Vision-Radio Architecture for xApps
CONVERGE fuses camera and radio sensing inside O-RAN xApps via a multi-agent architecture, reporting under-one-millisecond sensing delay for real-time blockage-driven RAN control.
-
MFP-VTON: Enhancing Mask-Free Person-to-Person Virtual Try-On via Diffusion Transformer
A mask-free person-to-person virtual try-on model built on FLUX-Fill-dev, trained with pseudo data generated by IDM and a Focus Attention loss.
Discussion (0). Continue with ORCID to comment.