Pith. sign in

REVIEW 9 cited by

StableGarment: Garment-Centric Generation via Stable Diffusion

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.10783 v1 pith:MCE7GTQ3 submitted 2024-03-16 cs.CV

classification cs.CV
keywords try-ongarmentgarment-centricgenerationstablegarmentstylizedtext-to-imagevirtual
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper, we introduce StableGarment, a unified framework to tackle garment-centric(GC) generation tasks, including GC text-to-image, controllable GC text-to-image, stylized GC text-to-image, and robust virtual try-on. The main challenge lies in retaining the intricate textures of the garment while maintaining the flexibility of pre-trained Stable Diffusion. Our solution involves the development of a garment encoder, a trainable copy of the denoising UNet equipped with additive self-attention (ASA) layers. These ASA layers are specifically devised to transfer detailed garment textures, also facilitating the integration of stylized base models for the creation of stylized images. Furthermore, the incorporation of a dedicated try-on ControlNet enables StableGarment to execute virtual try-on tasks with precision. We also build a novel data engine that produces high-quality synthesized data to preserve the model's ability to follow prompts. Extensive experiments demonstrate that our approach delivers state-of-the-art (SOTA) results among existing virtual try-on methods and exhibits high flexibility with broad potential applications in various garment-centric image generation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Dress&Dance: Dress up and Dance as You Like It - Technical Preview

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A video diffusion framework that unifies text, image, and video conditioning through attention to produce high-resolution virtual try-on videos with reference-driven motion.

  2. Stable-Hair v2: Real-World Hair Transfer via Multiple-View Diffusion Model

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A multi-view diffusion hair-transfer system transfers a reference hairstyle onto a portrait and renders the edited person from many consistent viewpoints.

  3. ID-Cloak: Crafting Identity-Specific Cloaks Against Personalized Text-to-Image Generation

    cs.CV 2025-02 conditional novelty 6.0 of 10

    ID-Cloak creates a single universal image perturbation from a few photos that degrades personalized text-to-image generation of that identity across unseen images.

  4. Any2AnyTryon: Leveraging Adaptive Position Embeddings for Versatile Virtual Clothing Tasks

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A single DiT-based model with adaptive position embeddings performs virtual try-on, garment reconstruction, model-free try-on, and layered try-on from text and variable-size image inputs.

  5. DreamFit: Garment-Centric Human Generation via a Lightweight Anything-Dressing Encoder

    cs.CV 2024-12 conditional novelty 6.0 of 10

    DreamFit generates human images from a garment reference and text by encoding the reference through LoRA-activated layers of a frozen Stable Diffusion UNet and injecting features with adaptive attention.

  6. FashionComposer: Compositional Fashion Image Generation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A single diffusion framework composes multiple garment and face references into one fashion image using an asset library and subject-binding attention.

  7. CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation

    cs.CV 2025-01 conditional novelty 5.0 of 10

    CatV2TON unifies image and video virtual try-on in one diffusion transformer, using temporal garment-person concatenation and clip-based inference with AdaCN for long, consistent try-on videos.

  8. AnyDressing: Customizable Multi-Garment Virtual Dressing via Latent Diffusion Models

    cs.CV 2024-12 conditional novelty 5.0 of 10

    AnyDressing combines a parallel garment encoder with localized attention to generate a person wearing multiple specified garments from a text prompt.

  9. CONVERGE: A Multi-Agent Vision-Radio Architecture for xApps

    cs.NI 2025-08 reject novelty 4.0 of 10

    CONVERGE fuses camera and radio sensing inside O-RAN xApps via a multi-agent architecture, reporting under-one-millisecond sensing delay for real-time blockage-driven RAN control.

Pith tools