Pith. sign in

REVIEW 8 cited by

Wear-Any-Way: Manipulable Virtual Try-on via Sparse Correspondence Alignment

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.12965 v1 pith:4FKWGEHA submitted 2024-03-19 cs.CV

classification cs.CV
keywords wear-any-waytry-onmakestylevirtualalignmentcontrolcorrespondence
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper introduces a novel framework for virtual try-on, termed Wear-Any-Way. Different from previous methods, Wear-Any-Way is a customizable solution. Besides generating high-fidelity results, our method supports users to precisely manipulate the wearing style. To achieve this goal, we first construct a strong pipeline for standard virtual try-on, supporting single/multiple garment try-on and model-to-model settings in complicated scenarios. To make it manipulable, we propose sparse correspondence alignment which involves point-based control to guide the generation for specific locations. With this design, Wear-Any-Way gets state-of-the-art performance for the standard setting and provides a novel interaction form for customizing the wearing style. For instance, it supports users to drag the sleeve to make it rolled up, drag the coat to make it open, and utilize clicks to control the style of tuck, etc. Wear-Any-Way enables more liberated and flexible expressions of the attires, holding profound implications in the fashion industry.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. VTBench: Comprehensive Benchmark Suite Towards Real-World Virtual Try-on Models

    cs.CV 2025-05 conditional novelty 7.0 of 10

    VTBench is a multi-dimensional benchmark with novel unpaired metrics and human preference data for evaluating image-based virtual try-on models, though the human-alignment evidence is incomplete.

  2. PromptDresser: Improving the Quality and Controllability of Virtual Try-On via Generative Textual Prompt and Prompt-aware Mask

    cs.CV 2024-12 conditional novelty 6.0 of 10

    PromptDresser improves text-editable virtual try-on by combining LMM-generated structured captions with a prompt-aware adaptive mask.

  3. DiffusionTrend: A Minimalist Approach to Virtual Fashion Try-On

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A training-free virtual try-on pipeline that blends DDIM-inverted garment latents into masked model latents, guided by a lightweight CNN apparel mask.

  4. FashionComposer: Compositional Fashion Image Generation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A single diffusion framework composes multiple garment and face references into one fashion image using an asset library and subject-binding attention.

  5. FitDiT: Advancing the Authentic Garment Details for High-fidelity Virtual Try-on

    cs.CV 2024-11 conditional novelty 6.0 of 10

    FitDiT applies a customized Diffusion Transformer to image-based virtual try-on, adding a garment feature evolution stage, a frequency-domain loss, and a relaxed mask strategy to improve texture and size fidelity.

  6. CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation

    cs.CV 2025-01 conditional novelty 5.0 of 10

    CatV2TON unifies image and video virtual try-on in one diffusion transformer, using temporal garment-person concatenation and clip-based inference with AdaCN for long, consistent try-on videos.

  7. PEMF-VTO: Point-Enhanced Video Virtual Try-on via Mask-free Paradigm

    cs.CV 2024-12 conditional novelty 5.0 of 10

    A mask-free video virtual try-on model that uses sparse point correspondences between garment and frames, plus frame-to-frame tracking, to improve garment transfer and temporal coherence.

  8. Advancing Myopia To Holism: Fully Contrastive Language-Image Pre-training

    cs.CV 2024-11 conditional novelty 5.0 of 10

    Holistic CLIP trains a multi-branch image encoder with multi-to-multi contrastive learning on multiple VLM-generated captions per image and reports consistent gains over one-to-one and one-to-multi CLIP variants.

Pith tools