REVIEW 18 cited by
IMAGDressing-v1: Customizable Virtual Dressing
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Latest advances have achieved realistic virtual try-on (VTON) through localized garment inpainting using latent diffusion models, significantly enhancing consumers' online shopping experience. However, existing VTON technologies neglect the need for merchants to showcase garments comprehensively, including flexible control over garments, optional faces, poses, and scenes. To address this issue, we define a virtual dressing (VD) task focused on generating freely editable human images with fixed garments and optional conditions. Meanwhile, we design a comprehensive affinity metric index (CAMI) to evaluate the consistency between generated images and reference garments. Then, we propose IMAGDressing-v1, which incorporates a garment UNet that captures semantic features from CLIP and texture features from VAE. We present a hybrid attention module, including a frozen self-attention and a trainable cross-attention, to integrate garment features from the garment UNet into a frozen denoising UNet, ensuring users can control different scenes through text. IMAGDressing-v1 can be combined with other extension plugins, such as ControlNet and IP-Adapter, to enhance the diversity and controllability of generated images. Furthermore, to address the lack of data, we release the interactive garment pairing (IGPair) dataset, containing over 300,000 pairs of clothing and dressed images, and establish a standard pipeline for data assembly. Extensive experiments demonstrate that our IMAGDressing-v1 achieves state-of-the-art human image synthesis performance under various controlled conditions. The code and model will be available at https://github.com/muzishen/IMAGDressing.
Forward citations
Cited by 18 Pith papers
-
Trajectory Map-Matching in Urban Road Networks Based on RSS Measurements
An HMM-based method that fits a signal-propagation model and decodes vehicle positions on a road graph from raw 5G RSS measurements achieves about 12-15 m trajectory error on two city datasets.
-
DreamFit: Garment-Centric Human Generation via a Lightweight Anything-Dressing Encoder
DreamFit generates human images from a garment reference and text by encoding the reference through LoRA-activated layers of a frozen Stable Diffusion UNet and injecting features with adaptive attention.
-
PromptDresser: Improving the Quality and Controllability of Virtual Try-On via Generative Textual Prompt and Prompt-aware Mask
PromptDresser improves text-editable virtual try-on by combining LMM-generated structured captions with a prompt-aware adaptive mask.
-
FashionComposer: Compositional Fashion Image Generation
A single diffusion framework composes multiple garment and face references into one fashion image using an asset library and subject-binding attention.
-
Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation
Simignore improves multimodal LLM complex question answering on ScienceQA by masking image tokens whose embeddings have low cosine similarity to the text prompt.
-
Controllable Human Image Generation with Personalized Multi-Garments
BootComp bootstraps large synthetic multi-garment training data with a decomposition network, then trains a frozen-generator diffusion model that generates humans wearing multiple reference garments with higher report...
-
Rethinking Vision Transformer for Large-Scale Fine-Grained Image Retrieval
EET speeds up vision transformers for fine-grained image retrieval by pruning background tokens and using teacher-student distillation to preserve accuracy, cutting latency by 42.7% with little or no drop in retrieval...
-
CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation
CatV2TON unifies image and video virtual try-on in one diffusion transformer, using temporal garment-person concatenation and clip-based inference with AdaCN for long, consistent try-on videos.
-
Re-Attentional Controllable Video Diffusion Editing
ReAtCo improves text-guided video editing by using attention-map gradients to place edited objects in user-specified regions and by re-injecting the original background during diffusion sampling.
-
AnyDressing: Customizable Multi-Garment Virtual Dressing via Latent Diffusion Models
AnyDressing combines a parallel garment encoder with localized attention to generate a person wearing multiple specified garments from a text prompt.
-
Hybrid Compact Least-Squares and Central Weighted Essentially Non-Oscillatory Schemes for Hyperbolic Conservation Laws on Structured Curvilinear Grids
No verifiable result: the abstract and body address unrelated topics, so the claimed CLS-CWENO schemes appear without derivation, experiments, or benchmarks.
-
Enhancing, Refining, and Fusing: Towards Robust Multi-Scale and Dense Ship Detection
CASS-Det, a YOLOX-based SAR ship detector with a center-enhancement module, a neighbor attention module, and a cross-connected feature pyramid network, achieves state-of-the-art mAP on SSDD, HRSID, and LS-SSDD.
-
Dialogue Director: Bridging the Gap in Dialogue Visualization for Multimodal Storytelling
Dialogue Director converts dialogue scripts into multi-view storyboards using GPT-4-based script analysis, multi-view diffusion, and cinematic layout planning, with mixed quantitative gains over baselines.
-
Fab-ME: A Vision State-Space and Attention-Enhanced Framework for Fabric Defect Detection
Fab-ME modifies YOLOv8s with a VMamba-based state-space module in the neck and an enhanced channel attention module, reporting 59.4 percent mAP@0.5 on the Tianchi fabric defect dataset versus a 57.4 percent baseline.
-
CCi-YOLOv8n: Enhanced Fire Detection with CARAFE and Context-Guided Modules
A YOLOv8n variant combining CARAFE, Context-Guided Downsampling, and iRMB reports small accuracy gains over YOLOv8n on two fire-detection datasets.
-
ICAS: IP Adapter and ControlNet-based Attention Structure for Multi-Subject Style Transfer Optimization
ICAS is a style-transfer pipeline that freezes IP-Adapter's style path, lightly tunes its content path, and adds ControlNet structure conditioning to preserve multi-subject layouts.
-
Cross-modal Context Fusion and Adaptive Graph Convolutional Network for Multimodal Conversational Emotion Recognition
A model combining co-attention transformers, a BiGRU, and graph convolution for speaker relationships reports higher accuracies on IEMOCAP and MELD, though the evaluation has significant gaps.
-
First-place Solution for Streetscape Shop Sign Recognition Competition
A team reports winning a street-view shop sign recognition competition with a multi-stage OCR pipeline built from known components, but provides no code, data, or rigorous ablations.
Discussion (0). Continue with ORCID to comment.