REVIEW 16 cited by
IMAGGarment: Fine-Grained Garment Generation for Controllable Fashion Design
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
This paper presents IMAGGarment, a fine-grained garment generation (FGG) framework that enables high-fidelity garment synthesis with precise control over silhouette, color, and logo placement. Unlike existing methods that are limited to single-condition inputs, IMAGGarment addresses the challenges of multi-conditional controllability in personalized fashion design and digital apparel applications. Specifically, IMAGGarment employs a two-stage training strategy to separately model global appearance and local details, while enabling unified and controllable generation through end-to-end inference. In the first stage, we propose a global appearance model that jointly encodes silhouette and color using a mixed attention module and a color adapter. In the second stage, we present a local enhancement model with an adaptive appearance-aware module to inject user-defined logos and spatial constraints, enabling accurate placement and visual consistency. To support this task, we release GarmentBench, a large-scale dataset comprising over 180K garment samples paired with multi-level design conditions, including sketches, color references, logo placements, and textual prompts. Extensive experiments demonstrate that our method outperforms existing baselines, achieving superior structural stability, color fidelity, and local controllability performance. Code, models, and datasets are publicly available at https://github.com/muzishen/IMAGGarment.
Forward citations
Cited by 16 Pith papers
-
LaP-Forensics: Latent-Pixel Consistency Guided Multimodal Reasoning for Deepfake Detection
A dual-stream deepfake forensic model that adds DDIM reconstruction residuals to RGB features improves artifact localization and cross-generator detection in evaluations, with honest caveats about text faithfulness.
-
StyleTailor: Towards Personalized Fashion Styling via Hierarchical Negative Feedback
A two-agent fashion styling framework that iteratively corrects garment retrieval and virtual try-on using hierarchical vision-language negative feedback.
-
FGAA-FPN: Foreground-Guided Angle-Aware Feature Pyramid Network for Oriented Object Detection
FGAA-FPN combines box-supervised foreground masks with orientation-biased attention in a feature pyramid and reports +3.9 mAP over FPN on DOTA v1.5, while its claimed SOTA on DOTA v1.0 is contradicted by its own table.
-
ACM-UNet: Adaptive Integration of CNNs and Mamba for Efficient Medical Image Segmentation
ACM-UNet, a UNet variant combining pretrained CNN and Mamba backbones via lightweight adapters and a wavelet decoder module, reports 85.12% Dice on Synapse and 92.29% on ACDC.
-
Edge-Aware and Content-Adaptive Infrared Gas Leak Detection for Industrial Safety Monitoring
PEG-DRNet reports 29.8% AP / 84.3% AP50 on the IIG infrared gas dataset and 36.3% AP / 68.5% AP50 on LangGas, beating RT-DETR-R18 by 3.0/6.5 AP/AP50 points.
-
PQ-DAF: Pose-driven Quality-controlled Data Augmentation for Data-scarce Driver Distraction Detection
PQ-DAF uses pose-conditioned diffusion generation plus CogVLM filtering to augment few-shot driver distraction training data, and reports large accuracy gains that are compromised by a non-standard train/test protocol.
-
Hybrid Compact Least-Squares and Central Weighted Essentially Non-Oscillatory Schemes for Hyperbolic Conservation Laws on Structured Curvilinear Grids
No verifiable result: the abstract and body address unrelated topics, so the claimed CLS-CWENO schemes appear without derivation, experiments, or benchmarks.
-
FashionPose: Unified Text-Driven Fashion Synthesis with Joint Geometric and Photometric Control
A single caption can drive pose generation, person-image synthesis, and relighting through a three-stage FashionPose pipeline, with reported text-to-pose gains on DF-PASS that are undermined by inconsistent tables.
-
DiffFit: Disentangled Garment Warping and Texture Refinement for Virtual Try-On
DiffFit synthesizes virtual try-on images by separately warping the garment geometry and then refining texture with a conditional diffusion model, reporting SOTA metrics on VITON-HD and DressCode but with inconsistent...
-
O2Former:Direction-Aware and Multi-Scale Query Enhancement for SAR Ship Instance Segmentation
O2Former adds a multi-scale query generator and an orientation-aware module to Mask2Former and reports improved SAR ship instance segmentation on SSDD and HRSID.
-
R-Genie: Reasoning-Guided Generative Image Editing
R-Genie couples a multimodal LLM with a discrete diffusion model to perform image edits that require commonsense reasoning, and introduces a 1,070-triple benchmark called REditBench.
-
Dual Attention Residual U-Net for Accurate Brain Ultrasound Segmentation in IVH Detection
A residual U-Net with CBAM and a dual-branch sparse/dense attention layer reports Dice 89.04 and IoU 81.84 on brain ultrasound ventricle segmentation.
-
YOLO-FDA: Integrating Hierarchical Attention and Detail Enhancement for Surface Defect Detection
A YOLOv5 variant with BiFPN, directional detail enhancement, and two attention fusion modules reports state-of-the-art mAP on GC10-DET and DAGM2007.
-
MCFNet: A Multimodal Collaborative Fusion Network for Fine-Grained Semantic Classification
MCFNet fuses ALBERT text features and ViT image features with dropout, L1/L2 regularization, hybrid self/cross attention, and multi-loss training, claiming state-of-the-art accuracy on Con-Text and Drink Bottle.
-
YOLO-SPCI: Enhancing Remote Sensing Object Detection via Selective-Perspective-Class Integration
YOLO-SPCI, a YOLOv8 variant with a three-branch attention module at backbone stages P3 and P5, reports 92.0% mAP50 on NWPU VHR-10 versus 88.9% for the baseline.
-
FreqU-FNet: Frequency-Aware U-Net for Imbalanced Medical Image Segmentation
FreqU-FNet mixes frequency-domain filters and adaptive upsampling inside a U-Net, and claims better minority-class segmentation, although key reported results are internally inconsistent.
Discussion (0). Sign in to comment.