The authors claim that LLM prompt refinement plus a CLIP-based weak supervision filter improves diffusion-based fashion image generation, but the evidence is unverifiable and internally inconsistent.
FIRST: A Million-Entry Dataset for Text-Driven Fashion Synthesis and Design
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Text-driven fashion synthesis and design is an extremely valuable part of artificial intelligence generative content(AIGC), which has the potential to propel a tremendous revolution in the traditional fashion industry. To advance the research on text-driven fashion synthesis and design, we introduce a new dataset comprising a million high-resolution fashion images with rich structured textual(FIRST) descriptions. In the FIRST, there is a wide range of attire categories and each image-paired textual description is organized at multiple hierarchical levels. Experiments on prevalent generative models trained over FISRT show the necessity of FIRST. We invite the community to further develop more intelligent fashion synthesis and design systems that make fashion design more creative and imaginative based on our dataset. The dataset will be released soon.
citation-role summary
citation-polarity summary
fields
cs.CL 1years
2025 1verdicts
REJECT 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Cross-Cultural Fashion Design via Interactive Large Language Models and Diffusion Models
The authors claim that LLM prompt refinement plus a CLIP-based weak supervision filter improves diffusion-based fashion image generation, but the evidence is unverifiable and internally inconsistent.