Pith. sign in

REVIEW 1 cited by

UniFashion: A Unified Vision-Language Model for Multimodal Fashion Retrieval and Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.11305 v2 pith:OQLZA73H submitted 2024-08-21 cs.CV cs.AI

classification cs.CVcs.AI
keywords generationmultimodaltasksfashionretrievaldomainmodelsunifashion
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The fashion domain encompasses a variety of real-world multimodal tasks, including multimodal retrieval and multimodal generation. The rapid advancements in artificial intelligence generated content, particularly in technologies like large language models for text generation and diffusion models for visual generation, have sparked widespread research interest in applying these multimodal models in the fashion domain. However, tasks involving embeddings, such as image-to-text or text-to-image retrieval, have been largely overlooked from this perspective due to the diverse nature of the multimodal fashion domain. And current research on multi-task single models lack focus on image generation. In this work, we present UniFashion, a unified framework that simultaneously tackles the challenges of multimodal generation and retrieval tasks within the fashion domain, integrating image generation with retrieval tasks and text generation tasks. UniFashion unifies embedding and generative tasks by integrating a diffusion model and LLM, enabling controllable and high-fidelity generation. Our model significantly outperforms previous single-task state-of-the-art models across diverse fashion tasks, and can be readily adapted to manage complex vision-language tasks. This work demonstrates the potential learning synergy between multimodal generation and retrieval, offering a promising direction for future research in the fashion domain. The source code is available at https://github.com/xiangyu-mm/UniFashion.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Agent-based Condition Monitoring Assistance with Multimodal Industrial Database Retrieval Augmented Generation

    cs.LG 2025-06 conditional novelty 5.0 of 10

    MindRAG retrieves similar historical vibration recordings and maintenance annotations, then uses LLM agents to generate fault predictions and alarm recommendations for industrial condition monitoring.

Pith tools