Pith. sign in

REVIEW 4 cited by

Non-confusing Generation of Customized Concepts in Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.06914 v1 pith:MKHTGQH4 submitted 2024-05-11 cs.CV

classification cs.CV
keywords conceptgenerationconceptscustomizedvisualfine-tuningclifclip
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We tackle the common challenge of inter-concept visual confusion in compositional concept generation using text-guided diffusion models (TGDMs). It becomes even more pronounced in the generation of customized concepts, due to the scarcity of user-provided concept visual examples. By revisiting the two major stages leading to the success of TGDMs -- 1) contrastive image-language pre-training (CLIP) for text encoder that encodes visual semantics, and 2) training TGDM that decodes the textual embeddings into pixels -- we point that existing customized generation methods only focus on fine-tuning the second stage while overlooking the first one. To this end, we propose a simple yet effective solution called CLIF: contrastive image-language fine-tuning. Specifically, given a few samples of customized concepts, we obtain non-confusing textual embeddings of a concept by fine-tuning CLIP via contrasting a concept and the over-segmented visual regions of other concepts. Experimental results demonstrate the effectiveness of CLIF in preventing the confusion of multi-customized concept generation.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TAP: Parameter-efficient Task-Aware Prompting for Adverse Weather Removal

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A two-stage prompt-tuning method with low-rank and contrastive prompt enhancement claims all-in-one adverse weather removal at 2.75M parameters.

  2. HOComp: Interaction-Aware Human-Object Composition

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A diffusion-transformer method that composes a foreground object into a human image with MLLM-chosen interaction regions, pose keypoint supervision, and appearance/background consistency losses, plus a new paired dataset.

  3. Show and Polish: Reference-Guided Identity Preservation in Face Video Restoration

    cs.CV 2025-07 conditional novelty 6.0 of 10

    IP-FVR restores degraded face videos with consistent identity by conditioning a video diffusion model on a reference photo of the same person.

  4. IRBridge: Solving Image Restoration Bridge with Pre-trained Generative Diffusion Models

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A Gaussian-path transition equation lets a pretrained Stable Diffusion model serve as the denoiser inside image restoration bridges, cutting per-task training to a lightweight ControlNet.

Pith tools