Pith. sign in

REVIEW 6 cited by

Direct Consistency Optimization for Robust Customization of Text-to-Image Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.12004 v2 pith:ETOOI2J2 submitted 2024-02-19 cs.CV

classification cs.CV
keywords modelsfine-tunedconsistencyfine-tuningstylesubjectfidelitymethod
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Text-to-image (T2I) diffusion models, when fine-tuned on a few personal images, can generate visuals with a high degree of consistency. However, such fine-tuned models are not robust; they often fail to compose with concepts of pretrained model or other fine-tuned models. To address this, we propose a novel fine-tuning objective, dubbed Direct Consistency Optimization, which controls the deviation between fine-tuning and pretrained models to retain the pretrained knowledge during fine-tuning. Through extensive experiments on subject and style customization, we demonstrate that our method positions itself on a superior Pareto frontier between subject (or style) consistency and image-text alignment over all previous baselines; it not only outperforms regular fine-tuning objective in image-text alignment, but also shows higher fidelity to the reference images than the method that fine-tunes with additional prior dataset. More importantly, the models fine-tuned with our method can be merged without interference, allowing us to generate custom subjects in a custom style by composing separately customized subject and style models. Notably, we show that our approach achieves better prompt fidelity and subject fidelity than those post-optimized for merging regular fine-tuned models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Noise Consistency Regularization for Improved Subject-Driven Image Synthesis

    cs.GR 2025-06 conditional novelty 6.0 of 10

    Adding consistency-to-pretrained and multiplicative-noise consistency losses to fine-tuning improves subject identity and background diversity over DreamBooth on a 30-subject benchmark.

  2. Transformed Low-rank Adaptation via Tensor Decomposition and Its Applications to Text-to-image Models

    cs.LG 2025-01 conditional novelty 6.0 of 10

    TLoRA combines a tensor-ring-matrix transform with a tensor-ring residual to fine-tune text-to-image models, achieving better or comparable performance than LoRA with far fewer parameters.

  3. DECOR:Decomposition and Projection of Text Embeddings for Text-to-Image Customization

    cs.CV 2024-12 conditional novelty 6.0 of 10

    DECOR suppresses undesired word-token semantics in text embeddings via orthogonal projection, reducing prompt misalignment and content leakage in LoRA-customized text-to-image models.

  4. Safety Alignment Backfires: Preventing the Re-emergence of Suppressed Concepts in Fine-tuned Text-to-Image Diffusion Models

    cs.AI 2024-11 conditional novelty 6.0 of 10

    Fine-tuning text-to-image diffusion models on benign data can reactivate suppressed unsafe concepts, and training the task adapter separately from a frozen safety LoRA prevents this.

  5. Style-Friendly SNR Sampler for Style-Driven Generation

    cs.CV 2024-11 conditional novelty 6.0 of 10

    Fine-tuning diffusion models for style-driven generation is improved by shifting the training noise-level distribution toward high-noise steps, where stylistic features are shown to emerge.

  6. Improving Multi-Subject Consistency in Open-Domain Image Generation with Isolation and Reposition Attention

    cs.CV 2024-11 conditional novelty 5.0 of 10

    IR-Diffusion adds two attention masks, Isolation and Reposition, that stop subjects in an image from blending into each other and align reference features to target positions, improving multi-subject consistency witho...

Pith tools