REVIEW 6 cited by
Direct Consistency Optimization for Robust Customization of Text-to-Image Diffusion Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Text-to-image (T2I) diffusion models, when fine-tuned on a few personal images, can generate visuals with a high degree of consistency. However, such fine-tuned models are not robust; they often fail to compose with concepts of pretrained model or other fine-tuned models. To address this, we propose a novel fine-tuning objective, dubbed Direct Consistency Optimization, which controls the deviation between fine-tuning and pretrained models to retain the pretrained knowledge during fine-tuning. Through extensive experiments on subject and style customization, we demonstrate that our method positions itself on a superior Pareto frontier between subject (or style) consistency and image-text alignment over all previous baselines; it not only outperforms regular fine-tuning objective in image-text alignment, but also shows higher fidelity to the reference images than the method that fine-tunes with additional prior dataset. More importantly, the models fine-tuned with our method can be merged without interference, allowing us to generate custom subjects in a custom style by composing separately customized subject and style models. Notably, we show that our approach achieves better prompt fidelity and subject fidelity than those post-optimized for merging regular fine-tuned models.
Forward citations
Cited by 6 Pith papers
-
Noise Consistency Regularization for Improved Subject-Driven Image Synthesis
Adding consistency-to-pretrained and multiplicative-noise consistency losses to fine-tuning improves subject identity and background diversity over DreamBooth on a 30-subject benchmark.
-
Transformed Low-rank Adaptation via Tensor Decomposition and Its Applications to Text-to-image Models
TLoRA combines a tensor-ring-matrix transform with a tensor-ring residual to fine-tune text-to-image models, achieving better or comparable performance than LoRA with far fewer parameters.
-
DECOR:Decomposition and Projection of Text Embeddings for Text-to-Image Customization
DECOR suppresses undesired word-token semantics in text embeddings via orthogonal projection, reducing prompt misalignment and content leakage in LoRA-customized text-to-image models.
-
Safety Alignment Backfires: Preventing the Re-emergence of Suppressed Concepts in Fine-tuned Text-to-Image Diffusion Models
Fine-tuning text-to-image diffusion models on benign data can reactivate suppressed unsafe concepts, and training the task adapter separately from a frozen safety LoRA prevents this.
-
Style-Friendly SNR Sampler for Style-Driven Generation
Fine-tuning diffusion models for style-driven generation is improved by shifting the training noise-level distribution toward high-noise steps, where stylistic features are shown to emerge.
-
Improving Multi-Subject Consistency in Open-Domain Image Generation with Isolation and Reposition Attention
IR-Diffusion adds two attention masks, Isolation and Reposition, that stop subjects in an image from blending into each other and align reference features to target positions, improving multi-subject consistency witho...
Discussion (0). Continue with ORCID to comment.