Pith. sign in

REVIEW 1 cited by

Simple Drop-in LoRA Conditioning on Attention Layers Will Improve Your Diffusion Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.03958 v3 pith:5RONDWIK submitted 2024-05-07 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords layersconditioningattentiondiffusiongenerationlorau-netconvolutional
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Current state-of-the-art diffusion models employ U-Net architectures containing convolutional and (qkv) self-attention layers. The U-Net processes images while being conditioned on the time embedding input for each sampling step and the class or caption embedding input corresponding to the desired conditional generation. Such conditioning involves scale-and-shift operations to the convolutional layers but does not directly affect the attention layers. While these standard architectural choices are certainly effective, not conditioning the attention layers feels arbitrary and potentially suboptimal. In this work, we show that simply adding LoRA conditioning to the attention layers without changing or tuning the other parts of the U-Net architecture improves the image generation quality. For example, a drop-in addition of LoRA conditioning to EDM diffusion model yields FID scores of 1.91/1.75 for unconditional and class-conditional CIFAR-10 generation, improving upon the baseline of 1.97/1.79.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CytoDiff: AI-Driven Cytomorphology Image Synthesis for Medical Diagnostics

    cs.CV 2025-07 conditional novelty 4.0 of 10

    Adding 5,000 CytoDiff-generated synthetic white blood cell images per class is reported to improve ResNet-50 accuracy from 27% to 78% and CLIP accuracy from 62% to 77% on the Munich AML dataset.

Pith tools