REVIEW 2 cited by
DEADiff: An Efficient Stylization Diffusion Model with Disentangled Representations
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The diffusion-based text-to-image model harbors immense potential in transferring reference style. However, current encoder-based approaches significantly impair the text controllability of text-to-image models while transferring styles. In this paper, we introduce DEADiff to address this issue using the following two strategies: 1) a mechanism to decouple the style and semantics of reference images. The decoupled feature representations are first extracted by Q-Formers which are instructed by different text descriptions. Then they are injected into mutually exclusive subsets of cross-attention layers for better disentanglement. 2) A non-reconstructive learning method. The Q-Formers are trained using paired images rather than the identical target, in which the reference image and the ground-truth image are with the same style or semantics. We show that DEADiff attains the best visual stylization results and optimal balance between the text controllability inherent in the text-to-image model and style similarity to the reference image, as demonstrated both quantitatively and qualitatively. Our project page is https://tianhao-qi.github.io/DEADiff/.
Forward citations
Cited by 2 Pith papers
-
StyleSSP: Sampling StartPoint Enhancement for Training-free Diffusion-based Method for Style Transfer
StyleSSP improves training-free diffusion style transfer by tuning the sampling startpoint through frequency filtering and negative guidance during inversion.
-
Enhancing Low-Cost Video Editing with Lightweight Adaptors and Temporal-Aware Inversion
GE-Adapter combines a temporal smoothness loss, bilateral-filtered DDIM inversion, and shared plus frame-specific prompt tokens to improve text-to-video editing, though the reported evidence is inconsistent.
Discussion (0). Continue with ORCID to comment.