Pith. sign in

REVIEW 1 cited by

Training-free Diffusion Model Adaptation for Variable-Sized Text-to-Image Synthesis

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.08645 v2 pith:YCHF54GB submitted 2023-06-14 cs.CV cs.LGeess.IV

classification cs.CVcs.LGeess.IV
keywords imagesmodelsattentiondiffusioninformationresolutionspatialsynthesis
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Diffusion models (DMs) have recently gained attention with state-of-the-art performance in text-to-image synthesis. Abiding by the tradition in deep learning, DMs are trained and evaluated on the images with fixed sizes. However, users are demanding for various images with specific sizes and various aspect ratio. This paper focuses on adapting text-to-image diffusion models to handle such variety while maintaining visual fidelity. First we observe that, during the synthesis, lower resolution images suffer from incomplete object portrayal, while higher resolution images exhibit repetitively disordered presentation. Next, we establish a statistical relationship indicating that attention entropy changes with token quantity, suggesting that models aggregate spatial information in proportion to image resolution. The subsequent interpretation on our observations is that objects are incompletely depicted due to limited spatial information for low resolutions, while repetitively disorganized presentation arises from redundant spatial information for high resolutions. From this perspective, we propose a scaling factor to alleviate the change of attention entropy and mitigate the defective pattern observed. Extensive experimental results validate the efficacy of the proposed scaling factor, enabling models to achieve better visual effects, image quality, and text alignment. Notably, these improvements are achieved without additional training or fine-tuning techniques.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Pippo: High-Resolution Multi-View Humans from a Single Image

    cs.CV 2025-02 conditional novelty 6.0 of 10

    A single-image multi-view diffusion transformer generates 1K-resolution turnaround views of humans, with attention biasing for many views and a new reprojection-error metric.

Pith tools