Pith. sign in

REVIEW 3 cited by

FouriScale: A Frequency Perspective on Training-Free High-Resolution Image Synthesis

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.12963 v1 pith:G47K24AC submitted 2024-03-19 cs.CV

classification cs.CV
keywords fouriscalegenerationhigh-resolutionimagesmethodmodelsstructuralconsistency
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this study, we delve into the generation of high-resolution images from pre-trained diffusion models, addressing persistent challenges, such as repetitive patterns and structural distortions, that emerge when models are applied beyond their trained resolutions. To address this issue, we introduce an innovative, training-free approach FouriScale from the perspective of frequency domain analysis. We replace the original convolutional layers in pre-trained diffusion models by incorporating a dilation technique along with a low-pass operation, intending to achieve structural consistency and scale consistency across resolutions, respectively. Further enhanced by a padding-then-crop strategy, our method can flexibly handle text-to-image generation of various aspect ratios. By using the FouriScale as guidance, our method successfully balances the structural integrity and fidelity of generated images, achieving an astonishing capacity of arbitrary-size, high-resolution, and high-quality generation. With its simplicity and compatibility, our method can provide valuable insights for future explorations into the synthesis of ultra-high-resolution images. The code will be released at https://github.com/LeonHLJ/FouriScale.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CineScale: Free Lunch in High-Resolution Cinematic Visual Generation

    cs.CV 2025-08 conditional novelty 6.0 of 10

    CineScale extends pre-trained diffusion models to 8k image and 4k video generation with mostly tuning-free inference plus a small LoRA adaptation for video.

  2. Revisiting Diffusion Models: From Generative Pre-training to One-Step Generation

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Fine-tuning a pretrained diffusion model with a GAN objective and most weights frozen yields a one-step generator that matches or beats prior distillation methods on several datasets.

  3. APT: Improving Diffusion Models for High Resolution Image Generation with Adaptive Path Tracing

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A training-free add-on for latent diffusion models that fixes patch statistics and re-schedules noise, improving detail in high-resolution images while reducing sampling steps.

Pith tools