REVIEW 11 cited by
A Fourier Space Perspective on Diffusion Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Diffusion models are state-of-the-art generative models on data modalities such as images, audio, proteins and materials. These modalities share the property of exponentially decaying variance and magnitude in the Fourier domain. Under the standard Denoising Diffusion Probabilistic Models (DDPM) forward process of additive white noise, this property results in high-frequency components being corrupted faster and earlier in terms of their Signal-to-Noise Ratio (SNR) than low-frequency ones. The reverse process then generates low-frequency information before high-frequency details. In this work, we study the inductive bias of the forward process of diffusion models in Fourier space. We theoretically analyse and empirically demonstrate that the faster noising of high-frequency components in DDPM results in violations of the normality assumption in the reverse process. Our experiments show that this leads to degraded generation quality of high-frequency components. We then study an alternate forward process in Fourier space which corrupts all frequencies at the same rate, removing the typical frequency hierarchy during generation, and demonstrate marked performance improvements on datasets where high frequencies are primary, while performing on par with DDPM on standard imaging benchmarks.
Forward citations
Cited by 11 Pith papers
-
A First-Principles Theory of Slow Thinking and Active Perception
Active lifting of data distributions via latent-sequence sampling and max-rate uncertainty reduction formally derives slow-thinking LLMs and places them on representation and sampler hierarchies that can be climbed.
-
Enhancing Membership Inference Attacks on Diffusion Models from a Frequency-Domain Perspective
Removing high-frequency components from reconstruction-error scores improves membership inference attacks on diffusion models, demonstrated on DDIM and Stable Diffusion.
-
Commitment Before Realization: When Classifier-Free Guidance Becomes Unnecessary in Masked Diffusion Language Models
A prompt-specific commitment horizon, identified by comparing guided versus base-only continuations, marks an early point where classifier-free guidance can be removed with little loss in constraint success.
-
CSGen: A Multi-Domain Curvilinear Structure Generation Model via Hierarchical Multimodal Diffusion
CSGen generates curvilinear structure images from layout and text conditions, improving structural fidelity and downstream segmentation over ControlNet and LoRA baselines.
-
WaiT for the Signal: Simple Frequency-Aware Flow-Matching
WaiT delays high-frequency wavelet bands in flow-matching image generation until coarse structure emerges, improving quality and cutting compute, with a reported SOTA FID of 1.30 on ImageNet 512.
-
Manifold-Constrained Noise Optimization for Diverse Diffusion Sampling
By optimizing each new starting noise on a fixed-radius, low-frequency sphere, MoNO recovers per-prompt diversity in distilled text-to-image models while keeping image quality roughly stable.
-
End-to-End Training for Autoregressive Video Diffusion via Self-Resampling
Resampling Forcing trains autoregressive video diffusion models on self-resampled degraded histories with a causal mask, achieving stable long-horizon generation without a teacher or discriminator.
-
Adaptive Transition State Refinement with Learned Equilibrium Flows
AEFM is a learned, structure-only refinement method that iteratively improves low-fidelity transition state geometries toward DFT-quality structures.
-
LSSGen: Leveraging Latent Space Scaling in Flow and Diffusion for Efficient Text to Image Generation
A latent-space scaling framework that replaces pixel-space upscaling with a trainable latent upsampler and noise compensation, yielding faster high-resolution text-to-image generation.
-
Cloud Diffusion Part 1: Theory and Motivation
Replacing white noise with scale-invariant noise tuned to an image set's power-law statistics could make diffusion models faster, sharper, and more controllable, this theory paper argues.
-
Perceptually Aligning Representations of Music via Noise-Augmented Autoencoders
Noise-augmented training of an audio autoencoder organizes representations so that perceptually salient information survives in coarse structures, improving musical surprisal estimates and EEG prediction.
Discussion (0). Continue with ORCID to comment.