REVIEW 14 cited by
Breaking the Curse of Dimensionality: Diffusion Models Efficiently Learn Low-Dimensional Distributions
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Despite their empirical success across a wide range of generative tasks, the fundamental principles underlying the ability of diffusion models to learn data distributions are poorly understood. In this work, we develop a new mathematical framework that explains how diffusion models can effectively learn low-dimensional distributions from a finite number of training samples without suffering from the curse of dimensionality. Specifically, motivated by the intrinsic low-dimensional structure of image data, we theoretically analyze a setting in which the data distribution is modeled as a mixture of low-rank Gaussians. Under suitable network parameterization, we show that optimizing the training objective of diffusion models is equivalent to solving the canonical subspace clustering problem over the training samples, where each subspace basis corresponds to the low-rank covariance of a Gaussian component. This equivalence allows us to show that the sample complexity for learning the underlying distribution scales linearly with the intrinsic dimension of the data, rather than exponentially with the ambient dimension. Our theoretical findings are further supported by empirical evidence that demonstrates phase transition phenomena in generalization on both synthetic and real-world image datasets. Moreover, we establish a correspondence between the learned subspace bases and semantic attributes of image data, providing a principled foundation for controllable image generation.
Forward citations
Cited by 14 Pith papers
-
Low-dimensional adaptation of diffusion models: Convergence in total variation
Under exact score functions and a covering-number notion of intrinsic dimension, DDIM and DDPM reach TV error epsilon in O-tilde(k/epsilon) iterations.
-
An analytic theory of creativity in convolutional diffusion models
Convolutional diffusion models generate novel images by assembling locally consistent patch mosaics of training patches, and this mechanism is captured by an analytic score machine that predicts individual model outputs.
-
ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features
ConceptAttention shows that linear projections in the output space of DiT attention layers yield sharper concept-localizing saliency maps than cross-attention maps, reaching state-of-the-art zero-shot segmentation.
-
Diffusion Models Adapt to Low-Dimensional Structure Under Flexible Coefficient Choices
For a broad class of coefficients, diffusion models achieve Õ(k/ε) iteration complexity for ε-accurate TV sampling under low-dimensional structure, independent of ambient dimension.
-
Deep Image Prototype Learning with Geometric Heat-Kernel Priors
Introduces a geometry-aware EM algorithm using heat-kernel-weighted graphs and medoids to enforce on-manifold prototypes in variational models for medical imaging cohorts.
-
The Effect of Training Task Diversity on In-Context Learning through the Lens of Low-Dimensional Subspaces
A low-rank Gaussian mixture model shows that training task diversity measured by non-overlapping subspace columns improves ICL generalization and shortens learning plateaus for linear attention, with empirical extensi...
-
On the Collapse Errors Induced by the Deterministic Sampler for Diffusion Models
Deterministic diffusion samplers induce collapse errors, where samples overly concentrate locally, caused by low-noise score learning degrading high-noise score accuracy.
-
Attention-Only Transformers via Unrolled Subspace Denoising
An attention-only transformer, derived as unrolled subspace denoising, provably multiplies token signal-to-noise ratio by a fixed factor per layer and roughly matches GPT-2 and ViT on small benchmarks.
-
The Linear Geometry of Interpretable Tokens: Jailbreaking Attacks and Defenses for Unlearned Diffusion Models
Attack token embeddings learned as non-negative sums of CLIP vocabulary tokens can jailbreak unlearned diffusion models, and projecting those embeddings out of the vocabulary reduces attack success while preserving im...
-
Subspace Langevin Monte Carlo
SLMC generalizes random-coordinate and preconditioned Langevin Monte Carlo by projecting updates onto random eigenblocks of a preconditioner, with coupling-based error bounds.
-
CCS: Controllable and Constrained Sampling with Diffusion Models via Initial Noise Perturbation
A training-free diffusion sampling method exploits an observed linear relation between initial noise perturbations and output changes to control the sample mean and diversity around a target image.
-
Adaptivity and Convergence of Probability Flow ODEs in Diffusion Generative Models
With accurate score estimates, the probability flow ODE sampler reaches O(k/T) total-variation error, where k is the intrinsic dimension of the target distribution.
-
Decentralized Diffusion Models
An ensemble of expert diffusion models trained in isolation on disjoint data clusters, combined by a learned router, matches the global flow-matching objective and outperforms a monolithic model at equal FLOPs.
-
Principles and Practice of Deep Representation Learning: or a Mathematical Theory of Memory
The book presents principles from optimization and information theory to explain deep network architectures and enable new interpretable models.
Discussion (0). Continue with ORCID to comment.