Pith. sign in

REVIEW 4 cited by

How Compositional Generalization and Creativity Improve as Diffusion Models are Trained

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.12089 v3 pith:DA56ANVE submitted 2025-02-17 stat.ML cs.LG

classification stat.MLcs.LG
keywords datamodelsdiffusionclusteringcompositioncontextfeatureshierarchical
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Natural data is often organized as a hierarchical composition of features. How many samples do generative models need in order to learn the composition rules, so as to produce a combinatorially large number of novel data? What signal in the data is exploited to learn those rules? We investigate these questions in the context of diffusion models both theoretically and empirically. Theoretically, we consider a simple probabilistic context-free grammar - a tree-like graphical model used to represent the hierarchical and compositional structure of data such as language and images. We demonstrate that diffusion models learn the grammar's composition rules with the sample complexity required for clustering features with statistically similar context, a process similar to the word2vec algorithm. However, this clustering emerges hierarchically: higher-level features associated with longer contexts require more data to be identified. This mechanism leads to a sample complexity that scales polynomially with the said context size. As a result, diffusion models trained on an intermediate dataset size generate data coherent up to a certain scale, but lacking global coherence. We test these predictions across different domains and find remarkable agreement: both generated texts and images achieve progressively larger coherence lengths as the training time or dataset size grows. We discuss connections between the hierarchical clustering mechanism we introduce here and the renormalization group in physics.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A theory of learning data statistics in diffusion models, from easy to hard

    stat.ML 2026-03 unverdicted novelty 6.0 of 10

    Diffusion models exhibit a distributional simplicity bias, learning pairwise input statistics at linear sample complexity while fourth-order cumulants require cubic complexity unless sharing correlated latent structure.

  2. Bigger Isn't Always Memorizing: Early Stopping Overparameterized Diffusion Models

    cs.LG 2025-05 conditional novelty 6.0 of 10

    In overparameterized diffusion models, generalization happens first and memorization starts later, with the memorization time growing linearly with dataset size.

  3. Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures

    cs.LG 2025-05 conditional novelty 6.0 of 10

    Convolutional networks trained on a random hierarchical grammar improve twice as fast with data as transformers, because weight sharing reuses the statistical signal across all positions.

  4. A solvable generative model with a linear, one-step denoiser

    cs.LG 2024-11 reject novelty 5.0 of 10

    The paper derives a closed-form KL divergence for a one-step linear diffusion model on Gaussian data, reports a sample-size threshold at n=d, and gives a heuristic argument that more diffusion steps improve quality.

Pith tools