REVIEW 4 major objections 4 minor 38 references
Diffusion models and stochastic quantisation in lattice field theory
T0 review · 4 major / 4 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read Diffusion models generate lattice fields via stochastic quantisation
desk verdict A clear, honest proceedings summary of the authors' earlier diffusion-model work; the stochastic-quantisation analogy is sound, the tail-coverage caveat is real and explicitly acknowledged, and there is no new result here. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the score function $\nabla_\phi \log p_\tau(\phi)$, the gradient of the logarithm of the time-dependent probability density, which appears in the backward process and is approximated by a neural network. In stochastic quantisation the corresponding drift is $-\nabla S(\phi)$ for a known action, so the identification $p(\phi,\tau) = \exp[-S(\phi,\tau)]/Z$ converts the diffusion backward equation into a Langevin equation with a time-dependent drift and noise amplitude. The variance-exploding scheme, with no drift in the forward process and $g(t)=\sigma t/T$, makes the forward dynamics a pure noising process and the backward dynamics a denoising Langevin evolution over finite time intervals. This identification is what turns a generative model into a sampler: the neural network learns the drift needed to steer noise back onto the target distribution.
What would settle it
Train a variance-exploding diffusion model on a $32^2$ $\lambda\phi^4$ ensemble in the broken phase whose training data contains only configurations in one of the two wells, then generate a fresh ensemble and measure the fraction of configurations in the other well; if the generated ensemble does not reproduce the exact two-peak distribution, or equivalently if the learned score at large $|\phi|$ deviates measurably from $-\nabla S$, the backward process fails to sample the target distribution.
Extended reading notes
Core claim
On the paper's own terms, the central discovery is the formal identity between the backward process of a diffusion model and the stochastic quantisation Langevin equation. Writing the target distribution as $p(\phi,t) = (1/Z) \exp[-S(\phi,t)]$, the backward SDE becomes $\partial_\tau \phi(x,\tau) = -g^2(T-\tau) \nabla S(\phi,T-\tau) + g(T-\tau) \eta(x,\tau)$, which matches the stochastic quantisation equation $\partial_\tau \phi(x,\tau) = -\nabla S(\phi,\tau) + \sqrt{2}\,\eta(x,\tau)$ up to time-dependent noise normalisation. Because of this match, a diffusion model trained on existing lattice configurations, where the score $\nabla_\phi \log p_\tau(\phi)$ is learned by a neural network, can generate new configurations representative of the target theory. The paper demonstrates this concretely for $\lambda\phi^4$ theory on a $32^2$ lattice in both symmetric and broken phases, checking susceptibility, Binder cumulant, and higher-order cumulants, and reporting that generated configurations can serve as proposals with reduced auto-correlation.
Load-bearing premise
The entire scheme rests on the learned score $\nabla \log p_\tau(\phi)$ being accurate wherever the target distribution has support, but the model can only learn the score where training data exists, so rare field configurations contribute little to the training objective and may be generated with the wrong weight.
Editorial extensions
If this is right
- Diffusion models trained on existing ensembles generate new configurations that approximate the target distribution $p(\phi) \propto e^{-S(\phi)}$.
- Because each run starts from fresh noise, generated configurations are not serially correlated like HMC chains; they can be used as proposals in a Markov chain with reduced auto-correlation.
- The learned time-dependent drift and effective action can be read off and compared with the exact action, giving a diagnostic of how well the model has learned the theory.
- The same framework extends to U(1) gauge theories and, in progress, to theories with a complex action via complex Langevin, where the target distribution is not known a priori.
- Making the algorithm exact requires an accept-reject step; the paper reports encouraging acceptance rates in the detailed studies.
Reading between the lines
- The formal identity means that improvements to diffusion-model score estimation, such as better architectures or exact score matching, should transfer directly to better samplers for lattice theories; the stochastic-quantisation literature on kernels may offer a principled way to reduce finite-time bias in the backward process.
- The paper's caveat that the model can only learn where data is available suggests a concrete failure mode for theories with rare topological sectors or metastable phases: a training ensemble that under-samples a sector will produce a biased ensemble, and this could be tested by measuring the fraction of configurations in each sector against the exact value.
- Because the noising time $T$ is finite, the backward process is a biased sampler unless a Metropolis accept-reject step is included; this suggests that diffusion samplers in lattice field theory will in practice be used as proposal generators, not standalone samplers, unless the bias can be quantified and controlled.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This proceedings contribution argues that the backward process of a variance-exploding diffusion model, Eq. (4), is a time-dependent analogue of the stochastic quantisation equation, Eq. (5), and uses this connection to motivate generating configurations of a two-dimensional scalar field theory by denoising. The paper reviews the forward/backward formalism, shows denoising snapshots in the broken phase, presents a one-degree-of-freedom toy model in which the learned drift and effective action are compared with exact results, and outlines future directions including gauge theories, fermions, and theories with a complex action. Most quantitative lattice results are deferred to Refs. [20] and [23], and the paper repeatedly acknowledges its limitations, including finite-time runs and the fact that the diffusion model can only learn where data is available.
Significance. The proposed connection between diffusion models and stochastic quantisation is conceptually appealing and, if made quantitative, could provide a practical route to fast, decorrelated lattice configurations. The paper is honest about the main caveats and points to code and data in Ref. [20], which is a strength. As a standalone contribution, however, it does not establish that generated ensembles are representative of the target Boltzmann distribution: the key distance and acceptance evidence is in companion papers, and the self-reported deviations at large field values are load-bearing. If the open issues of tail coverage and exactness are addressed, the programme would be a significant contribution to machine-learning-based lattice sampling.
major comments (4)
- [Section 2, Eqs. (4)-(5)] The formal analogy between the backward diffusion equation and stochastic quantisation is correct at the level of the SDE, but Eq. (4) is not a stochastic-quantisation sampler in the usual sense. In Eq. (5) the drift is known and time-independent and the long-time limit guarantees convergence to the stationary distribution; in Eq. (4) the drift is time-dependent and learned, and the backward process runs for a finite interval T. Consequently, an error in the learned score at any intermediate time is inherited directly by the generated samples, with no asymptotic correction available. The bullet list in Section 2 notes the time dependence and finite-time runs, but the paper should state explicitly that Eq. (4) alone does not imply p_T(phi) approximately equal to p(phi); establishing that requires a separate, quantitative bound on the score error or an explicit validation of the final ensemble.
- [Section 3, Fig. 4] The paper's own toy-model diagnostic shows that the learned drift and effective action deviate strongly for large |phi|, and the accompanying text says 'the diffusion model can only learn where data is available'. These are precisely the tail regions that contribute to the Binder cumulant and to higher-order cumulants in the broken phase, so the deviation is not a peripheral caveat. The two-dimensional susceptibility and cumulant comparisons are only cited to Refs. [20,23] and are not shown here; this proceedings therefore contains no quantitative evidence that the tail error is controlled. Please add a quantitative tail diagnostic, for example a histogram ratio, a tail-restricted moment, or a comparison of the learned and exact actions on the full support, or alternatively state explicitly that representativeness on rare configurations is assumed rather than demonstrated.
- [Section 3, training and validation] The diffusion model is trained on HMC configurations and then checked against observables computed on the same HMC ensemble. Such a check largely tests the model's ability to reproduce its training distribution, not whether it correctly samples p(phi) proportional to exp[-S(phi)] in regions that are under-represented in the training data. This is not a formal circularity in the derivation, but it limits the evidential force of the reported agreement. An out-of-sample test, such as a different lattice volume, a different coupling, or an observable that was not used during training, would materially strengthen the central claim that the generated configurations are representative.
- [Section 4 and Fig. 2] The Outlook correctly states that making the algorithm exact by an accept-reject step is important and that work in this direction is in progress, and Fig. 2 hedges with 'if all algorithms are working well'. However, the abstract and parts of the introduction describe the method as a way to 'generate configurations' without the same qualification. Until an accept-reject correction or an equivalent exactness mechanism is included, direct samples from the trained model should be described as approximate proposals rather than as configurations drawn from the target distribution; the current wording overstates what is demonstrated.
minor comments (4)
- [Abstract and Section 1] The abbreviation 'LTFs' is used for 'lattice field theories'; the standard abbreviation is 'LFT' and should be made consistent.
- [Section 2] The name 'DALLE-E' appears twice; the correct product name is 'DALL-E'.
- [Fig. 4] The bottom row would be more informative if it included a quantitative comparison, such as sample means, variances, or a distance metric, rather than only overlaid histograms.
- [Eq. (6) and Fig. 4] The action in Eq. (6) is given with mu^2 = +/-1, but the figure caption refers to single-well and double-well cases; please state explicitly which mu^2 value corresponds to each column.
Circularity Check
No circularity: the central relation between the diffusion backward process and stochastic quantisation is derived in the text from the score identity, not imported as a fitted result; the empirical validation is cited to the authors' own prior work but is not load-bearing for the formal claim.
full rationale
The paper's main structural claim is the formal correspondence between Eq. (4), the backward diffusion equation obtained by substituting the score into the reverse SDE, and Eq. (5), the stochastic quantisation Langevin equation. This derivation is self-contained: Eq. (2) is the standard reverse process, Eq. (3) defines the action from the time-dependent density, and Eq. (4) follows by substitution. The paper explicitly lists the differences (time-dependent vs. constant drift, learned vs. known action), so it does not assert an equivalence by construction. The numerical demonstrations are delegated to the authors' own Refs. [20,23] ('The first two questions are answered positively in Ref. [20]'), which is a normal self-citation for a proceedings summary and is not used to forbid alternatives or to import a uniqueness theorem. The empirical consistency between generated and HMC ensembles is expected because the model is trained on HMC data; however, this is presented as a learned approximation of the target distribution, not as a prediction from first principles, so it does not constitute a fitted-input-called-prediction step. The paper's own caveat, 'the diffusion model can only learn where data is available, which explains the deviations seen for larger values of |φ|', identifies a limitation in the accuracy of the learned score on low-density tails. That is a correctness/robustness concern about the approximation, not a circularity in the derivation of the stochastic-quantisation correspondence. No specific reduction of a claimed output to its input by definition or self-citation chain is present in the text.
Assumptions & free parameters
free parameters (3)
- noise scale σ =
σ ≫ 1 (no exact value given)
- total diffusion time T
- neural network architecture and training hyperparameters
assumptions (3)
- standard math Reverse-time SDE for a diffusion process (Eq. (2) given score).
- domain assumption Existence of a time-dependent Boltzmann-like distribution p(φ,t) = Z^{-1} exp[-S(φ,t)] for all intermediate times.
- domain assumption The trained backward process samples from the target distribution p0(φ) once the score is learned.
Cite this review
Pith. "Pith review of Diffusion models and stochastic quantisation in lattice field theory." pith.science (2026). https://pith.science/paper/LVGVIC7U
@misc{pith2026241213704,
author = {Pith},
title = {Pith review of: Diffusion models and stochastic quantisation in lattice field theory},
year = {2026},
howpublished = {\url{https://pith.science/paper/LVGVIC7U}},
note = {Machine review of arXiv:2412.13704}
}
read the original abstract
Diffusion models are currently the leading generative AI approach used for image generation in e.g. DALL-E and Stable Diffusion. In this talk we relate diffusion models to stochastic quantisation in field theory and employ it to generate configurations for scalar fields on a two-dimensional lattice. We end with some speculations on possible applications.
Figures
Figures from the paper (1 more)
Reference graph
Works this paper leans on
-
[20]
L. Wang, G. Aarts and K. Zhou,Diffusion models as stochastic quantization in lattice field theory, JHEP05(2024) 060 [2309.17082]
arXiv 2024
-
[23]
G. Aarts, D.E. Habibi, L. Wang and K. Zhou,On learning higher-order cumulants in diffusion models, in38th conference on Neural Information Processing Systems, 2024 [2410.21212]
arXiv 2024
-
[1]
D. Boyda et al.,Applications of Machine Learning to Lattice Quantum Field Theory, in Snowmass 2021, 2022 [2202.05838]
arXiv 2021
-
[2]
K. Cranmer, G. Kanwar, S. Racanière, D.J. Rezende and P.E. Shanahan,Advances in machine-learning-based sampling motivated by lattice quantum chromodynamics, Nature Rev. Phys.5 (2023) 526 [2309.01156]
arXiv 2023
-
[3]
G. Kanwar,Flow-based sampling for lattice field theories, in40th International Symposium on Lattice Field Theory, 2024 [2401.01297]
arXiv 2024
-
[4]
Duane, A.D
S. Duane, A.D. Kennedy, B.J. Pendleton and D. Roweth,Hybrid Monte Carlo,Phys. Lett. B 195 (1987) 216
1987
-
[5]
M.S. Albergo, G. Kanwar and P.E. Shanahan,Flow-based generative models for Markov chain Monte Carlo in lattice field theory, Phys. Rev. D100 (2019) 034515 [1904.12072]
arXiv 2019
-
[6]
G. Kanwar, M.S. Albergo, D. Boyda, K. Cranmer, D.C. Hackett, S. Racanière et al., Equivariant Flow-Based Sampling for Lattice Gauge Theory, Phys. Rev. Lett.125 (2020) 121601 [2003.06413]
arXiv 2020
Show all 38 references
-
[7]
Rezende and S
D.J. Rezende and S. Mohamed,Variational Inference with Normalizing Flows, in International conference on machine learning, pp. 1530–1538, PMLR, 2015 [1505.05770]
2015 arXiv
-
[8]
F. Noé, S. Olsson, J. Köhler and H. Wu,Boltzmann generators: Sampling equilibrium states of many-body systems with deep learning, Science365 (2019) eaaw1147 [1812.01729]
2019 arXiv
-
[9]
Nicoli, S
K.A. Nicoli, S. Nakajima, N. Strodthoff, W. Samek, K.-R. Müller and P. Kessel, Asymptotically unbiased estimation of physical observables with neural samplers, Phys. Rev. E 101 (2020) 023304 [1910.13496]
2020 arXiv
-
[10]
Nicoli, C.J
K.A. Nicoli, C.J. Anders, L. Funcke, T. Hartung, K. Jansen, P. Kessel et al.,Estimation of thermodynamic observables in lattice field theories with deep generative models, Phys. Rev. Lett.126(2021) 032001 [2007.07115]
2021 arXiv
-
[11]
Del Debbio, J
L. Del Debbio, J. Marsh Rossney and M. Wilson,Efficient modeling of trivializing maps for lattice 𝜙4 theory using normalizing flows: A first look at scalability, Phys. Rev. D104 (2021) 094507 [2105.12481]. 7 Diffusion models and stochastic quantisation Gert Aarts
2021 arXiv
-
[12]
Nicoli, C.J
K.A. Nicoli, C.J. Anders, T. Hartung, K. Jansen, P. Kessel and S. Nakajima,Detecting and Mitigating Mode-Collapse for Flow-based Sampling of Lattice Field Theories, 2302.14082
-
[13]
R.T. Chen, Y. Rubanova, J. Bettencourt and D.K. Duvenaud,Neural ordinary differential equations,Advances in neural information processing systems31(2018) [1806.07366]
2018 arXiv
-
[14]
de Haan, C
P. de Haan, C. Rainone, M.C.N. Cheng and R. Bondesan,Scaling Up Machine Learning For Quantum Field Theory with Equivariant Continuous Flows, 2110.02673
-
[15]
Gerdes, P
M. Gerdes, P. de Haan, C. Rainone, R. Bondesan and M.C.N. Cheng,Learning lattice quantum field theories with equivariant continuous flows,SciPost Phys.15(2023) 238 [2207.00283]
2023 arXiv
-
[16]
Caselle, E
M. Caselle, E. Cellini and A. Nada,Sampling the lattice Nambu-Goto string using Continuous Normalizing Flows, Journal of High Energy Physics02 (2024) 048 [2307.01107]
2024 arXiv
-
[17]
H. Wu, J. Köhler and F. Noé,Stochastic normalizing flows, Advances in Neural Information Processing Systems33(2020) 5933 [2002.06707]
2020 arXiv
-
[18]
Caselle, E
M. Caselle, E. Cellini, A. Nada and M. Panero,Stochastic normalizing flows as non-equilibrium transformations,JHEP 07(2022) 015 [2201.08862]
2022 arXiv
-
[19]
K. Zhou, G. Endrődi, L.-G. Pang and H. Stöcker,Regressive and generative neural networks for scalar field theory, Phys. Rev. D100 (2019) 011501
2019
-
[21]
L. Wang, G. Aarts and K. Zhou,Generative Diffusion Models for Lattice Field Theory, in 37th Conference on Neural Information Processing Systems, 2023 [2311.03578]
2023 arXiv
-
[22]
Q. Zhu, G. Aarts, W. Wang, K. Zhou and L. Wang,Diffusion models for lattice gauge field simulations, in38th conference on Neural Information Processing Systems, 2024 [2410.19602]
2024 arXiv
-
[24]
Habibi, G
D.E. Habibi, G. Aarts, L. Wang and K. Zhou,Diffusion models learn distributions generated by complex Langevin dynamics, in41st International Symposium on Lattice Field Theory, 2024 [2412.01919]
2024 arXiv
-
[25]
Hirono, A
Y. Hirono, A. Tanaka and K. Fukushima,Understanding Diffusion Models by Feynman’s Path Integral, 2403.11262
-
[26]
Fukushima and S
K. Fukushima and S. Kamata,Stochastic quantization and diffusion models, 2411.11297. 8 Diffusion models and stochastic quantisation Gert Aarts
-
[27]
Ramesh, P
A. Ramesh, P. Dhariwal, A. Nichol, C. Chu and M. Chen,Hierarchical Text-Conditional Image Generation with CLIP Latents, 2204.06125
-
[28]
Rombach, A
R. Rombach, A. Blattmann, D. Lorenz, P. Esser and B. Ommer,High-Resolution Image Synthesis with Latent Diffusion Models,Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)(2022) 10684 [2112.10752]
2022 arXiv
-
[29]
Sohl-Dickstein, E.A
J. Sohl-Dickstein, E.A. Weiss, N. Maheswaranathan and S. Ganguli,Deep unsupervised learning using nonequilibrium thermodynamics, inProc. 32nd Int. Conf. Int. Conf. Mach. Learn. - Vol. 37, pp. 2256–2265, 2015 [1503.03585]
2015 arXiv
-
[30]
G.ParisiandY.S.Wu, Perturbationtheorywithoutgaugefixing ,Sci.China,A 24(1980)483
1980
-
[31]
Damgaard and H
P.H. Damgaard and H. Hüffel,Stochastic quantization,Phys. Rept.152 (1987) 227
1987
-
[32]
Alvestad, R
D. Alvestad, R. Larsen and A. Rothkopf,Towards learning optimized kernels for complex Langevin,JHEP 04(2023) 057 [2211.15625]
2023 arXiv
-
[33]
D.Boyda,G.Kanwar,S.Racanière,D.J.Rezende,M.S.Albergo,K.Cranmeretal., Sampling using 𝑆𝑈(𝑁) gauge equivariant flows,Phys. Rev. D103 (2021) 074504 [2008.05456]
2021 arXiv
-
[34]
Favoni, A
M. Favoni, A. Ipp, D.I. Müller and D. Schuh,Lattice Gauge Equivariant Convolutional Neural Networks, Phys. Rev. Lett.128 (2022) 032003 [2012.12901]
2022 arXiv
-
[35]
Parisi,On complex probabilities, Physics Letters B131(1983) 393
G. Parisi,On complex probabilities, Physics Letters B131(1983) 393
1983
-
[36]
Aarts and I.-O
G. Aarts and I.-O. Stamatescu,Stochastic quantization at finite chemical potential,JHEP 09 (2008) 018 [0807.1597]
2008 arXiv
-
[37]
Seiler, D
E. Seiler, D. Sexty and I.-O. Stamatescu,Gauge cooling in complex Langevin for QCD with heavy quarks, Phys. Lett. B723 (2013) 213 [1211.3709]
2013 arXiv
-
[38]
Aarts,Introductory lectures on lattice QCD at nonzero baryon number,J
G. Aarts,Introductory lectures on lattice QCD at nonzero baryon number,J. Phys. Conf. Ser.706 (2016) 022004 [1512.05145]. 9
2016 arXiv
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.