{"id":"b27ada6a-d448-4f2e-9ef1-1b9decd4203c","arxiv_id":"2412.13704","paper_version":1,"verdict":"UNVERDICTED","confidence":"MODERATE","novelty_score":2.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Diffusion models, whose backward denoising step resembles stochastic quantisation, can learn from HMC data to generate configurations for 2D scalar lattice field theory.","lead":"This conference talk explains how diffusion models, the AI technique behind image generators, are mathematically similar to stochastic quantisation in field theory, and shows their use in generating configurations for a scalar field on a two-dimensional lattice. It is a summary of the authors' own earlier work, with speculations about future applications to gauge theories and complex actions.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Generated ensembles are representative only if the learned score is accurate on low-density tails; the paper's own Fig. 4 shows that is where it deviates, so the central claim is conditional on unquantified tail coverage.","rationale":"The reader's weakest assumption identifies exactly the concern that matters here: the validity of generated configurations depends on the learned score being accurate on the full support of the target distribution, including the tails. The paper's Section 3 and Fig. 4 explicitly concede that deviations are largest in the low-density regions, so this is not a manufactured objection but a limitation the authors themselves flag. I checked the formal analogy between Eq. (4) and Eq. (5): with the variance-exploding reverse-time SDE and the identification ∇ log p = -∇S, the sign and coefficients are consistent, so there is no internal mathematical error in the core correspondence. The weakness is in the interpretational step from formal equivalence to actual representative sampling. Because the paper is a conference proceedings summary that defers quantitative validation, acceptance rates, and detailed comparisons to Refs. [20,23], the concern does not overturn the reader's UNVERDICTED verdict; rather, it sharpens why the verdict should remain UNVERDICTED rather than ACCEPT. A training-size scaling test is the most direct way to determine whether finite-data tail coverage is the dominant source of bias in the generated ensembles.","tokens_in":6093,"tokens_out":4394,"duration_ms":44571,"concrete_test":"Retrain the variance-exploding diffusion model for the two-dimensional φ^4 theory at the same volume L=32 with training ensembles of size 5120, 10240, and 20480, keeping the U-net architecture and all hyperparameters fixed, then compute the Binder cumulant and the fourth cumulant of the generated ensembles and compare them with the HMC reference values. If these observables drift monotonically with training-set size toward the HMC values, the residual bias is due to finite-data tail coverage and the representativeness claim is not yet validated; if the observables are stable and agree within errors, the tail-coverage concern is mitigated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that the backward diffusion process, Eq. (4), has the same form as stochastic quantisation, Eq. (5), and therefore a trained diffusion model generates configurations representative of p(φ) ∝ e^{-S(φ)}—is conditionally true, but the condition is load-bearing. Eq. (4) is exact only when the learned score equals ∇ log p_t(φ) for every t and every φ in the support of the noised distribution. Unlike Eq. (5), which converges to equilibrium for any initial condition with a known drift, the finite-time reverse process has no asymptotic correction: any score bias is inherited by the generated ensemble. The paper itself states in Section 3, with Fig. 4, that 'the diffusion model can only learn where data is available,' and the plotted deviations of the learned drift and effective action are largest at large |φ|. These tail regions are precisely where the Binder cumulant and higher-order cumulants in the broken phase receive important contributions. Since the results shown are direct generations rather than proposals corrected by an accept-reject step (acceptance analysis is only cited to Refs. [20,23]), the representativeness assertion is not established by the evidence in this manuscript. This is not an internal inconsistency, but it means the central claim rests on an unquantified approximation regarding rare field values.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This proceedings contribution argues that the backward process of a variance-exploding diffusion model, Eq. (4), is a time-dependent analogue of the stochastic quantisation equation, Eq. (5), and uses this connection to motivate generating configurations of a two-dimensional scalar field theory by denoising. The paper reviews the forward/backward formalism, shows denoising snapshots in the broken phase, presents a one-degree-of-freedom toy model in which the learned drift and effective action are compared with exact results, and outlines future directions including gauge theories, fermions, and theories with a complex action. Most quantitative lattice results are deferred to Refs. [20] and [23], and the paper repeatedly acknowledges its limitations, including finite-time runs and the fact that the diffusion model can only learn where data is available.","tokens_in":6325,"tokens_out":4929,"duration_ms":45755,"significance":"The proposed connection between diffusion models and stochastic quantisation is conceptually appealing and, if made quantitative, could provide a practical route to fast, decorrelated lattice configurations. The paper is honest about the main caveats and points to code and data in Ref. [20], which is a strength. As a standalone contribution, however, it does not establish that generated ensembles are representative of the target Boltzmann distribution: the key distance and acceptance evidence is in companion papers, and the self-reported deviations at large field values are load-bearing. If the open issues of tail coverage and exactness are addressed, the programme would be a significant contribution to machine-learning-based lattice sampling.","major_comments":[{"comment":"The formal analogy between the backward diffusion equation and stochastic quantisation is correct at the level of the SDE, but Eq. (4) is not a stochastic-quantisation sampler in the usual sense. In Eq. (5) the drift is known and time-independent and the long-time limit guarantees convergence to the stationary distribution; in Eq. (4) the drift is time-dependent and learned, and the backward process runs for a finite interval T. Consequently, an error in the learned score at any intermediate time is inherited directly by the generated samples, with no asymptotic correction available. The bullet list in Section 2 notes the time dependence and finite-time runs, but the paper should state explicitly that Eq. (4) alone does not imply p_T(phi) approximately equal to p(phi); establishing that requires a separate, quantitative bound on the score error or an explicit validation of the final ensemble.","section":"Section 2, Eqs. (4)-(5)"},{"comment":"The paper's own toy-model diagnostic shows that the learned drift and effective action deviate strongly for large |phi|, and the accompanying text says 'the diffusion model can only learn where data is available'. These are precisely the tail regions that contribute to the Binder cumulant and to higher-order cumulants in the broken phase, so the deviation is not a peripheral caveat. The two-dimensional susceptibility and cumulant comparisons are only cited to Refs. [20,23] and are not shown here; this proceedings therefore contains no quantitative evidence that the tail error is controlled. Please add a quantitative tail diagnostic, for example a histogram ratio, a tail-restricted moment, or a comparison of the learned and exact actions on the full support, or alternatively state explicitly that representativeness on rare configurations is assumed rather than demonstrated.","section":"Section 3, Fig. 4"},{"comment":"The diffusion model is trained on HMC configurations and then checked against observables computed on the same HMC ensemble. Such a check largely tests the model's ability to reproduce its training distribution, not whether it correctly samples p(phi) proportional to exp[-S(phi)] in regions that are under-represented in the training data. This is not a formal circularity in the derivation, but it limits the evidential force of the reported agreement. An out-of-sample test, such as a different lattice volume, a different coupling, or an observable that was not used during training, would materially strengthen the central claim that the generated configurations are representative.","section":"Section 3, training and validation"},{"comment":"The Outlook correctly states that making the algorithm exact by an accept-reject step is important and that work in this direction is in progress, and Fig. 2 hedges with 'if all algorithms are working well'. However, the abstract and parts of the introduction describe the method as a way to 'generate configurations' without the same qualification. Until an accept-reject correction or an equivalent exactness mechanism is included, direct samples from the trained model should be described as approximate proposals rather than as configurations drawn from the target distribution; the current wording overstates what is demonstrated.","section":"Section 4 and Fig. 2"}],"minor_comments":[{"comment":"The abbreviation 'LTFs' is used for 'lattice field theories'; the standard abbreviation is 'LFT' and should be made consistent.","section":"Abstract and Section 1"},{"comment":"The name 'DALLE-E' appears twice; the correct product name is 'DALL-E'.","section":"Section 2"},{"comment":"The bottom row would be more informative if it included a quantitative comparison, such as sample means, variances, or a distance metric, rather than only overlaid histograms.","section":"Fig. 4"},{"comment":"The action in Eq. (6) is given with mu^2 = +/-1, but the figure caption refers to single-well and double-well cases; please state explicitly which mu^2 value corresponds to each column.","section":"Eq. (6) and Fig. 4"}],"recommendation":"major_revision","confidential_remarks":"This is a conference proceedings paper, and much of the quantitative support for the central claim lives in Refs. [20,23]. My major-revision recommendation is intended to ask for a scoped statement of what is proven versus assumed, together with a quantitative tail diagnostic, rather than to dispute the underlying idea, which is plausible and honestly presented. The main risk is that the proceedings text, as written, may be read as claiming more than the evidence in this manuscript establishes."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a LATTICE proceeding that summarises work the same group already published, chiefly Ref. [20]. If you read it expecting new results, you will be disappointed: the figures are republished, the equations are the same, and the quantitative evidence is deferred. That is normal for a talk write-up, and judged on that basis it does its job.\n\nWhat the paper does well is exposition. The relation between the backward diffusion process, Eq. (4), and stochastic quantisation, Eq. (5), is presented cleanly, and the bullet-point comparison between the two is genuinely useful for a reader who has not followed the ML/lattice literature. The paper is also honest: it says explicitly that the diffusion model learns only where data is available, and the toy-model figure shows the learned drift deviating from the exact drift at large |phi|. The original papers are cited throughout, and the data/code statement points to Ref. [20], so the authors are not hiding the ball.\n\nThe soft spot is exactly the one the stress-test flags. Eq. (4) and Eq. (5) looking alike does not by itself establish that a trained diffusion model generates configurations representative of exp(-S). That requires the learned score to be accurate over the support of the target distribution, including low-density tails, and on the evidence shown here those tails are where the model deviates. The Binder cumulant and higher-order cumulants in the broken phase are sensitive to exactly those regions. The paper acknowledges the limitation in words but does not quantify it, and the acceptance/rejection analysis that would restore representativeness is only cited, not shown. So the central claim is conditional, not demonstrated. I would not call this a load-bearing flaw in a proceedings, because the authors are careful to say their generation runs are direct generations and to point to the detailed papers for acceptance rates. But a reader relying solely on this paper cannot judge how representative the generated ensemble really is.\n\nThe citation practice looks fine: it points to the prior JHEP paper and related normalising-flow work, and self-citation here is appropriate because the work is being summarised. There is no sign of circular reasoning beyond what the method inherently involves, namely training on HMC data and then comparing against HMC observables, and that caveat is also stated.\n\nWho is this for? A lattice or ML grad student wanting a quick, readable entry point into diffusion models for lattice field theory, not someone looking for a new result. It deserves a serious referee in the sense that a proceedings referee should check that the summary accurately represents the underlying papers and that the claims are appropriately hedged; both hold. I would not cite it in my own work, since Ref. [20] is the citable source, but I would happily put it on a reading-group list.","headline":"A clear, honest proceedings summary of the authors' earlier diffusion-model work; the stochastic-quantisation analogy is sound, the tail-coverage caveat is real and explicitly acknowledged, and there is no new result here.","tokens_in":6889,"tokens_out":1549,"would_cite":false,"duration_ms":17257,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["81T80","60H10","81T25"],"pacs":["11.15.Ha","05.10.Ln"],"model":"deepseek-v4-flash","headline":"Diffusion models generate lattice fields via stochastic quantisation","keywords":["diffusion models","stochastic quantisation","lattice field theory","Langevin dynamics","score-based generative models","scalar field theory","Hybrid Monte Carlo","phi^4 theory"],"falsifier":"Train a variance-exploding diffusion model on a $32^2$ $\\lambda\\phi^4$ ensemble in the broken phase whose training data contains only configurations in one of the two wells, then generate a fresh ensemble and measure the fraction of configurations in the other well; if the generated ensemble does not reproduce the exact two-peak distribution, or equivalently if the learned score at large $|\\phi|$ deviates measurably from $-\\nabla S$, the backward process fails to sample the target distribution.","tokens_in":5852,"feed_emoji":"⚛️","tokens_out":5699,"duration_ms":45121,"temperature":0.7,"pith_summary":"The paper argues that the backward (denoising) process of a diffusion model is, in form, the same stochastic differential equation used in stochastic quantisation, with the learned score playing the role of the drift given by the action gradient. It then demonstrates that a diffusion model trained on Hybrid Monte Carlo ensembles of a two-dimensional scalar $\\lambda\\phi^4$ theory generates new configurations whose ensembles reproduce the target distribution $p(\\phi) \\propto e^{-S(\\phi)}$. The practical stakes are that generative models could supplement or replace traditional samplers in lattice field theory, producing configurations faster and with reduced auto-correlation. The paper is a high-level overview of a programme reported in detail in earlier papers, rather than a self-contained derivation.","feed_headline":"Diffusion models generate lattice fields via stochastic quantisation","feed_subtitle":"The backward denoising process matches the Langevin equation, so generated fields reproduce the target distribution.","key_machinery":"The load-bearing object is the score function $\\nabla_\\phi \\log p_\\tau(\\phi)$, the gradient of the logarithm of the time-dependent probability density, which appears in the backward process and is approximated by a neural network. In stochastic quantisation the corresponding drift is $-\\nabla S(\\phi)$ for a known action, so the identification $p(\\phi,\\tau) = \\exp[-S(\\phi,\\tau)]/Z$ converts the diffusion backward equation into a Langevin equation with a time-dependent drift and noise amplitude. The variance-exploding scheme, with no drift in the forward process and $g(t)=\\sigma t/T$, makes the forward dynamics a pure noising process and the backward dynamics a denoising Langevin evolution over finite time intervals. This identification is what turns a generative model into a sampler: the neural network learns the drift needed to steer noise back onto the target distribution.","core_discovery":"On the paper's own terms, the central discovery is the formal identity between the backward process of a diffusion model and the stochastic quantisation Langevin equation. Writing the target distribution as $p(\\phi,t) = (1/Z) \\exp[-S(\\phi,t)]$, the backward SDE becomes $\\partial_\\tau \\phi(x,\\tau) = -g^2(T-\\tau) \\nabla S(\\phi,T-\\tau) + g(T-\\tau) \\eta(x,\\tau)$, which matches the stochastic quantisation equation $\\partial_\\tau \\phi(x,\\tau) = -\\nabla S(\\phi,\\tau) + \\sqrt{2}\\,\\eta(x,\\tau)$ up to time-dependent noise normalisation. Because of this match, a diffusion model trained on existing lattice configurations, where the score $\\nabla_\\phi \\log p_\\tau(\\phi)$ is learned by a neural network, can generate new configurations representative of the target theory. The paper demonstrates this concretely for $\\lambda\\phi^4$ theory on a $32^2$ lattice in both symmetric and broken phases, checking susceptibility, Binder cumulant, and higher-order cumulants, and reporting that generated configurations can serve as proposals with reduced auto-correlation.","pith_inferences":["The formal identity means that improvements to diffusion-model score estimation, such as better architectures or exact score matching, should transfer directly to better samplers for lattice theories; the stochastic-quantisation literature on kernels may offer a principled way to reduce finite-time bias in the backward process.","The paper's caveat that the model can only learn where data is available suggests a concrete failure mode for theories with rare topological sectors or metastable phases: a training ensemble that under-samples a sector will produce a biased ensemble, and this could be tested by measuring the fraction of configurations in each sector against the exact value.","Because the noising time $T$ is finite, the backward process is a biased sampler unless a Metropolis accept-reject step is included; this suggests that diffusion samplers in lattice field theory will in practice be used as proposal generators, not standalone samplers, unless the bias can be quantified and controlled."],"forward_implications":["Diffusion models trained on existing ensembles generate new configurations that approximate the target distribution $p(\\phi) \\propto e^{-S(\\phi)}$.","Because each run starts from fresh noise, generated configurations are not serially correlated like HMC chains; they can be used as proposals in a Markov chain with reduced auto-correlation.","The learned time-dependent drift and effective action can be read off and compared with the exact action, giving a diagnostic of how well the model has learned the theory.","The same framework extends to U(1) gauge theories and, in progress, to theories with a complex action via complex Langevin, where the target distribution is not known a priori.","Making the algorithm exact requires an accept-reject step; the paper reports encouraging acceptance rates in the detailed studies."],"supporting_citations":[{"why":"Establishes the diffusion-model/stochastic-quantisation relation and provides the scalar-field results this talk summarizes.","marker":"[20]"},{"why":"Supplies the diffusion-model framework of forward noising and backward denoising as a non-equilibrium thermodynamic process.","marker":"[29]"},{"why":"Introduces stochastic quantisation, the Langevin equation the backward process is compared with.","marker":"[30]"},{"why":"Review of stochastic quantisation that supplies the kernel-generalisation footnote and the stationary-distribution argument.","marker":"[31]"},{"why":"Analysis of higher-order cumulants for variance-exploding and variance-preserving schemes, used to validate generated ensembles.","marker":"[23]"},{"why":"Hybrid Monte Carlo, the baseline sampler that produces the training data and the comparison chain.","marker":"[4]"},{"why":"Flow-based generative sampling for lattice field theory, the main alternative generative approach the paper compares against.","marker":"[5]"}],"fun_headline_variants":["Diffusion denoising is stochastic quantisation","Backward diffusion matches stochastic quantisation","Lattice fields from diffusion via stochastic quantisation","Generating lattice configs with diffusion models","Diffusion models and stochastic quantisation link"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire scheme rests on the learned score $\\nabla \\log p_\\tau(\\phi)$ being accurate wherever the target distribution has support, but the model can only learn the score where training data exists, so rare field configurations contribute little to the training objective and may be generated with the wrong weight.","fun_headline_variants_meta":{"raw":{"variants":["Diffusion denoising is stochastic quantisation","Backward diffusion matches stochastic quantisation","Lattice fields from diffusion via stochastic quantisation","Generating lattice configs with diffusion models","Diffusion models and stochastic quantisation link"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000203,"raw_usage":{"total_tokens":1322,"prompt_tokens":819,"completion_tokens":503,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":435,"completion_tokens_details":{"reasoning_tokens":436}},"tokens_in":435,"tokens_out":503,"duration_ms":5137,"temperature":1.0,"reasoning_tokens":436,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T12:52:37.340235+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train a variance-exploding diffusion model on a $32^2$ $\\lambda\\phi^4$ ensemble in the broken phase whose training data contains only configurations in one of the two wells, then generate a fresh ensemble and measure the fraction of configurations in the other well; if the generated ensemble does not reproduce the exact two-peak distribution, or equivalently if the learned score at large $|\\phi|$ deviates measurably from $-\\nabla S$, the backward process fails to sample the target distribution.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces stochastic quantisation, the Langevin equation the backward process is compared with."},{"cited_title":"Damgaard and H","cited_arxiv_id":null,"evidence_quote":"Review of stochastic quantisation that supplies the kernel-generalisation footnote and the stationary-distribution argument."}],"review_version":1}