{"id":"3734aa6b-b448-47c6-b877-0650ed7a60ae","arxiv_id":"2607.08505","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":4,"one_line_summary":"Fully convolutional diffusion models trained on small lattices transfer to unseen larger volumes for 2D/3D phi^4 sampling across phases, matching or beating same-size training on most observables.","lead":"Diffusion models can generate lattice field configurations for phi^4 theory near criticality, and a model trained only on small lattices can sample much larger ones without retraining. This offers a practical route around critical slowing-down in lattice simulations by learning the score from cheap small-volume data.","discovery_kind":"extension","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified","rationale":"The paper’s strongest claim is carefully scoped: multi-L training transfers the propagator and most scalars to unseen L=64, improves several critical observables relative to in-distribution L=64, and leaves a residual susceptibility excess mainly in the broken phase. That claim is backed by direct HMC comparisons, single-size controls, unfiltered tables, MALA diagnostics at the unseen size, and released code. The residual zero-mode bias is real and is the softest point, but it is already quantified rather than assumed away; optional exact MALA refinement further reduces it. No hidden inconsistency or untested premise is required for the stated result to hold. The reader’s weakest_assumption correctly identifies the IR/zero-mode issue, yet the evidence already addresses it sufficiently for an ACCEPT verdict. A modest σ-recalibration check would still be worth running, but a negative outcome would only tighten the claim, not reverse it. Therefore the verdict remains ACCEPT with no change.","tokens_in":54689,"tokens_out":500,"duration_ms":5927,"concrete_test":"Recompute the L=64 near-critical and broken-phase rows of Tab. IV (and unfiltered Tab. VI) after replacing the coarse pairwise-distance σ choices with a single σ calibrated on the multi-L training set only (no L=64 data); if the multi-L advantage over in-distribution L=64 for χ and the higher zero-mode cumulants disappears, the transfer claim would need qualification.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that multi-L training of a fully convolutional score network transfers to unseen L=64 and can improve critical observables relative to in-distribution L=64 training—is supported by the paper’s own evidence rather than resting on an untested leap. Sec. VI and Figs. 16–19 show the multi-L propagator tracks HMC from UV to k_min; Tab. IV and the unfiltered Tab. VI show most scalar observables (including several critical zero-mode entries) closer to HMC than the in-distribution baseline; single-size baselines isolate the multi-volume contribution; residual susceptibility excess and action-density bias are quantified and partially corrected by exact MALA. The reader’s weakest_assumption correctly flags residual IR sensitivity, but that residual is already measured and does not overturn the transfer result as stated.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The manuscript studies variance-exploding score-based diffusion models as generative samplers for lattice φ⁴ theory in D=2 and D=3 across the symmetric, near-critical, and broken phases. Ensembles from reverse-SDE sampling are validated against FA-HMC–Wolff references on scalar observables, joint (magnetization, action-density) distributions, single-site cumulants, and the momentum-space propagator G(|k|). Residual bias is reported mainly in the zero-mode (especially susceptibility in the broken phase) and, in 3D, the action density. The authors introduce a small-t score diagnostic, a MALA acceptance diagnostic, and an HMC-referenced MSE-based ESS. Using a fully convolutional U-Net with circular padding and shared weights, they demonstrate cross-volume training: a 3D model trained on L∈{4,8,16,32} samples the unseen L=64 lattice and, at criticality, improves several observables relative to in-distribution L=64 training, establishing multi-L score transfer as a route to large-volume sampling.","tokens_in":54910,"tokens_out":1147,"duration_ms":25376,"significance":"Critical slowing down remains a central bottleneck in lattice field theory. Demonstrating that a local, fully convolutional score trained only on cheaper small volumes can generate usable ensembles at an unseen larger volume—with propagator agreement from UV to IR and competitive or better critical observables—is a concrete and practically relevant result for generative sampling. The multi-L versus single-size extrapolation comparison isolates the value of multi-volume finite-size information. The local diagnostics and HMC-referenced ESS are reusable tools for the community. Residual IR and action-density biases are quantified rather than hidden, and code is released. If the cross-L transfer continues to hold under gauge-equivariant and fermion-aware extensions, the approach would matter for full QCD volume scaling.","major_comments":[{"comment":"App. A (and the 3D rows of Tabs. III–VI): the two support filters (site field range and magnetization support) are defined from the target-volume L=64 HMC ensemble and reject up to ~200/512 in-distribution broken-phase samples. The central practical claim—that the score from cheap small lattices transfers to the target volume without retraining—is weakened if production sampling still needs large-L HMC to define acceptance windows. The unfiltered Tabs. V–VI and the exact MALA refinement rows help, but the main text (Sec. IV, Sec. VI, abstract) should state filter rates, that filters use target-volume HMC, and which claims survive without them or after MALA alone.","section":null},{"comment":"Sec. VI and Tab. IV (broken phase, 3D): the residual susceptibility excess at unseen L=64 remains the main exception even for the largest multi-L model (χ≈9.3(6) filtered vs HMC 4.9(1); worse unfiltered). The paper correctly attributes this to intra-sector zero-mode width. Because the abstract’s transfer claim is otherwise strong, the main text should state more sharply which observables are reliable after pure reverse-SDE transfer and which still require exact MALA (score or −∇S) before the ensembles can be used for precision IR physics.","section":null}],"minor_comments":[{"comment":"Sec. III B / sampling appendices: reverse-SDE uses 2000 EM steps (log grid in 2D, linear in 3D). A short wall-clock or NFE comparison to FA-HMC–Wolff and a note on whether EDM/DPM-style fewer-step solvers preserve the reported G(|k|) would help readers assess cost.","section":null},{"comment":"Sec. VI: cross-L noise scales σ are described as coarse pairwise-distance choices and differ by training set and κ. A one-paragraph sensitivity check (or fixed-σ ablation) would strengthen the claim that multi-L transfer is not an artifact of σ retuning.","section":null},{"comment":"Fig. 1 caption and Tab. III: the broken-phase peak broadening is said to be invisible by eye yet quantified by χ; consider adding a one-line inset or quoting σ_DM/σ_HMC in the caption for readability.","section":null},{"comment":"App. E: residual addition inside ResBlocks is omitted for stability. A brief note on whether this choice affects conservativeness (App. G) or MALA acceptance would connect architecture to the diagnostics.","section":null},{"comment":"Notation: Σ(t) vs Σ_max and the normalized field ˆϕ appear in several places; a short symbol table or consistent first-use definitions in Sec. III would reduce cross-referencing load.","section":null}],"recommendation":"minor_revision","confidential_remarks":"Solid, carefully validated hep-lat ML paper; the multi-L transfer result is the real contribution and is backed by tables and single-size baselines. Support filters are the main presentational/practical soft spot, not a fatal circularity, because unfiltered and MALA-refined results are already in the appendices. Suitable for the journal after a focused revision that surfaces filter dependence and IR caveats in the main text. No novelty or citation concerns."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The real news here is the quantitative multi-L transfer: a 3D score network trained only on L in {4,8,16,32} generates L=64 ensembles that match the FA-HMC–Wolff propagator from UV to k_min and improve several critical zero-mode observables relative to an in-distribution L=64 model. That is new relative to the existing diffusion-for-lattice literature (including prior work by overlapping authors), and the paper backs it with tables, unfiltered appendices, and single-size extrapolation baselines that isolate the multi-volume contribution.\n\nWhat they do well is the validation discipline. They compare against a strong reference (Fourier-accelerated HMC plus Wolff), report scalar observables, joint (M,s) distributions, single-site cumulants, and the full G(|k|), and introduce three useful diagnostics: small-t score vs −∇S, MALA acceptance of the learned drift, and an HMC-referenced MSE-based ESS that folds bias and variance together. Residual problems—broken-phase susceptibility excess (intra-sector zero-mode width), 3D action-density shift from single-site moment cancellations, and the need for support filters on the in-distribution runs—are stated and measured, not papered over. Code is released. The math and citation pattern look clean; the free-field ESS check and non-conservative score appendix are careful extras.\n\nSoft spots are real but proportionate. Zero-mode structure remains the weak point; single-size L=4→L=64 fails badly, so multi-L finite-size information is doing the work, and residual IR bias still needs optional exact MALA (or better schedules) for full accuracy. Noise scales σ for the cross-L runs are coarse, and the support filters, while honestly documented, are an engineering crutch. None of that overturns the central transfer claim as written.\n\nThis is for people who care about generative sampling near criticality and about whether learned local scores can be reused across volumes. It deserves a serious referee. I would engage with it and expect it to influence how people train large-volume samplers.","headline":"Solid, carefully validated demonstration that multi-L convolutional score training transfers to larger unseen lattices for φ⁴, with residual IR bias quantified rather than hidden.","tokens_in":55502,"tokens_out":566,"would_cite":true,"duration_ms":19024,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"A diffusion score trained only on small lattices can generate accurate large-volume φ⁴ ensembles without retraining.","keywords":["lattice field theory","diffusion models","score-based generative models","φ⁴ theory","critical sampling","cross-volume generalization","reverse SDE","Metropolis-adjusted Langevin"],"falsifier":"Generate unfiltered reverse-SDE ensembles at L=64 from a model trained only on L∈{4,8,16,32} and compare the broken-phase and critical susceptibilities and zero-mode cumulants to FA-HMC–Wolff; a growing, uncorrectable excess as the volume ratio increases would falsify clean transfer.","tokens_in":55581,"feed_emoji":"⚛️","tokens_out":887,"duration_ms":22338,"temperature":0.7,"pith_summary":"This paper shows that generative diffusion models can sample two- and three-dimensional lattice φ⁴ theory across the symmetric, near-critical, and broken phases. The reverse stochastic process, driven by a learned score, reproduces scalar observables and the momentum-space propagator when checked against high-quality Monte Carlo references, with leftover bias mainly in the zero mode. The decisive result is architectural: a fully convolutional score network trained on several small volumes can be evaluated at a larger lattice never seen in training. In three dimensions a model trained on L up to 32 matches or improves critical observables at L=64 relative to a model trained at L=64 itself. That transfer means expensive large-volume training data are not always required for large-volume sampling.","feed_headline":"Small-lattice scores sample large-volume field theory","feed_subtitle":"A multi-volume diffusion model matches or beats target-size training at L=64 without seeing that size","key_machinery":"The fully convolutional U-Net score network with circular padding and volume-shared weights; it learns a local drift that is evaluated at any lattice size, and is used inside the reverse variance-exploding SDE (with optional Metropolis-adjusted Langevin correction) to transport noise to the Boltzmann measure.","core_discovery":"Cross-volume generalization works for score-based diffusion samplers of lattice φ⁴ theory: a fully convolutional network whose weights are shared across volumes, when trained on many cheap small-lattice ensembles, produces reverse-SDE samples at an unseen larger volume that reproduce the propagator and most scalar observables, and at criticality can outperform an otherwise identical model trained only at the target size.","pith_inferences":["Multi-volume training appears to act as an infrared regularizer: seeing the zero-mode at several finite sizes constrains long-wavelength score components more effectively than a single large volume.","The same transfer logic, if gauge-equivariant and fermion-aware scores can be trained, would let small-lattice QCD ensembles seed large-volume generation.","Residual action-density bias in three dimensions is amplified by on-site cancellations; architectures that match single-site moments more tightly should suppress it without extra volume.","A controlled volume-ratio scan beyond factor-of-two linear size would map where multi-L transfer remains quantitative versus where infrared degradation reappears."],"forward_implications":["Large-volume scalar ensembles can be produced from reference data generated only on cheaper small lattices, without retraining at the target size.","Independent reverse trajectories replace long autocorrelated Markov chains once the multi-volume score is learned, reducing critical-slowing-down cost in practice.","In-distribution training at the largest volume is not automatically optimal; multi-L training can improve critical infrared observables.","Optional exact MALA refinement with the action gradient can remove residual zero-mode bias while retaining most of the equilibration gain from the diffusion proposal."],"fun_headline_variants":["Small-lattice scores transfer to large-volume φ⁴ sampling","Cross-volume diffusion matches L=64 without seeing that size","Shared convolutional scores sample unseen large φ⁴ lattices","Many cheap small volumes train scores for large lattice field theory","Reverse-SDE scores from multi-size training hit critical large lattices"],"cache_read_input_tokens":40192,"weakest_assumption_plain":"That multi-volume training on several small lattices supplies enough finite-size infrared information for the shared local score to control the zero-mode distribution on a substantially larger unseen lattice.","fun_headline_variants_meta":{"raw":{"variants":["Small-lattice scores transfer to large-volume φ⁴ sampling","Cross-volume diffusion matches L=64 without seeing that size","Shared convolutional scores sample unseen large φ⁴ lattices","Many cheap small volumes train scores for large lattice field theory","Reverse-SDE scores from multi-size training hit critical large lattices"]},"model":"grok-4.5","effort":"low","cost_usd":0.004094,"raw_usage":{"total_tokens":1263,"prompt_tokens":823,"num_sources_used":0,"completion_tokens":86,"cost_in_usd_ticks":40940000,"prompt_tokens_details":{"text_tokens":823,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":354,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":823,"tokens_out":86,"duration_ms":4410,"temperature":1.0,"reasoning_tokens":354,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-10T06:19:12.522115+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Generate unfiltered reverse-SDE ensembles at L=64 from a model trained only on L∈{4,8,16,32} and compare the broken-phase and critical susceptibilities and zero-mode cumulants to FA-HMC–Wolff; a growing, uncorrectable excess as the volume ratio increases would falsify clean transfer.","supporting_citations":[],"review_version":1}