REVIEW 3 major objections 7 minor 39 references
Future Requirements of Lattice Field Theory Calculations on European High-Performance Computing Facilities
T0 review · 3 major / 7 minor · reviewed 2026-07-31 · grok-4.5
Pith's one-line read Lattice QCD will keep advancing on European supercomputers only if machines prioritise memory bandwidth, fast interconnects and stable double precision—not peak FLOPs alone.
desk verdict Solid EuroHPC requirements brief: accurate on today’s Dirac/bandwidth profile, thin on how durable that profile stays if ML sampling or FP64 emulation mature. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Repeated application of the lattice Dirac operator inside iterative sparse linear solvers (in Hybrid Monte Carlo configuration generation and in large-scale measurements). These bandwidth-bound, nearest-neighbour stencil kernels—not peak FLOPs—set the hardware, software and scaling requirements the paper derives.
What would settle it
If EuroHPC systems chosen after including lattice-style bandwidth and communication benchmarks still deliver the same strong-scaling efficiency and double-precision sustained fraction for Dirac solvers as current AI-oriented machines, or if multi-year lattice software investment does not rise once a public roadmap exists, the policy half of the claim fails.
Extended reading notes
Core claim
Sustained progress of lattice field theory on European HPC requires balanced architectures—high memory bandwidth, low-latency high-bandwidth interconnects, large on-device memory, and stable double-precision performance—plus portable software ecosystems and dedicated human expertise; peak floating-point performance and AI-driven low-precision hardware alone are a poor match for Dirac-solver-dominated workloads.
Load-bearing premise
That putting community-specific benchmarks into procurement and publishing a clearer European post-exascale roadmap will actually change which machines get built and how groups invest in long-lived code.
Editorial extensions
If this is right
- Procurement suites that include lattice-like bandwidth, latency and strong-scaling tests would favour more balanced memory-to-compute ratios.
- Lattice codes should sit in HPC centres’ continuous testing pipelines so toolchain and hardware regressions are caught early.
- Stable posts for research software engineers become a prerequisite for using next-generation machines effectively.
- Mixed precision can accelerate inner solver iterations, but double (and sometimes extended) precision remains mandatory for observables and long trajectories.
- Shared gauge-field ensembles continue to amortise the high fixed cost of configuration generation across many physics analyses.
Reading between the lines
- Other stencil-heavy, communication-bound fields named in the paper—hydrodynamics and numerical gravity—would gain from the same balanced procurement criteria.
- If accelerator roadmaps keep cutting native double precision, verified high-precision emulation becomes a shared infrastructure problem, not a lattice-only patch.
- Without European-scale co-design support, lattice groups risk fragmenting effort across many small, architecture-specific codes as platforms diversify.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This is a white-paper-style contribution to the EuroHPC User Days 2026 proceedings, authored by representatives of the European lattice field theory community (EuroLFT). It describes the computational profile of lattice QCD — HMC-based gauge-field generation dominated by sparse Dirac-operator inversions, measurement pipelines limited by signal-to-noise degradation, and emerging machine-learning-assisted workflows — and argues that these workloads are bandwidth- and latency-bound rather than FLOP-bound. From this it derives hardware requirements (high memory bandwidth, low-latency high-bandwidth interconnects, large on-device memory, sustained double-precision performance, mixed-precision support, topology-aware networks), software needs (performance portability, continuous integration at computing centres), human-resource needs (dedicated performance-engineering support, PASC-like European structures), and policy recommendations (community benchmarks in procurement, a European post-exascale roadmap). The descriptive material is standard and well referenced (HMC, critical slowing down, multigrid, FLAG, ILDG), and the resource figures (Fig. 1; Fig. 2a–b) are illustrative rather than derived results.
Significance. If taken up by EuroHPC stakeholders, the paper provides a credible, field-backed statement of what bandwidth-bound scientific workloads need from post-exascale European systems, and it does so with concrete evidence rather than assertion: per-workload resource-investment data from the OpenLAT collaboration (Fig. 1), documented order-of-magnitude speedups from multigrid solvers and GPU offloading in tmLQCD (Fig. 2a), and HiRep scaling measurements to 1024 GPUs on LUMI-G and JUWELS Booster. The articulation of why peak-FLOP and AI-oriented low-precision procurement trends are a poor match for Dirac-solver-dominated workloads is the most useful contribution for a policy-facing venue. The requirements are falsifiable in principle (benchmarkable against delivered systems) and are consistent with the cited algorithmic literature.
major comments (3)
- [Fig. 2b and §2.3] Figure 2b (§2.3) appears to embed multiple full pages of ref [38] (Drach et al., Comput. Phys. Commun. 322 (2026) 110061) — including its Table 6, Fig. 31, §§9.3–10 and Eqs. (138)–(144), with the same block repeated several times — rather than a clean strong/weak-scaling plot for HiRep. If this is not an artifact of the arXiv PDF text layer, the figure as submitted is illegible, reproduces substantial material from another publication, and cannot serve its stated evidentiary role (demonstrating 'excellent strong-scaling performance'). Please replace it with a self-contained plot of the HiRep scaling data with proper attribution, and verify the production version carefully.
- [§2.2 / §3.1 (double-precision requirement)] The hardware requirement that 'high and stable double precision performance remains essential' (§3.1, fourth bullet; echoed in §2.2 and the Abstract) is the most forward-looking and procurement-relevant claim in the paper, but its durability is asserted rather than argued. The manuscript itself cites the two trends that could erode it: generative models for sampling (§2.1, refs [25]–[30]) and emulated high-precision arithmetic (§2.2: 'may be feasible, if performant and sufficiently verified'). If either matures, AI-oriented low-precision-heavy hardware would become substantially more adequate for lattice workloads. For a requirements document this does not invalidate the claim — current production workloads are unambiguously DP-critical — but the paper should quantify or at least delineate the robust core: e.g., what fraction of the projected ensemble-generation and measurement cycle (lo
- [§3.4 (policy recommendations)] The two main policy proposals — incorporating community-specific benchmarks into EuroHPC procurement and publishing a 5/10-year post-exascale roadmap — are stated as self-evidently beneficial, with no evidence or precedent that benchmark inclusion changes awarded system balance, or that roadmap visibility (rather than funding level or vendor roadmaps) is the binding constraint on long-term code investment. The PRACE scientific case (ref [40]) is cited but not used to support the causal claim. Either cite concrete precedents where representative application benchmarks demonstrably shaped procurement outcomes (e.g., US DOE/NNSA benchmark-driven procurements, or the role of lattice codes in BlueGene co-design, which the Conclusions themselves invoke), or temper the language from prescription to recommendation.
minor comments (7)
- [Fig. 1 caption and axis labels] The ordinate label '[Mch]CPU or [Knh]GPU' is ambiguous: CPU million-core-hours and GPU thousand-node-hours are mixed in one plot, and the legend includes a single 'US' series with no definition or source. Please state units per series, the conversion (if any) used to make CPU and GPU investment comparable, and what the US bar represents.
- [Fig. 2a] The axis legend reads 'JuQueen'; the standard spelling is 'JUQUEEN'. Similarly check 'Juwels-Booster' (JURECA/JUWELS Booster).
- [Keywords and typos] Missing space ('Keywords:Lattice'); 'high- performance' broken across lines in the Abstract; inconsistent capitalisation of collaboration names ('OpenLAT' vs 'openLAT' in §2.3).
- [§2.3, collaboration list] The collaboration list includes 'TWEXT'; please verify this is the intended name, as it may be unfamiliar to some readers (a reference or expansion would help).
- [§1, muon g−2 motivation] The statement that g−2 'showed a tension between experimental measurements and theoretical predictions' should be updated or nuanced in light of ref [1] (the 2025 Theory Initiative update) and the most recent experimental/lattice situation, since the paper cites precisely that update.
- [References] Ref [6] is listed as Phys. Rev. D 113 (2026) — please check the year/volume; ref [38] uses an https DOI in the doi field; ref [40] lacks an author/organisation field (PRACE).
- [§2.1, lattice sizes and degrees of freedom] The claim 'typically leads to lattice spacings a~0.04–0.1 fm' and 'order 10^10 degrees of freedom' would benefit from one clarifying sentence on how the 12 spin-colour-per-site factor enters, for the non-lattice HPC audience this proceedings addresses.
Circularity Check
No circular derivation: descriptive HPC-requirements paper; algorithmic bottlenecks and resource plots are empirical/structural, not self-defined predictions.
full rationale
This is a community position/requirements paper for EuroHPC, not a first-principles derivation of a target observable from fitted parameters. The central claims (Dirac-solver kernels are bandwidth- and communication-bound; future systems need high memory bandwidth, low-latency interconnects, large on-device memory, sustained double precision, portable software, and dedicated human expertise) follow from the stated structure of HMC plus iterative sparse Dirac solves, multigrid/deflation, and nearest-neighbour halo exchange—standard lattice practice illustrated by external and community scaling data (e.g. tmLQCD cost curves, HiRep LUMI-G/JUWELS scaling). Figure 1 reports invested OpenLAT resources as historical fact, not a fit re-labeled as prediction. Self-citations document the authors’ codes and allocations but do not supply a uniqueness theorem or force the procurement argument by construction. No step reduces a claimed prediction to its own inputs. Forward-looking caveats (ML sampling, emulated FP64) are hedges about durability of assumptions, not circularity.
Assumptions & free parameters
assumptions (5)
- domain assumption Physical observables in lattice QCD are estimated by Monte Carlo sampling of gauge configurations with cost dominated by iterative solves of the sparse Dirac operator.
- domain assumption Dirac-operator and multigrid/deflated solver kernels are memory-bandwidth and communication bound rather than peak-FLOP bound on modern architectures.
- domain assumption Reducing lattice spacing and approaching physical quark masses increases volume and autocorrelation costs enough that leadership-class sustained access remains necessary.
- domain assumption Reliable double-precision (or verified emulated high precision) remains necessary for many observables and long MD trajectories despite mixed-precision inner solvers.
- ad hoc to paper Including lattice-like community benchmarks in EuroHPC procurement and publishing a multi-year post-exascale roadmap would improve system fitness for bandwidth-bound science.
Cite this review
Pith. "Pith review of Future Requirements of Lattice Field Theory Calculations on European High-Performance Computing Facilities." pith.science (2026). https://pith.science/paper/X4MIEJTB
@misc{pith2026260725061,
author = {Pith},
title = {Pith review of: Future Requirements of Lattice Field Theory Calculations on European High-Performance Computing Facilities},
year = {2026},
howpublished = {\url{https://pith.science/paper/X4MIEJTB}},
note = {Machine review of arXiv:2607.25061}
}
read the original abstract
Lattice field theory provides a first-principles framework for studying properties of strongly interacting quantum field theories in elementary particle physics. Researchers in lattice field theory are also among the largest and most efficient users of high- performance computing resources in fundamental science. In this contribution, we outline the computational profile of lattice QCD, from gauge-field generation to large-scale measurements, and discuss the main hardware, software, and human resource requirements needed to sustain progress on current and future European HPC infrastructures.
Figures
Reference graph
Works this paper leans on
- [38]
-
[25]
K. Cranmer, G. Kanwar, S. Racani `ere, D. J. Rezende, P. E. Shanahan, Advances in machine-learning-based sampling motivated by lattice quantum chromodynamics, Nature Rev. Phys. 5 (9) (2023) 526–535.doi:10.1038/s42254-023-00616-w
-
[30]
S. Caron, et al., Strategic white paper on AI infrastructure for particle, nuclear, and astroparticle physics: insights from JENA and EuCAIF, Mach. Learn. Sci. Tech. 7 (1) (2026) 013002.doi:10.1088/2632-2153/ae35cd
-
[40]
The scientific and innovation case for computing in europe 2026–2034, iSBN: 9789464597738 (Apr. 2026). URLhttps://prace-ri.eu/scientific-case/the-scientific-and-innovation-case-for-computing-in-europe-2026-2034/
2026
-
[1]
Aliberti, et al., The anomalous magnetic moment of the muon in the Standard Model: an update, Phys
R. Aliberti, et al., The anomalous magnetic moment of the muon in the Standard Model: an update, Phys. Rept. 1143 (2025) 1–158.doi: 10.1016/j.physrep.2025.08.002
-
[2]
J. de Blas, M. Dunford, E. Bagnaschi, A. Freitas, P. P. Giardino, et al., Physics Briefing Book: Input for the 2026 update of the European Strategy for Particle Physicsdoi:10.17181/CERN.35CH.2O2P
-
[3]
G. Aarts, et al., Phase Transitions in Particle Physics: Results and Perspectives from Lattice Quantum Chromo-Dynamics, Prog. Part. Nucl. Phys. 133 (2023) 104070.doi:10.1016/j.ppnp.2023.104070
arXiv 2023
-
[4]
Aoki, et al., FLAG Review 2019: Flavour Lattice Averaging Group (FLAG), Eur
S. Aoki, et al., FLAG Review 2019: Flavour Lattice Averaging Group (FLAG), Eur. Phys. J. C 80 (2) (2020) 113.doi:10.1140/epjc/ s10052-019-7354-7
doi:10.1140/epjc/ 2019
Show all 39 references
-
[5]
Aoki, et al., FLAG Review 2021, Eur
Y . Aoki, et al., FLAG Review 2021, Eur. Phys. J. C 82 (10) (2022) 869.doi:10.1140/epjc/s10052-022-10536-1
2021 doi
-
[6]
Aoki, et al., FLAG review 2024, Phys
Y . Aoki, et al., FLAG review 2024, Phys. Rev. D 113 (1) (2026) 014508.doi:10.1103/nfzp-p5dn. G. Aarts et al./00 (2026) 1–1010
2024 doi
-
[7]
C. T. H. Davies, A. C. Irving, R. D. Kenway, C. M. Maynard, International lattice data grid, Nucl. Phys. B Proc. Suppl. 119 (2003) 225–226. doi:10.1016/S0920-5632(03)01509-3
2003 doi
-
[8]
A. C. Irving, R. D. Kenway, C. M. Maynard, T. Yoshie, Progress in building an International Lattice Data Grid, Nucl. Phys. B Proc. Suppl. 129 (2004) 159–163.doi:10.1016/S0920-5632(03)02517-9
2004 doi
-
[10]
M. G. Beckett, B. Joo, C. M. Maynard, D. Pleiter, O. Tatebe, T. Yoshie, Building the International Lattice Data Grid, Comput. Phys. Commun. 182 (2011) 1208–1214.doi:10.1016/j.cpc.2011.01.027
2011 doi
-
[11]
M. D. Wilkinson, et al., The FAIR Guiding Principles for scientific data management and stewardship, Sci. Data 3 (1) (2016) 160018. doi:10.1038/sdata.2016.18
2016 doi
-
[12]
Karsch, H
F. Karsch, H. Simma, T. Yoshie, The International Lattice Data Grid – towards FAIR data, PoS LATTICE2022 (2023) 244.doi:10.22323/ 1.430.0244
2023
-
[13]
Duane, A
S. Duane, A. D. Kennedy, B. J. Pendleton, D. Roweth, Hybrid Monte Carlo, Phys. Lett. B 195 (1987) 216–222.doi:10.1016/ 0370-2693(87)91197-X
1987
-
[14]
Del Debbio, G
L. Del Debbio, G. M. Manca, E. Vicari, Critical slowing down of topological modes, Phys. Lett. B 594 (2004) 315–323.doi:10.1016/j. physletb.2004.05.038
2004 doi
-
[15]
Schaefer, R
S. Schaefer, R. Sommer, F. Virotta, Critical slowing down and error analysis in lattice QCD simulations, Nucl. Phys. B 845 (2011) 93–119. doi:10.1016/j.nuclphysb.2010.11.020
2011 doi
-
[16]
Hasenbusch, Speeding up the hybrid Monte Carlo algorithm for dynamical fermions, Phys
M. Hasenbusch, Speeding up the hybrid Monte Carlo algorithm for dynamical fermions, Phys. Lett. B 519 (2001) 177–182.doi:10.1016/ S0370-2693(01)01102-9
2001
-
[17]
L ¨uscher, Solution of the Dirac equation in lattice QCD using a domain decomposition method, Comput
M. L ¨uscher, Solution of the Dirac equation in lattice QCD using a domain decomposition method, Comput. Phys. Commun. 156 (2004) 209–220.doi:10.1016/S0010-4655(03)00486-7
2004 doi
-
[18]
L ¨uscher, Local coherence and deflation of the low quark modes in lattice QCD, JHEP 07 (2007) 081.doi:10.1088/1126-6708/2007/ 07/081
M. L ¨uscher, Local coherence and deflation of the low quark modes in lattice QCD, JHEP 07 (2007) 081.doi:10.1088/1126-6708/2007/ 07/081
2007 doi
-
[19]
Babich, J
R. Babich, J. Brannick, R. C. Brower, M. A. Clark, T. A. Manteuffel, S. F. McCormick, J. C. Osborn, C. Rebbi, Adaptive multigrid algorithm for the lattice Wilson-Dirac operator, Phys. Rev. Lett. 105 (2010) 201602.doi:10.1103/PhysRevLett.105.201602
2010 doi
-
[20]
Boyda, et al., Applications of Machine Learning to Lattice Quantum Field Theory, in: Snowmass 2021, 2022
D. Boyda, et al., Applications of Machine Learning to Lattice Quantum Field Theory, in: Snowmass 2021, 2022
2021
-
[21]
Holland, A
K. Holland, A. Ipp, D. I. M ¨uller, U. Wenger, Machine-Learned Renormalization-Group-Improved Gauge Actions and Classically Perfect Gradient Flows, Phys. Rev. Lett. 136 (3) (2026) 031901.doi:10.1103/k41k-2pnc
2026 doi
-
[22]
Favoni, A
M. Favoni, A. Ipp, D. I. M ¨uller, D. Schuh, Lattice Gauge Equivariant Convolutional Neural Networks, Phys. Rev. Lett. 128 (3) (2022) 032003. doi:10.1103/PhysRevLett.128.032003
2022 doi
-
[23]
Catumba, A
G. Catumba, A. Ramos, Stochastic automatic differentiation and the signal to noise problem, Eur. Phys. J. C 85 (9) (2025) 1037.doi: 10.1140/epjc/s10052-025-14690-0
2025 doi
-
[24]
Abbott, D
R. Abbott, D. Boyda, Y . Fu, D. C. Hackett, G. Kanwar, F. Romero-L ´opez, P. E. Shanahan, J. M. Urban, Variance reduction in lattice QCD observables via normalizing flows (2026).arXiv:2603.02984
2026
-
[26]
Aarts, D
G. Aarts, D. E. Habibi, A. Ipp, D. I. M ¨uller, T. R. Ranner, L. Wang, W. Wang, Q. Zhu, Generalizable Equivariant Diffusion Models for Non-Abelian Lattice Gauge Theory (2026).arXiv:2601.19552
2026
-
[27]
Abbott, D
R. Abbott, D. Boyda, G. Kanwar, F. Romero-L ´opez, D. C. Hackett, P. E. Shanahan, J. M. Urban, Progress in Normalizing Flows for 4d Gauge Theories, PoS LATTICE2024 (2025) 066.doi:10.22323/1.466.0066
2025 doi
-
[28]
Caselle, E
M. Caselle, E. Cellini, A. Nada, M. Panero, Stochastic normalizing flows as non-equilibrium transformations, JHEP 07 (2022) 015.doi: 10.1007/JHEP07(2022)015
2022 doi
-
[29]
L. Wang, G. Aarts, K. Zhou, Diffusion models as stochastic quantization in lattice field theory, JHEP 05 (2024) 060.doi:10.1007/ JHEP05(2024)060
2024
-
[31]
K. K. Szabo, L. Lellouch, Z. Fodor, F. Stokes, B. C. Toth, G. Wang, New physics in the muon magnetic moment?, Procedia Comput. Sci. 240 (2024) 91–98.doi:10.1016/j.procs.2024.07.012
2024 doi
-
[32]
Alexandrou, et al., Large-scale simulations of lattice QCD for nucleon structure using Nf=2+1+1 flavors of twisted mass fermions, Procedia Comput
C. Alexandrou, et al., Large-scale simulations of lattice QCD for nucleon structure using Nf=2+1+1 flavors of twisted mass fermions, Procedia Comput. Sci. 267 (2025) 92–101.doi:10.1016/j.procs.2025.08.236
2025 doi
-
[33]
Bors ´anyi, Z
S. Bors ´anyi, Z. Fodor, J. N. Guenther, P. Parotto, A. P´asztor, L. Pirelli, K. K. Szab´o, C. H. Wong, The QCD crossover line in a finite volume, Procedia Comput. Sci. 267 (2025) 4–13.doi:10.1016/j.procs.2025.08.229
2025 doi
-
[34]
Finkenrath, Review on Algorithms for dynamical fermions, PoS LATTICE2022 (2023) 227.doi:10.22323/1.430.0227
J. Finkenrath, Review on Algorithms for dynamical fermions, PoS LATTICE2022 (2023) 227.doi:10.22323/1.430.0227
2023 doi
-
[35]
Finkenrath, Future trends in lattice QCD simulations, PoS EuroPLEx2023 (2024) 009.doi:10.22323/1.451.0009
J. Finkenrath, Future trends in lattice QCD simulations, PoS EuroPLEx2023 (2024) 009.doi:10.22323/1.451.0009
2024 doi
-
[36]
Alexandrou, S
C. Alexandrou, S. Bacchio, J. Finkenrath, A. Frommer, K. Kahl, M. Rottmann, Adaptive Aggregation-based Domain Decomposition Multi- grid for Twisted Mass Fermions, Phys. Rev. D 94 (11) (2016) 114509.doi:10.1103/PhysRevD.94.114509
2016 doi
-
[37]
Kostrzewa, S
B. Kostrzewa, S. Bacchio, J. Finkenrath, M. Garofalo, F. Pittler, S. Romiti, C. Urbach, Twisted mass ensemble generation on GPU machines, PoS LATTICE2022 (2023) 340.doi:10.22323/1.430.0340
2023 doi
-
[39]
G. I. Egri, Z. Fodor, C. Hoelbling, S. D. Katz, D. Nogradi, K. K. Szabo, Lattice QCD as a video game, Comput. Phys. Commun. 177 (2007) 631–639.doi:10.1016/j.cpc.2007.06.005
2007 doi
Reviewed July 31, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.