{"id":"35325c88-b358-4ecd-a0e6-3e611b6ceefc","arxiv_id":"2607.25061","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":2.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"Lattice QCD on European HPC needs bandwidth- and communication-balanced machines, sustained double precision, portable software, and dedicated human expertise more than peak FLOPs alone.","lead":"European lattice QCD groups map how gauge-field generation and measurements stress memory bandwidth, interconnects, and double precision on current and future HPC systems. The paper is a community requirements brief for EuroHPC procurement, software support, and research-software careers.","discovery_kind":"review","skeptic_critique":{"model":"moonshotai/kimi-k3","headline":"The \"future requirements\" claim assumes the workload stays Dirac-solver/DP-dominated; two trends the paper itself cites (ML-based sampling, emulated FP64) could erode the premise that native balanced/DP-heavy hardware is the binding need.","rationale":"The reader located the soft spot in the untested §3.4 policy premise (benchmarks/roadmap changing procurement). I agree that premise is unsupported, but it is a recommendation layered on top of the central claim, not a load-bearing assumption of it — the central claim would stand even if the policy mechanism failed. The more load-bearing assumption sits one level deeper: the future-requirements profile presumes the workload remains DP-dominated Dirac-solver work, while the paper's own §2.1–2.2 describe two active developments (ML sampling, FP64 emulation) that could invalidate exactly that premise on the 5–10 year horizon the paper addresses. This is a durability risk, not a correctness defect: the descriptive profile of current workloads is accurate and well-supported, and for a proceedings requirements document the hedged language (\"could be suboptimal,\" \"may be feasible, if performant and sufficiently verified\") is honest. Hence the reader's ACCEPT with low correctness risk stands; the concern warrants a quantified follow-up (the emulation benchmark above) rather than a verdict change. Confidence in this assessment is moderate-high: the technical facts about arithmetic intensity and mixed-precision practice are standard, and the uncertainty is about future workload mix, which is genuinely open.","tokens_in":23742,"tokens_out":1677,"duration_ms":55695,"concrete_test":"Benchmark a production multigrid Dirac solver (e.g., QUDA or tmLQCD/QUDA offload as in Fig. 2a) using emulated FP64 via tensor-core Ozaki-style schemes versus native FP64 on the same GPU, for a physical-mass ensemble solve; measure sustained time-to-solution and verify a large-time-separation correlator is statistically unchanged. If emulated DP reaches ≥80% of native DP sustained performance with identical physics output, the \"native DP hardware essential\" requirement in §3.1 weakens and the procurement argument needs revision.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The descriptive core of the paper is solid: the Dirac operator is a low-arithmetic-intensity stencil, HMC plus iterative solvers dominate cost, multigrid/deflation shift the bottleneck further toward bandwidth and latency — all consistent with decades of lattice practice and the cited algorithmic literature. The load-bearing part of the *central* claim, however, is forward-looking: §2.2 and §3.1 argue that \"high and stable double precision performance remains essential\" and that AI-oriented low-precision GPU trends \"could be suboptimal for lattice workloads.\" This holds only if (a) production workloads remain dominated by DP-critical Dirac solves and long MD trajectories, and (b) no performant substitute for native FP64 emerges. The paper itself supplies the countervailing evidence: §2.1 notes generative models (flows, diffusion) aimed at replacing HMC sampling, and §2.2/§3.1 concede emulated high-precision arithmetic \"may be feasible, if performant and sufficiently verified.\" If either matures — 4D production-scale ML sampling, or Ozaki-style FP64 emulation on tensor cores matching native DP throughput — the procurement argument for DP-balanced architectures weakens substantially, since the very AI-oriented hardware trends the paper cautions against would become adequate. The paper hedges verbally but never quantifies how much of the projected workload mix is robustly DP-bound, so the durability of its headline requirement is asserted rather than demonstrated. This is a soft spot in the normative thrust, not an error in the descriptive analysis.","agreement_with_reader":"partial"},"referee_report":{"model":"moonshotai/kimi-k3","summary":"This is a white-paper-style contribution to the EuroHPC User Days 2026 proceedings, authored by representatives of the European lattice field theory community (EuroLFT). It describes the computational profile of lattice QCD — HMC-based gauge-field generation dominated by sparse Dirac-operator inversions, measurement pipelines limited by signal-to-noise degradation, and emerging machine-learning-assisted workflows — and argues that these workloads are bandwidth- and latency-bound rather than FLOP-bound. From this it derives hardware requirements (high memory bandwidth, low-latency high-bandwidth interconnects, large on-device memory, sustained double-precision performance, mixed-precision support, topology-aware networks), software needs (performance portability, continuous integration at computing centres), human-resource needs (dedicated performance-engineering support, PASC-like European structures), and policy recommendations (community benchmarks in procurement, a European post-exascale roadmap). The descriptive material is standard and well referenced (HMC, critical slowing down, multigrid, FLAG, ILDG), and the resource figures (Fig. 1; Fig. 2a–b) are illustrative rather than derived results.","tokens_in":23951,"tokens_out":2690,"duration_ms":28835,"significance":"If taken up by EuroHPC stakeholders, the paper provides a credible, field-backed statement of what bandwidth-bound scientific workloads need from post-exascale European systems, and it does so with concrete evidence rather than assertion: per-workload resource-investment data from the OpenLAT collaboration (Fig. 1), documented order-of-magnitude speedups from multigrid solvers and GPU offloading in tmLQCD (Fig. 2a), and HiRep scaling measurements to 1024 GPUs on LUMI-G and JUWELS Booster. The articulation of why peak-FLOP and AI-oriented low-precision procurement trends are a poor match for Dirac-solver-dominated workloads is the most useful contribution for a policy-facing venue. The requirements are falsifiable in principle (benchmarkable against delivered systems) and are consistent with the cited algorithmic literature.","major_comments":[{"comment":"Figure 2b (§2.3) appears to embed multiple full pages of ref [38] (Drach et al., Comput. Phys. Commun. 322 (2026) 110061) — including its Table 6, Fig. 31, §§9.3–10 and Eqs. (138)–(144), with the same block repeated several times — rather than a clean strong/weak-scaling plot for HiRep. If this is not an artifact of the arXiv PDF text layer, the figure as submitted is illegible, reproduces substantial material from another publication, and cannot serve its stated evidentiary role (demonstrating 'excellent strong-scaling performance'). Please replace it with a self-contained plot of the HiRep scaling data with proper attribution, and verify the production version carefully.","section":"Fig. 2b and §2.3"},{"comment":"The hardware requirement that 'high and stable double precision performance remains essential' (§3.1, fourth bullet; echoed in §2.2 and the Abstract) is the most forward-looking and procurement-relevant claim in the paper, but its durability is asserted rather than argued. The manuscript itself cites the two trends that could erode it: generative models for sampling (§2.1, refs [25]–[30]) and emulated high-precision arithmetic (§2.2: 'may be feasible, if performant and sufficiently verified'). If either matures, AI-oriented low-precision-heavy hardware would become substantially more adequate for lattice workloads. For a requirements document this does not invalidate the claim — current production workloads are unambiguously DP-critical — but the paper should quantify or at least delineate the robust core: e.g., what fraction of the projected ensemble-generation and measurement cycle (lo","section":"§2.2 / §3.1 (double-precision requirement)"},{"comment":"The two main policy proposals — incorporating community-specific benchmarks into EuroHPC procurement and publishing a 5/10-year post-exascale roadmap — are stated as self-evidently beneficial, with no evidence or precedent that benchmark inclusion changes awarded system balance, or that roadmap visibility (rather than funding level or vendor roadmaps) is the binding constraint on long-term code investment. The PRACE scientific case (ref [40]) is cited but not used to support the causal claim. Either cite concrete precedents where representative application benchmarks demonstrably shaped procurement outcomes (e.g., US DOE/NNSA benchmark-driven procurements, or the role of lattice codes in BlueGene co-design, which the Conclusions themselves invoke), or temper the language from prescription to recommendation.","section":"§3.4 (policy recommendations)"}],"minor_comments":[{"comment":"The ordinate label '[Mch]CPU or [Knh]GPU' is ambiguous: CPU million-core-hours and GPU thousand-node-hours are mixed in one plot, and the legend includes a single 'US' series with no definition or source. Please state units per series, the conversion (if any) used to make CPU and GPU investment comparable, and what the US bar represents.","section":"Fig. 1 caption and axis labels"},{"comment":"The axis legend reads 'JuQueen'; the standard spelling is 'JUQUEEN'. Similarly check 'Juwels-Booster' (JURECA/JUWELS Booster).","section":"Fig. 2a"},{"comment":"Missing space ('Keywords:Lattice'); 'high- performance' broken across lines in the Abstract; inconsistent capitalisation of collaboration names ('OpenLAT' vs 'openLAT' in §2.3).","section":"Keywords and typos"},{"comment":"The collaboration list includes 'TWEXT'; please verify this is the intended name, as it may be unfamiliar to some readers (a reference or expansion would help).","section":"§2.3, collaboration list"},{"comment":"The statement that g−2 'showed a tension between experimental measurements and theoretical predictions' should be updated or nuanced in light of ref [1] (the 2025 Theory Initiative update) and the most recent experimental/lattice situation, since the paper cites precisely that update.","section":"§1, muon g−2 motivation"},{"comment":"Ref [6] is listed as Phys. Rev. D 113 (2026) — please check the year/volume; ref [38] uses an https DOI in the doi field; ref [40] lacks an author/organisation field (PRACE).","section":"References"},{"comment":"The claim 'typically leads to lattice spacings a~0.04–0.1 fm' and 'order 10^10 degrees of freedom' would benefit from one clarifying sentence on how the 12 spin-colour-per-site factor enters, for the non-lattice HPC audience this proceedings addresses.","section":"§2.1, lattice sizes and degrees of freedom"}],"recommendation":"minor_revision","confidential_remarks":"This is a community position paper authored by representatives of the EuroLFT initiative advocating for the resource needs of their own field; the advocacy framing is transparent and appropriate for a EuroHPC User Days proceedings. Citation pattern is normal for the genre, with moderate self-citation to the authors' software efforts (tmLQCD, HiRep), which are also the natural examples. The apparent wholesale embedding of pages from ref [38] (a paper by a subset of the authors) into Fig. 2b should be checked against the publisher's version — if real, it raises a reproduction/copyright question the editor may wish to verify before production. Fit with the venue is good; I do not see grounds for rejection or major revision."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"This is a User Days requirements note, not a physics result. Read it as infrastructure strategy from people who actually run the codes.\n\nWhat it does well is the descriptive core. The workflow sketch—HMC, Dirac solves, multigrid/deflation shifting cost toward bandwidth and latency, signal-to-noise driving measurement volume—is standard lattice practice and matches the cited literature. Fig. 1 (OpenLAT spend) and Fig. 2 (tmLQCD cost drop; HiRep strong/weak scaling on LUMI-G/JUWELS) are concrete illustrations, not hand-waving. The hardware ask is coherent: high memory bandwidth, low-latency interconnects, large on-device memory, sustained DP, mixed precision where it belongs, topology-aware networks. Software portability, CI at centres, and human expertise (PASC-style co-design) are the right second-order points. Circularity is low; they are not fitting a target to itself.\n\nNovelty is low by design. QUDA, tmLQCD, HiRep, critical slowing down, FLAG/ILDG context—community knowledge restated for EuroHPC. That is fine for the venue.\n\nSoft spots, in proportion. The normative thrust in §3.4 (community benchmarks and a clearer post-exascale roadmap will improve procurement) is asserted without evidence that those are the binding constraints versus funding or vendor roadmaps. More load-bearing: the headline “balanced DP-heavy hardware remains essential” assumes production stays Dirac/DP-dominated. The paper itself flags generative sampling and emulated high precision as possible outs, but never quantifies how much of the projected mix is robustly native-DP-bound. If 4D ML sampling or performant FP64 emulation lands, the caution against AI-oriented low-precision hardware weakens. That is a durability gap in the policy claim, not a flaw in the present-day bottleneck analysis.\n\nWho it is for: facility people, EuroHPC programme staff, and lattice groups writing resource cases. Not a reading-group paper for new physics. I would send it to peer review for this proceedings class—clear, usable, no technical red flags—and engage if you care about European system balance. I would cite the bottleneck framing in an infrastructure note; I would not cite it for algorithms or results.","headline":"Solid EuroHPC requirements brief: accurate on today’s Dirac/bandwidth profile, thin on how durable that profile stays if ML sampling or FP64 emulation mature.","tokens_in":24694,"tokens_out":577,"would_cite":true,"duration_ms":16357,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Lattice QCD will keep advancing on European supercomputers only if machines prioritise memory bandwidth, fast interconnects and stable double precision—not peak FLOPs alone.","keywords":["Lattice field theory","QCD","Computational cost","Efficiency","HPC","Exascale","Dirac solvers","Memory bandwidth"],"falsifier":"If EuroHPC systems chosen after including lattice-style bandwidth and communication benchmarks still deliver the same strong-scaling efficiency and double-precision sustained fraction for Dirac solvers as current AI-oriented machines, or if multi-year lattice software investment does not rise once a public roadmap exists, the policy half of the claim fails.","tokens_in":24504,"feed_emoji":"⚛️","tokens_out":824,"duration_ms":30379,"temperature":0.7,"pith_summary":"This paper maps what lattice field theory actually needs from European high-performance computing. Lattice QCD workloads are dominated by repeated sparse solves of the Dirac operator during gauge-field generation and measurements; those kernels are limited by memory bandwidth and communication, not by peak arithmetic speed. The authors argue that future systems must therefore supply high bandwidth, low-latency interconnects, large on-device memory and reliable double-precision performance, together with portable software and specialist human support at computing centres. They warn that procurement and allocation trends that favour AI-oriented low-precision throughput risk starving these communication-bound codes. Sustained progress, they conclude, also needs community benchmarks in machine selection, a clearer post-exascale roadmap, and career paths that keep research software expertise inside the field.","feed_headline":"Lattice QCD needs bandwidth, not peak FLOPs","feed_subtitle":"European supercomputers must keep double precision, fast networks and specialist staff or progress stalls","key_machinery":"Repeated application of the lattice Dirac operator inside iterative sparse linear solvers (in Hybrid Monte Carlo configuration generation and in large-scale measurements). These bandwidth-bound, nearest-neighbour stencil kernels—not peak FLOPs—set the hardware, software and scaling requirements the paper derives.","core_discovery":"Sustained progress of lattice field theory on European HPC requires balanced architectures—high memory bandwidth, low-latency high-bandwidth interconnects, large on-device memory, and stable double-precision performance—plus portable software ecosystems and dedicated human expertise; peak floating-point performance and AI-driven low-precision hardware alone are a poor match for Dirac-solver-dominated workloads.","pith_inferences":["Other stencil-heavy, communication-bound fields named in the paper—hydrodynamics and numerical gravity—would gain from the same balanced procurement criteria.","If accelerator roadmaps keep cutting native double precision, verified high-precision emulation becomes a shared infrastructure problem, not a lattice-only patch.","Without European-scale co-design support, lattice groups risk fragmenting effort across many small, architecture-specific codes as platforms diversify."],"forward_implications":["Procurement suites that include lattice-like bandwidth, latency and strong-scaling tests would favour more balanced memory-to-compute ratios.","Lattice codes should sit in HPC centres’ continuous testing pipelines so toolchain and hardware regressions are caught early.","Stable posts for research software engineers become a prerequisite for using next-generation machines effectively.","Mixed precision can accelerate inner solver iterations, but double (and sometimes extended) precision remains mandatory for observables and long trajectories.","Shared gauge-field ensembles continue to amortise the high fixed cost of configuration generation across many physics analyses."],"fun_headline_variants":["Lattice QCD needs bandwidth over peak FLOPs","European HPC: bandwidth and staff over raw FLOPs","Lattice field theory stalls without balanced HPC","Double precision and fast nets key for lattice QCD","Peak FLOPs alone fail lattice QCD workloads"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"That putting community-specific benchmarks into procurement and publishing a clearer European post-exascale roadmap will actually change which machines get built and how groups invest in long-lived code.","fun_headline_variants_meta":{"raw":{"variants":["Lattice QCD needs bandwidth over peak FLOPs","European HPC: bandwidth and staff over raw FLOPs","Lattice field theory stalls without balanced HPC","Double precision and fast nets key for lattice QCD","Peak FLOPs alone fail lattice QCD workloads"]},"model":"grok-4.5","effort":"low","cost_usd":0.002642,"raw_usage":{"total_tokens":886,"prompt_tokens":614,"num_sources_used":0,"completion_tokens":51,"cost_in_usd_ticks":26424000,"prompt_tokens_details":{"text_tokens":614,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":221,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":614,"tokens_out":51,"duration_ms":4166,"temperature":1.0,"reasoning_tokens":221,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-31T02:18:06.945807+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"If EuroHPC systems chosen after including lattice-style bandwidth and communication benchmarks still deliver the same strong-scaling efficiency and double-precision sustained fraction for Dirac solvers as current AI-oriented machines, or if multi-year lattice software investment does not rise once a public roadmap exists, the policy half of the claim fails.","supporting_citations":[],"review_version":1}