{"id":"4ecb8d62-3685-4a31-ae4a-2a4bfeb7afae","arxiv_id":"2602.21617","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Using Tr M^-1 as both an input and a feature, a bias-corrected ML model predicts Tr M^-2..-4 and reproduces chiral-condensate cumulants with ~1% labeled data at ~26% of the original cost.","lead":"Lattice QCD measurements of chiral-condensate fluctuations are computationally expensive; this paper shows machine learning can predict most of the costly Dirac-operator trace terms from a cheap one, cutting the measurement budget to about a quarter. The preserved higher-order cumulants are exactly the observables used to hunt for the QCD critical endpoint.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Claimed consistency for susceptibility and skewness at ~1% labeled data is not shown; only kurtosis at κ_t is presented, so the abstract overstates Fin coverage.","rationale":"The paper's strongest claim names three observables, but the evidence section only supports kurtosis. The reader's weakest assumption points to TrM^-1 dominance, which is related but different: even if dominance holds, the manuscript still needs to show χ and S agree. I do not see an internal inconsistency in the Fin construction or the cost arithmetic as stated; the issue is that a load-bearing part of the abstract is unsubstantiated by the presented data. A concrete check is feasible because the pipeline and ensembles are described and code is linked: reproduce the omitted C_B maps. If they pass, the concern is resolved; if they fail or cannot be produced, the central claim should be conditional on kurtosis only, which is why the CONDITIONAL verdict remains appropriate rather than ACCEPT. I partially agree with the reader: the dominance assumption is also untested, but the more immediate gap is missing evidence for χ and S.","tokens_in":13480,"tokens_out":10065,"duration_ms":97980,"concrete_test":"Using the released Deborah.jl code and the same ensembles, run the Fin pipeline at R_LB=1% (and at least the R_TR values shown in Fig. 2) and compute C_B and (x,r) for χ and S against the fully measured baseline. If every C_B ≥0.95, the abstract's claim is supported; if these maps are absent or any C_B falls below 0.95 at the cited operating point, the claim should be narrowed to kurtosis and the cost-reduction statement revised accordingly.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim in the abstract is that with ~1% labeled data, susceptibility, skewness, and kurtosis all remain statistically consistent with full baselines at ~26% cost. The only quantitative evidence presented for the Fin setup (Fig. 2b, Fig. 3, and the surrounding discussion) is C_B for the kurtosis at the transition point K(κ_t). There is no displayed C_B map, x/r diagnostic, or table for χ=C_2/V or S=C_3/C_2^{3/2}, for Fin or Fex. The §4 sentence 'the cumulants are dominated by this observable' is an asserted explanation, not a tested property, and it is used to justify the cost claim. Even if TrM^-1 dominance is true, the abstract's three-observable claim requires χ and S to be consistent; without those results the 26% cost reduction is demonstrated only for kurtosis. This is a support gap in the strongest claim, not merely a missing generalization to other ensembles.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a bias-corrected machine-learning scheme to estimate the higher-power traces Tr M^{-n} (n=2,3,4) of the Wilson-clover Dirac operator, with two feature scenarios: Fin, which uses the exactly measured Tr M^{-1} both as a physical input to the cumulant construction and as an ML feature, and Fex, which uses only gauge observables (plaquette, rectangle). The traces are combined into cumulants of the chiral condensate and analyzed through multi-ensemble reweighting across five κ values. The main quantitative result is the Bhattacharyya-coefficient map for the kurtosis at the transition point (Fig. 2): Fin gives C_B ≈ 1 across almost the entire (R_LB, R_TR) scan, while Fex requires a sufficiently large labeled fraction and explicit bias correction to avoid multi-standard-deviation deviations. The abstract claims that susceptibility, skewness, and kurtosis are all statistically consistent with fully measured baselines at ~1% labeled data and ~26% cost.","tokens_in":13775,"tokens_out":6520,"duration_ms":56866,"significance":"If the claimed consistency holds for all three cumulants, the Fin scheme would reduce the measurement cost for higher-order cumulants from 400 CG solves per configuration to ~103 solves (25.75%), without statistically changing the results on this dataset. The bias-correction framework is well motivated by AMA and the paper provides reproducible code (Deborah.jl), which strengthens the work. The Fex result, showing that removing bias correction leads to large distortions once the ML outputs are fed into multi-stage reweighting, is a useful cautionary finding for the lattice-ML community.","major_comments":[{"comment":"The abstract and Sec. 4 claim that the Fin setup yields statistically consistent susceptibility, skewness, and kurtosis at ~1% labeled data. However, the only quantitative evidence displayed is the C_B map for the kurtosis K(κ_t) in Fig. 2(b) and the x/r diagnostics in Fig. 3. No C_B, x, or r results are shown for χ = C_2/V or S = C_3/C_2^{3/2} under either Fin or Fex. This is a support gap in the central claim: the 26% cost reduction and the 'statistically consistent' statement are made without qualification. The authors should either add the corresponding C_B maps/tables for χ and S, or revise the abstract and conclusion to state that consistency is demonstrated specifically for the kurtosis on this dataset.","section":"Abstract, Sec. 4, Fig. 2"},{"comment":"The cost figure 25.75% is arithmetic, but the fidelity claim relies on the assertion that 'the cumulants are dominated by this observable' (Tr M^{-1}). The paper does not quantify the relative contributions of the different trace terms to C_2, C_3, and C_4. In particular, the kurtosis expression for Q_4 (Eq. 2) contains the term -6 N_f Tr M^{-4}, and Fig. 1 shows that Tr M^{-4} has a substantially weaker correlation with Tr M^{-1} (0.78 and 0.41 for the two representative ensembles). The empirical C_B ≈ 1 in Fig. 2(b) suggests the approximation works for this dataset, but the conclusion presents the dominance as a general property. Please provide a quantitative decomposition of the cumulants' mean/variance by trace term, or explicitly state that the dominance is an assumption validated only for the kappa values, volume, and action used here.","section":"Sec. 4, Eq. (2), Fig. 1"},{"comment":"The evaluation metric C_B is based on a Gaussian approximation of the sampling distributions of the final reweighted observables. For higher-order cumulants such as skewness and kurtosis near a first-order transition, the distributions may be significantly non-Gaussian, and the Bhattacharyya coefficient computed from only the mean and variance may not capture the agreement faithfully. The paper does not test this Gaussian assumption for the observables in Figs. 2-3. This is load-bearing because C_B is the primary figure of merit throughout. Please either justify the Gaussianity (e.g., with quantile-quantile plots or higher-moment checks) or use a metric that does not rely on it.","section":"Sec. 2.7, Eq. (7)"}],"minor_comments":[{"comment":"The manuscript contains numerous typos, most notably '/github' appearing in the abstract and Sec. 1 instead of a full URL or DOI. The code link should be given in standard form.","section":"General"},{"comment":"The table formatting is garbled; the row identifiers (e.g., 'L12T4b1.60k13575') are merged with the column headers, making the ensemble parameters hard to read.","section":"Table 2"},{"comment":"The block-bootstrap procedure is described only verbally. Specify the block length selection, the number of bootstrap resamples, and whether the block length is scale-dependent.","section":"Sec. 2.6"},{"comment":"The caption says the left half of the panel shows x and the right half shows r, but the figure appears to contain two separate heatmaps with different color scales. Clarify the layout and label the color axes directly in the figure.","section":"Fig. 3 caption"},{"comment":"The statement that the problematic region 'largely disappears once R_LB ≳ 20%' is not fully consistent with the heatmap: several cells with R_LB < 10% show C_B ≥ 0.9, and some cells with R_LB between 15% and 20% at R_TR = 90% show C_B ≈ 0.85–0.94. Define the threshold used for 'problematic' or discuss the non-monotonicity.","section":"Sec. 3, Fig. 2(a)"}],"recommendation":"major_revision","confidential_remarks":"The paper is technically sound in its core Fin result for the kurtosis, but the abstract overstates the coverage to all three cumulants. Adding the missing C_B maps for χ and S, or narrowing the claims, should be straightforward given the data. The trace-dominance assumption in Sec. 4 also needs either quantitative support or an explicit caveat. The work is likely acceptable after these revisions; I do not see a basis for rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a useful methodological paper. The genuinely new piece is the systematic (R_LB, R_TR) scan of the bias-corrected ML estimator from Ref. [9] applied to multi-ensemble reweighting for Nf=4 Wilson-clover ensembles, plus the Fin/Fex comparison. The Fin result is convincing for the kurtosis at kappa_t: C_B is about 1 across nearly the whole scan, and the R_TR=100% column showing catastrophic failure without bias correction is a nice sanity check. The cost arithmetic is sound: keeping exact Tr M^-1 and predicting the three higher traces brings the CG budget to roughly 25.75% of full cost.\n\nThe soft spots are real but not fatal. The abstract overstates coverage: it says susceptibility, skewness, and kurtosis all remain statistically consistent at ~1% labeled data, but the displayed evidence is only kurtosis at the transition point. There are no C_B maps or x/r diagnostics for chi or S. That doesn't kill the central claim, but it needs either figures or a clear statement that the abstract is about expected rather than demonstrated consistency. Relatedly, the assertion that the cumulants are dominated by Tr M^-1 is reasonable but not stress-tested; Fig. 1 shows Tr M^-4 is weakly correlated, so on ensembles where higher traces matter more the Fin scheme could fail. The conclusion does acknowledge this caveat, but the abstract does not.\n\nReproducibility is a mixed bag: the code is on GitHub/Zenodo, which is good, but the ML architecture and hyperparameters are not specified in the text, and the gauge ensembles are not public. For a Lattice proceedings that is acceptable, though a would-be user needs more detail.\n\nThe Fex approach is more ambitious and honestly reported—low labeled fractions and no bias correction give bad results, and the Newton solver convergence issues are marked in the figures. That transparency is to the authors' credit.\n\nWho this is for: lattice QCD practitioners thinking about reducing stochastic measurement cost near the QCD critical endpoint. It deserves a serious referee. Send it to review, but ask the authors to either add the chi and S evidence or soften the abstract to match what is actually shown.","headline":"Fin scheme looks like a real cost saver for kurtosis at the transition point, but the abstract's claims about susceptibility and skewness are not supported by the shown plots.","tokens_in":14214,"tokens_out":2299,"would_cite":true,"duration_ms":60764,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["11.15.Ha","12.38.Gc"],"model":"deepseek-v4-flash","headline":"Machine learning can estimate the higher-order cumulants of the chiral condensate at roughly a quarter of the current measurement cost, using only about 1% of configurations as labeled data, provided the leading trace Tr M^{-1} is kept exac","keywords":["lattice QCD","chiral condensate","higher-order cumulants","machine learning","bias correction","multi-ensemble reweighting","stochastic trace estimation","critical endpoint"],"falsifier":"Run the Fin pipeline on an ensemble or at a parameter point where the Tr M^{-4} contribution to the kurtosis is sizable (for instance, closer to the critical endpoint or at a smaller volume), and compare the ML-reweighted kurtosis to the fully measured baseline; if the difference exceeds the statistical error, the dominance assumption fails.","tokens_in":13419,"feed_emoji":"⚛️","tokens_out":6137,"duration_ms":54342,"temperature":0.7,"pith_summary":"The paper tries to show that a bias-corrected machine-learning estimator can replace most of the expensive inversions of the lattice Dirac operator needed for the susceptibility, skewness, and kurtosis of the chiral condensate. In the main 'Fin' setup, the measured Tr M^{-1} is kept exact and used as a feature to predict Tr M^{-2}, Tr M^{-3}, and Tr M^{-4}; the higher traces are then combined with the exact trace to build cumulants. With as little as 1% of configurations labeled, the reweighted cumulants remain statistically consistent with fully measured baselines, cutting computational cost to about 26%. A second 'Fex' setup, which predicts all traces from gauge observables, works only if enough labeled data (about 20%) is available and if explicit bias correction is applied; without bias correction, the inferred kurtosis at the transition deviates by many standard deviations.","feed_headline":"Machine learning cuts quark-condensate measurement cost to 26%","feed_subtitle":"Even with ~1% of configurations labeled, susceptibility, skewness, and kurtosis stay consistent with full measurements.","key_machinery":"The load-bearing mechanism is the bias-corrected estimator P1: predictions from a supervised regression model on the unlabeled set are added to the mean of (exact value minus prediction) over a separate bias-correction set, which cancels systematic model bias without exact inversions on every configuration. This estimator is wrapped in a multi-ensemble Ferrenberg-Swendsen reweighting procedure that combines five ensembles at different quark masses (kappa values) to interpolate observables and the transition point. The cumulants of the chiral condensate are built from the quark-loop operators Q1-Q4, which are polynomial combinations of Tr M^{-n}; the paper defines the Fin setup, which keeps T","core_discovery":"The authors report that a bias-corrected supervised regression estimator, applied to stochastic estimates of Tr M^{-n} on Wilson-clover ensembles with the Iwasaki gauge action, reproduces the conventional full-measurement results for the cumulants of the chiral condensate. In the Fin configuration, where the full set of measured Tr M^{-1} values is retained and only higher powers are predicted, the multi-ensemble reweighting yields Bhattacharyya coefficients near 1 across the scanned labeled/training fractions, including the extreme case of 1% labeled data. This corresponds to reducing the dominant matrix-inversion cost to roughly 26% of the original measurement budget. In the Fex configurat","pith_inferences":["If the dominance of Tr M^{-1} persists on larger volumes and physical quark masses, the 26% figure could become a practical default for fluctuation measurements, but the current evidence is restricted to the ensembles studied.","The weak correlation of Tr M^{-4} with Tr M^{-1} in the lightest-quark ensemble suggests that the Fin approach would break down for observables where the fourth trace term is not subleading; a direct test would be to compute baryon-number cumulants or isospin fluctuations with the same pipeline.","The Fex result suggests a cheap screening strategy for identifying critical-endpoint candidates, but the large sensitivity to bias correction means one could test whether adding the Polyakov loop as a feature changes the convergence pattern and the poor low-label behavior."],"forward_implications":["If the Fin result holds, lattice QCD can compute the susceptibility, skewness, and kurtosis of the chiral condensate at roughly a quarter of the current cost without changing the statistical conclusions.","The 26% cost bound follows directly from the claim that Tr M^{-1} dominates the cumulants, so the exact measurement of the lowest trace cannot be avoided.","The Fex results imply that a fully feature-based ML pipeline for these observables would require roughly 20% labeled data and an explicit bias-correction stage to be reliable.","Bias correction is not optional when ML outputs feed into multi-ensemble reweighting: the no-bias-correction column shows multi-standard-deviation deviations in the kurtosis at the transition point.","The same bias-corrected estimator could be applied to other fermionic observables that enter higher-order fluctuation analyses near the critical endpoint."],"fun_headline_variants":["ML cuts quark-condensate cumulant cost to 26%","Bias-corrected ML keeps cumulants stable at 1% labels","Machine learning slashes QCD measurement budget to 26%","Reweighting with ML: 26% cost, consistent cumulants","Bias-corrected regression: 26% cost for chiral cumulants"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The cost and fidelity claims rest on the assumption that the cumulants are dominated by the exactly measured Tr M^{-1}, so errors in the machine-learned higher traces do not shift the result.","fun_headline_variants_meta":{"raw":{"variants":["ML cuts quark-condensate cumulant cost to 26%","Bias-corrected ML keeps cumulants stable at 1% labels","Machine learning slashes QCD measurement budget to 26%","Reweighting with ML: 26% cost, consistent cumulants","Bias-corrected regression: 26% cost for chiral cumulants"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000289,"raw_usage":{"total_tokens":1560,"prompt_tokens":807,"completion_tokens":753,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":551,"completion_tokens_details":{"reasoning_tokens":657}},"tokens_in":551,"tokens_out":753,"duration_ms":5939,"temperature":1.0,"reasoning_tokens":657,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T20:58:43.035773+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the Fin pipeline on an ensemble or at a parameter point where the Tr M^{-4} contribution to the kurtosis is sizable (for instance, closer to the critical endpoint or at a smaller volume), and compare the ML-reweighted kurtosis to the fully measured baseline; if the difference exceeds the statistical error, the dominance assumption fails.","supporting_citations":[],"review_version":1}