{"id":"7ea1cd4f-585a-48cf-8375-b09f6be285f0","arxiv_id":"2411.18170","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Gradient boosting with bias correction estimates Tr M^{-n} in lattice QCD from cheaper observables in favorable ensembles, but needs a substantial labeled fraction and fails for light-quark Tr M^{-3,4}.","lead":"This paper tests whether a gradient boosting machine learning model can estimate traces of inverse Dirac operators in lattice QCD from cheaper observables such as plaquettes and Polyakov loops, instead of computing each trace with a costly linear solver. The method reproduces the exact results within one sigma in several ensembles when enough labeled configurations are provided, with the best performance for heavier quarks and at a first-order phase transition.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The good regions are defined only for P2, which directly includes the exact CG average over the labeled set; without reporting the pure ML estimator P1, the central claim that GBDT estimates Tr M^{-n} is not tested.","rationale":"The reader's conditional verdict identifies the transfer of the bias correction from S_BC to S_UL as the weakest assumption. That is a legitimate sub-issue, but the more fundamental omission is that the pure ML estimator P1 is never reported or evaluated. Eq. (2) defines P2 as a weighted average of P1 and the exact labeled CG average, so at R_LB of 30-50 percent a large, model-independent component is automatically present. This does not by itself make P2 invalid as a cost-reduction estimator, but it means the paper's central claim that the GBDT model estimates Tr M^{-n} is not directly tested by Tables 3-5. The proposed check is inexpensive because the true Y values on the unlabeled set are available from the original CG production; one simply evaluates P1 against those values. If P1 passes, the reader's concern about bias transfer is resolved empirically and the conditional verdict can be upgraded. If P1 fails while P2 passes, the paper should be reframed as validating a hybrid ML-plus-CG estimator rather than ML estimation alone. The other issues noted by the reader, such as the post hoc R_sigma threshold and the bootstrap not obviously including training-set variance, are real but secondary relative to the absence of the P1 comparison. The verdict remains conditional because the recommended check, not rejection, is the appropriate response to the identified gap.","tokens_in":7609,"tokens_out":12510,"duration_ms":127339,"concrete_test":"Compute the P1 estimator from Eq. (1) for every entry in Tables 3-5 using the same splits, and compare Ybar_P1 to the true CG average over the unlabeled set with the same EC-1 and EC-2 criteria. If P1 fails in a region where P2 was reported good, the machine-learning component is not validated. Also report the R_TR=0 no-ML control on the same R_LB grid; if that control already reaches R_sigma <= 1.2 at the claimed boundaries, the improvement attributed to the GBDT model should be quantified against it.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing concern is that the paper's success criteria are applied only to the P2 estimator of Eq. (2), not to the pure machine-learning estimator P1 of Eq. (1). P2 is a weighted average of Ybar_P1 and the exact CG average over the labeled subset, Ybar^(LB), with weight N_LB/N. At the reported good regions, R_LB is 30-50 percent, so 30-50 percent of the final P2 value is taken directly from the original CG data and is independent of the GBDT model quality. The text says that R_TR=0 (no ML, just the labeled average) was monitored, but no control comparison is shown and no baseline is used to subtract this direct contribution. If Ybar_P1 is biased because the S_BC residual does not represent S_UL, or if it is noisy, P2 can still satisfy the EC-1 score-2 and R_sigma <= 1.2 criteria because the exact labeled component pulls the estimate toward the full CG result. Since results for Ybar_P1 are not reported anywhere, Tables 3-5 do not establish that the GBDT model itself reproduces Tr M^{-n}; they establish only that a hybrid estimator containing a substantial exact-CG component agrees with the full CG estimate. The title and the summary claim about machine-learning estimation therefore outrun the evidence.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript reports a preliminary machine-learning study of the trace of inverse Dirac operator, Tr M^{-n}, in 4-flavor Wilson-clover QCD on three 16^3 x 4 ensembles with different hopping parameters. Using LightGBM gradient-boosted decision trees, the authors train a model f(X) to map input observables X (Plaquette, Polyakov loop, or Tr M^{-m} with m<n) to Y = Tr M^{-n}. They construct a bias-corrected ML estimator \\bar Y_P1 (Eq. 1) and a hybrid estimator \\bar Y_P2 (Eq. 2) that averages \\bar Y_P1 with the exact CG estimate on the labeled subset. The paper defines two evaluation criteria: EC-1 (1-sigma agreement of means) and EC-2 (statistical error ratio R_sigma <= 1.2 relative to the original CG result). For each (X,Y) pair and each dataset, they report regions in the labeled-fraction R_LB and training-fraction R_TR plane where P2 achieves EC-1 score 2 and EC-2. The reported good regions are summarized in Tables 3-5, with the main qualitative conclusions that heavier quark mass and the first-order phase-transition ensemble (ID-1) give larger good regions, while Tr M^{-4} is generally harder.","tokens_in":7878,"tokens_out":4973,"duration_ms":45344,"significance":"If the claimed results held, the paper would demonstrate a useful cost-reduction strategy for stochastic trace estimation in lattice QCD, building on the bias-correction methodology of Yoon et al. The authors are honest about the preliminary nature of the work, explicitly list to-do items, fix the LightGBM hyperparameters, and provide correlation matrices that motivate the choice of inputs. However, the central claim is currently supported only for the hybrid estimator \\bar Y_P2, not for the pure machine-learning estimator \\bar Y_P1. Since \\bar Y_P2 contains an exact-CG component with weight N_LB/N (30-50% in the reported good regions), the agreement with the full CG result is partly built in by construction. The paper also does not account for the training-set split variance in the bootstrap errors and uses a post-hoc EC-2 threshold. These issues make the significance of the quantitative results uncertain, although the question is worthy of investigation and the manuscript provides a reasonable framework for a subsequent, more complete analysis.","major_comments":[{"comment":"The central claim that gradient boosting estimates Tr M^{-n} is evaluated only through the hybrid estimator \\bar Y_P2, not through the machine-learning estimator \\bar Y_P1 of Eq. (1). Since \\bar Y_P2 = (N_UL/N)\\bar Y_P1 + (N_LB/N)\\bar Y^{(LB)}, at the reported good regions R_LB is 30-50% (and 10% in Table 5), so a substantial fraction of \\bar Y_P2 is the exact CG average over the labeled subset. Because \\bar Y^{(LB)} estimates the same ensemble average as \\bar Y^{Orig} from a subset of the same configurations, EC-1 score 2 and R_sigma <= 1.2 can be satisfied even if \\bar Y_P1 is badly biased or noisy. No results for \\bar Y_P1 are reported anywhere, and no control with R_TR = 0 is shown as a baseline. Tables 3-5 therefore establish only that a weighted average containing an exact-CG component agrees with the full CG result; the title and the summary claim about ML estimation outrun this evidence. Please report \\bar Y_P1 (and its EC scores) for the same grid, or restrict the claims to \\bar Y_P2.","section":"§2, Eq. (1)"},{"comment":"The bias-correction step adds the mean residual on S_BC to the ML predictions on S_UL. This assumes that the model bias on the unlabeled configurations equals the bias on the labeled configurations. The paper itself cautions about \"unwanted fluctuations or weird bias due to the partial data usage\" (Section 2), yet no diagnostic is given: for example, one could hold out a second labeled slice and compare \\bar Y_P1 computed on it with the direct CG average, or examine residual distributions across bootstrap resamples. Without such a check, unbiasedness of \\bar Y_P1 is an assumption, and by extension the EC-1 agreements and the good regions in Tables 3-5 do not demonstrate that the ML estimate matches CG on unlabeled data.","section":"§2, Eq. (1)"},{"comment":"The bootstrap error for \\bar Y_P1 and \\bar Y_P2 is computed by resampling configurations with N_BS = 10,000, but the text does not state that the model is retrained on each bootstrap sample. With a fixed trained model, the bootstrap variance omits the training-set split variance, i.e., the randomness in choosing S_TR and S_BC. This makes sigma_Q too small and likely biases R_sigma below its true value. Please clarify whether retraining is performed inside the bootstrap; if not, use nested resampling or a separate Monte Carlo over splits, and report both variance components.","section":"§2, bootstrap error"},{"comment":"The success threshold R_sigma <= 1.2 is described as \"tentatively determined empirically, monitoring our preliminary results\" and is applied to the same datasets that were monitored. This makes the criterion post-hoc; the reported good regions may reflect selection of the threshold rather than a robust property of the estimator. Please provide a pre-specified or independently motivated threshold, or show that the good regions are insensitive to the threshold within a reasonable range (e.g., 1.1-1.5).","section":"§3, EC-2"}],"minor_comments":[{"comment":"The caption says the ID-1 row is bold, but in the manuscript text the row does not appear bold; please fix the formatting or the caption.","section":"Table 1"},{"comment":"The color scale for R_sigma in the EC-2 panels is not numerically labeled, so the reader cannot read off the actual R_sigma values; adding a quantitative color bar would improve reproducibility.","section":"Figs. 3-5"},{"comment":"The LightGBM hyperparameters (40 stages, depth 3, learning rate 0.1, subsampling 0.7) are fixed, but no validation or sensitivity study is reported; a sentence on how these values were selected would help assess robustness.","section":"§2, hyperparameters"}],"recommendation":"major_revision","confidential_remarks":"This is an honest preliminary proceedings paper, and I do not see a fundamental methodological error that would require rejection. The main gap is quite concrete: the evaluation must be reported for \\bar Y_P1, and the error accounting should include the training-split variance. If the authors can provide these, the paper could become acceptable. The EC-2 threshold issue should also be addressed, but it is secondary."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is an honest, clearly written feasibility study of gradient-boosted decision trees for estimating Tr M^{-n} in lattice QCD. What is genuinely new is the application to n=2,3,4 using plaquette and Polyakov loop inputs, with a systematic scan over labeled and training fractions. The finding that heavier-quark and phase-transition ensembles are easier to predict is plausible and potentially useful. The authors also deserve credit for reporting failure cases (ID-2, n=3,4) and for having an explicit to-do list. The bias-correction methodology follows Yoon et al. correctly, and the paper is transparent about its preliminary status.\n\nThe main soft spot is real and load-bearing. All the reported success criteria (EC-1 scores, R_sigma, and the good regions in Tables 3-5) apply only to the P2 estimator of Eq. (2), not the pure ML estimator P1 of Eq. (1). At the claimed good regions, R_LB is 30-50%, so 30-50% of P2 is just the exact CG average over the labeled subset. That component trivially pulls the estimate toward the full CG result. Since P1 results are never shown, the paper does not actually demonstrate that the GBDT model itself reproduces Tr M^{-n}. The title and summary overstate what is established. This is not a minor omission; it is the difference between testing ML and testing a weighted average that contains ML.\n\nTwo other issues are secondary but worth noting. First, the bootstrap errors are computed over gauge configurations with a fixed trained model, so the variance from the training-set split is not included. That likely makes sigma_Q too small and the R_sigma<=1.2 criterion easier to pass. Second, the R_sigma threshold is admittedly chosen post hoc on the same datasets. These are common in proceedings papers, but they further temper the strength of the claims.\n\nThe cost-reduction message is still credible: you avoid CG on the unlabeled configurations. But the effective saving is less than the R_LB=10% figure for the phase-transition ensemble if you only count the exact-CG fraction.\n\nWho is this for? Lattice practitioners thinking about cheap estimators for chiral observables. I would not cite it as a demonstration that ML alone works, but I would reference it as a preliminary study with an important methodological lesson: report P1 and P2 separately, and include training variance in the error budget.\n\nMy recommendation: send to peer review. A serious referee can ask for P1 results, a proper error budget that includes train/test variability, and a justified or pre-specified EC-2 threshold. The paper is coherent, honest, and worth engaging with, but it needs revision before the central claim is supported.","headline":"Useful, honest feasibility study, but the central claim is only tested for a hybrid estimator that includes up to half exact CG data; P1 results and a proper variance budget are missing.","tokens_in":8443,"tokens_out":3474,"would_cite":false,"duration_ms":31895,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["12.38.Gc","11.15.Ha"],"model":"deepseek-v4-flash","headline":"Given enough labeled data, a bias-corrected gradient boosting estimator reproduces conjugate-gradient trace results for $\\text{Tr}\\,M^{-n}$ within $1\\sigma$, and it works best at a first-order phase transition.","keywords":["lattice QCD","trace of inverse Dirac operator","gradient boosting decision tree","bias correction","stochastic estimator","conjugate gradient solver","first-order phase transition","Wilson clover action"],"falsifier":"A direct test is to rerun the analysis with the same $X,Y$ pairs and the same claimed-good values such as $R_{\\rm LB}=30\\%$, $R_{\\rm TR}=40\\%$ for $(X,Y)=({\\rm Plaquette},\\text{Tr}\\,M^{-3})$ on the ID-0 data, but with the CG traces already known for the unlabeled set, and then compare $\\bar{Y}^{P1}$ against the true mean on $S_{\\rm UL}$. The assumption fails if the difference is significantly larger than the mean residual on $S_{\\rm BC}$.","tokens_in":7410,"feed_emoji":"⚛️","tokens_out":21033,"duration_ms":158171,"temperature":0.7,"pith_summary":"Lattice QCD calculations often need the trace of an inverse power of the Dirac operator, $\\text{Tr}\\,M^{-n}$, which normally requires many conjugate-gradient solves. This paper tries to replace most of those solves with a gradient boosting decision tree regressor that learns $\\text{Tr}\\,M^{-n}$ from cheaper observables such as the plaquette, the Polyakov loop, or lower inverse powers of the Dirac operator. The proposed estimator applies a bias correction: the average residual on a labeled bias-correction set is added to the machine-learning predictions on the unlabeled set. The authors report that with a large enough labeled fraction this corrected estimator matches the original conjugate-gradient trace within $1\\sigma$ and with statistical error no more than 1.2 times the original, with the smallest labeled fraction ($R_{\\rm LB}\\gtrsim10\\%$, the fraction of configurations with known CG traces) needed on the dataset containing a first-order phase transition. If this holds, the expensive inversion would only be needed on a small labeled subset of configurations, and the rest of the ensemble would be predicted essentially for free.","feed_headline":"Cut lattice trace costs: boosted trees match CG with 10–40% of labels","feed_subtitle":"Bias-corrected gradient boosting reproduces conjugate-gradient traces within 1 sigma, even near a phase transition.","key_machinery":"The load-bearing mechanism is the gradient boosting decision tree regressor combined with the bias-correction prescription of Eqs. (1)--(2). Gradient boosting builds $f(X)$ by adding shallow decision trees sequentially, each new tree fitting the residual left by the previous ensemble; the implementation used here runs 40 depth-3 trees with learning rate 0.1 and sub-sampling 0.7. The bias correction is what lets the model be wrong configuration-by-configuration as long as it is right on average: $\\bar{Y}^{P1}$ transfers the mean residual computed on the labeled bias-correction set $S_{\\rm BC}$ to the mean prediction on the unlabeled set $S_{\\rm UL}$, and $\\bar{Y}^{P2}$ blends this with the direct CG average on all labeled data. The acceptance criteria are EC-1 (the two means agree within each other's $1\\sigma$ error, called score 2) and EC-2 ($R_\\sigma=\\sigma_Q/\\sigma_{\\rm Orig}\\le1.2$), with errors estimated by $10{,}000$ bootstrap resamples. The method's precondition, checked in the paper's correlation matrices, is a strong correlation between the chosen input observable and the target trace.","core_discovery":"The central claim is that the bias-corrected estimators of Eqs. (1)--(2), obtained with the gradient boosting decision tree regression, reproduce the original conjugate-gradient (CG) trace estimates for $\\text{Tr}\\,M^{-n}$ within $1\\sigma$ and with $R_\\sigma \\le 1.2$ in well-defined regions of the two resolution parameters: the labeled fraction $R_{\\rm LB}=N_{\\rm LB}/N$ and the training fraction $R_{\\rm TR}=N_{\\rm TR}/N_{\\rm LB}$. The first estimator $\\bar{Y}^{P1}$ adds the mean residual on the bias-correction set $S^{Y}_{\\rm BC}$ to the mean machine-learning prediction on the unlabeled set $S^{Y}_{\\rm UL}$; the second estimator $\\bar{Y}^{P2}$ takes the weighted average of $\\bar{Y}^{P1}$ with the direct CG mean on the labeled set. On the heaviest-quark (ID-0) data the method succeeds for $n=1,2,3$ at $R_{\\rm LB}\\gtrsim40\\%$, with examples such as $X={\\rm Plaquette}$, $Y=\\text{Tr}\\,M^{-3}$ succeeding already at $R_{\\rm LB}\\gtrsim30\\%$, $R_{\\rm TR}\\lesssim40\\%$; on the first-order-transition (ID-1) data good agreement starts at $R_{\\rm LB}\\gtrsim10\\%$ for $n=1,2,3$; on the lightest-quark (ID-2) data only $n=1,2$ work, at $R_{\\rm LB}\\gtrsim40\\%$. The trace $\\text{Tr}\\,M^{-4}$ is harder in every dataset, requiring $R_{\\rm LB}\\gtrsim40\\%$ even at the phase transition and working best with input $X=\\text{Tr}\\,M^{-3}$. The paper concludes that ML estimation of $\\text{Tr}\\,M^{-n}$ works better with heavier quark mass and works 'quite well' when the dataset contains the first-order phase transition.","pith_inferences":["The paper does not test whether the bias correction transfers across ensembles; a natural extension is to train on one $\\kappa$ value and bias-correct on another, which would reveal whether the good regions survive outside the single-ensemble setting.","Because Eq. (1) applies a single additive bias correction, any input-dependent error (for example, larger error in one phase than the other) survives; fitting a second regressor to the residuals and re-checking the good regions would test how much this matters.","The paper does not report wall-clock timings, so the claimed cost reduction is implicit; measuring CG solver time against training and prediction time would show whether the $R_{\\rm LB}=10\\%$--$40\\%$ regions actually save time in production.","The strong performance at the first-order phase transition suggests the model is learning the two-phase mixture; checking whether the predicted traces on $S_{\\rm UL}$ reproduce the bimodal distribution of the CG data would make that mechanism visible."],"forward_implications":["If the central claim holds, trace estimation can be reduced to labeling a fraction $R_{\\rm LB}$ of configurations with the CG solver, training the regressor, and predicting the rest, cutting the dominant cost of the calculation.","The method performs best on the dataset containing a first-order phase transition, where $R_{\\rm LB}\\gtrsim10\\%$ suffices for $n=1,2,3$, so the cheap estimator is most useful precisely where the trace signal is hardest to obtain.","For $\\text{Tr}\\,M^{-4}$ the labeled fraction has to stay near $R_{\\rm LB}\\gtrsim40\\%$ in all three datasets, and $X=\\text{Tr}\\,M^{-3}$ is the recommended input.","The paper's EC-1 score-2 plus $R_\\sigma\\le1.2$ rule gives a practical, tentative criterion for deciding when the ML estimator can replace the CG result.","The planned use of the method is to compute chiral order-parameter cumulants such as susceptibility, skewness, and kurtosis from the ML traces, which the paper lists as work in progress."],"supporting_citations":[{"why":"Supplies the stochastic noise estimator used to obtain the trace values that the ML method is trained to reproduce.","marker":"[1]"},{"why":"Provides the gradient boosting regression algorithm that the whole study is built on.","marker":"[5]"},{"why":"Supplies the training-versus-bias-correction split methodology from which Eqs. (1) and (2) are adapted.","marker":"[6]"},{"why":"Provides the three gauge ensembles and the conjugate-gradient trace data used for the analysis.","marker":"[8]"},{"why":"Supplies the specific gradient boosting implementation used, with the hyperparameters reported in the paper.","marker":"[14]"},{"why":"Supplies the Julia interface through which the gradient boosting implementation is run.","marker":"[15]"}],"fun_headline_variants":["Gradient boosting cuts Dirac trace cost, matches CG at 30% labels","Boosted trees match CG traces with 10-40% of stochastic sources","Bias-corrected ML: Dirac traces from 30% labels within 1 sigma","Phase transition data: GBDT traces match CG at 10% labels","Cheap trace estimation: gradient boosting rivals CG, bias-corrected"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The average model error on the labeled bias-correction set is assumed to equal the error on the unlabeled set, an assumption the paper itself flags as potentially suffering from 'unwanted fluctuations or weird bias due to the partial data usage'.","fun_headline_variants_meta":{"raw":{"variants":["Gradient boosting cuts Dirac trace cost, matches CG at 30% labels","Boosted trees match CG traces with 10-40% of stochastic sources","Bias-corrected ML: Dirac traces from 30% labels within 1 sigma","Phase transition data: GBDT traces match CG at 10% labels","Cheap trace estimation: gradient boosting rivals CG, bias-corrected"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000995,"raw_usage":{"total_tokens":4277,"prompt_tokens":1070,"completion_tokens":3207,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":686,"completion_tokens_details":{"reasoning_tokens":3104}},"tokens_in":686,"tokens_out":3207,"duration_ms":23141,"temperature":1.0,"reasoning_tokens":3104,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T11:27:13.792496+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A direct test is to rerun the analysis with the same $X,Y$ pairs and the same claimed-good values such as $R_{\\rm LB}=30\\%$, $R_{\\rm TR}=40\\%$ for $(X,Y)=({\\rm Plaquette},\\text{Tr}\\,M^{-3})$ on the ID-0 data, but with the CG traces already known for the unlabeled set, and then compare $\\bar{Y}^{P1}$ against the true mean on $S_{\\rm UL}$. The assumption fails if the difference is significantly larger than the mean residual on $S_{\\rm BC}$.","supporting_citations":[{"cited_title":"Friedman,Greedy function approximation: A gradient boosting machine., The Annals of Statistics29(2001) 1189","cited_arxiv_id":null,"evidence_quote":"Provides the gradient boosting regression algorithm that the whole study is built on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the specific gradient boosting implementation used, with the hyperparameters reported in the paper."},{"cited_title":"Blaom, F","cited_arxiv_id":null,"evidence_quote":"Supplies the Julia interface through which the gradient boosting implementation is run."}],"review_version":1}