{"id":"06f35963-38b3-4ccf-a2c5-6dcacafdca5c","arxiv_id":"2412.21147","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A source-time split with bias correction lets ML models infer lattice correlators at nearby mass parameters with per-configuration uncertainty estimates that match truth-level fit results.","lead":"This paper trains machine learning models to reconstruct lattice QCD correlation functions at one quark-mass parameter from correlators at another, using a source-time splitting scheme with a bias-correction step. The authors test four model types and a ratio-method benchmark on a single gauge ensemble and report that all give fitted energy and amplitude parameters close to the true values.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The unbiasedness of Eq. (1) rests on the unverified assumption that the ML bias on the 18 unlabeled source times equals the bias measured on the 5 bias-correction source times; the paper shows only an aggregate comparison, so a source-time-resolved bias check is needed.","rationale":"The paper's central claim is that the bias-corrected estimator in Eq. (1) provides per-configuration predictions of a target-mass correlator from which unbiased spectral parameters can be extracted, with statistical uncertainties computed without bootstrapping or repetitive training. The weakest link is the source-time generalization step: the correction term averages residuals over only 5 bias-correction source times, while the prediction term averages over 18 unlabeled source times. If the model's bias is not identical across source times, the correction is systematically wrong. The reader identified exactly this assumption. I agree, and I add a concrete way to test it: because the paper has truth-level O2 for all source times, the actual residual on the UD source times can be compared with the BC residual. If the difference is statistically zero, the concern is resolved. I also note that lattice translation invariance may make the assumption true, but the paper does not demonstrate that the trained model preserves it. The CNN excited-state deviations in Table 3 are consistent with a residual shape bias and strengthen the need for this check. The paper has useful structure and a sensible benchmark, but the missing source-time-resolved validation leaves the unbiasedness claim conditional. This does not change the reader's CONDITIONAL verdict; it sharpens the specific condition that must be verified.","tokens_in":5230,"tokens_out":15106,"duration_ms":157975,"concrete_test":"Use the truth-level O2 values already present in the dataset on the 18 UD source times to compute Delta_UD = mean_{i,j in UD}(O2_{i,j} - f(O1_{i,j})) and compare it with Delta_BC = mean_{i,k in BC}(O2_{i,k} - f(O1_{i,k})). If |Delta_UD - Delta_BC| is not consistent with zero at roughly the 1-sigma level, Eq. (1) is biased. A stronger version is leave-one-source-time-out: for each of the 24 source times, train on one, correct with five, and evaluate on a held-out sixth source time, repeating over all choices to map the bias as a function of source time.","verdict_should_be":"UNCHANGED","load_bearing_attack":"In Eq. (1), O_i = <O_pred>_UD + <O - O_pred>_BC. This estimator is unbiased only if E_UD[O_pred] + E_BC[O - O_pred] = E[O]. If the joint distribution of (O1, O2) is identical for every source time, the identity holds for any model, but the paper never establishes that this stationarity survives after training. The model is trained on a single source-time label per configuration, and the BC correction is a single average over 5 source times. The residual <O - O_pred> could vary with source time due to boundary effects, source-time-dependent noise correlations, or the model exploiting source-time-specific artifacts. Table 3 shows the CNN's excited-state parameters a1 and dE1 deviate from truth by roughly 2.4 sigma, which is consistent with an uncorrected shape bias. Figure 1 (right) shows only the aggregate relative correlated difference, averaged over source times, so it cannot detect such variation. Until a source-time-resolved residual comparison is provided, the central unbiasedness claim of Eq. (1) is not fully supported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript proposes a bias-corrected machine-learning estimator, Eq. (1), for inferring lattice two-point correlation functions at nearby quark masses. Source times are split into training, bias-correction, and unlabeled sets while preserving the configuration index, so the method yields per-configuration predictions that can be used with standard error propagation. The authors benchmark MLP, CNN, transformer, and tree-regressor models, as well as ratio-method variants, against truth-level correlators and spectral fits, reporting ground-state and first-excited-state amplitudes and energy splittings in Tables 3 and 4. The central claims are that the bias correction makes the estimator unbiased and that the per-configuration setup avoids bootstrapping or repeated training.","tokens_in":5487,"tokens_out":5402,"duration_ms":53985,"significance":"If the unbiasedness and efficiency claims hold, this is a useful analysis-level variance-reduction scheme: it replaces expensive repeated quark propagator computations with a one-time ML training and yields per-configuration observables that fit naturally into existing lattice analysis pipelines. The comparison across multiple architectures and against ratio methods is informative and the validation against held-out truth data is a genuine strength. However, the present evidence is incomplete in two load-bearing respects: the source-time stationarity of the ML residual that underpins Eq. (1) is not demonstrated, and the claimed computational efficiency is not quantified. The excited-state deviations in Table 3 indicate that the bias correction may leave a systematic shape error, so the significance of the method depends on resolving those points.","major_comments":[{"comment":"The unbiasedness of the estimator O_i = <O_pred>_UD + <O - O_pred>_BC requires that the expected residual E[O - O_pred] be identical on the UD and BC source-time sets. The manuscript does not establish this stationarity: the BC correction is a single average over five source times, while the UD set spans eighteen source times, and boundary effects or source-time-dependent noise correlations could make the residual vary. Figure 1 (right) shows only an aggregate relative correlated difference averaged over source times and cannot detect such variation. Please provide a source-time-resolved comparison of O_i(tau) - O_pred_i(tau) (or of fitted parameters obtained from each source-time slice) for the BC and UD subsets, or give a theoretical argument that the trained model's residual is stationary in the source-time index.","section":"Section 3, Eq. (1)"},{"comment":"For the CNN, the first excited-state amplitude and splitting deviate from the truth-level values by approximately 2.4 standard deviations (a1: 0.0740(18) vs 0.0689(10); dE1: 0.1938(36) vs 0.1838(21)). The quoted uncertainties are statistical only; no systematic uncertainty is added for model choice, training seed, or residual bias. Since the excited-state parameters are precisely where an uncorrected shape bias in the estimator would appear, the claim in Section 6 of 'good agreement' needs qualification. Please add a systematic uncertainty estimated from the spread across models/architectures or from the residual bias, or present a joint test that accounts for the number of fitted parameters.","section":"Table 3"},{"comment":"The title and Key Feature 1 claim that the method is 'efficient' and avoids 'intensive bootstrapping or repetitive training,' but no computational cost comparison is reported. The reader cannot assess whether the ML training overhead plus the need for O2 measurements on the BC set is actually cheaper than standard analysis or than the ratio method alone. Please provide wall-clock time or floating-point operation counts for training, inference, and bias correction relative to the cost of the original correlator computation, or soften the efficiency claim accordingly.","section":"Section 1, Key Features"}],"minor_comments":[{"comment":"The sentence 'The training, bias-correction, and unlabeled sets each consist of 1024 configurations selected for each of 1, 5, and 18 (24 total) time source labels' is ambiguous; please clarify how the 24 source times are divided among the three sets and whether the configuration sets overlap.","section":"Section 2"},{"comment":"The caption says 'Comparison of predicted correlators (MLP)' but the legend includes Ratio Method curves; please specify which model and which ratio variant corresponds to each curve.","section":"Figure 2"},{"comment":"The oscillating factor (-1)^{n(t+1)} is unusual; please clarify the staggered-state convention, since Eq. (4) is used for the fits in Table 3.","section":"Equation (4)"},{"comment":"In Eq. (6), the exponent alpha is 'tailored to minimize the uncertainty' of the boosted ratio; please state how alpha is determined (e.g., from the same data or a separate scan) and whether the quoted uncertainties in Table 4 include the uncertainty from estimating alpha.","section":"Section 5, Eq. (6)"}],"recommendation":"major_revision","confidential_remarks":"For the editor: the manuscript is a proceedings contribution, and the source-time-resolved bias check is feasible with the existing dataset. The main risk is that the excited-state deviations in Table 3 indicate a real shape bias, so the authors should address this before publication. I would not reject, but the central unbiasedness claim needs direct support and the efficiency claim needs quantitative backing."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The thing to know about this paper: it's a short proceedings-style methods note on using ML to infer two-point lattice correlators at nearby masses, and the central bias-corrected estimator seems to work for the ground state. The new piece is the source-time split into training, bias-correction, and unlabeled sets while preserving the configuration axis, so you get per-configuration predictions without retraining. That is a legitimate extension of the truncated solver method and Yoon et al.'s ML estimator, and it's clearly explained. The numerical comparison against truth-level data is the right test, and for a0 and dE0 all models agree with truth within uncertainties. That is the paper's real strength: a clean, honest benchmark on a realistic ensemble.\n\nThe soft spots are real but not fatal for a proceedings paper. First, the excited-state parameters a1 and dE1 deviate from truth by up to 2.4 sigma for the CNN, and the paper adds no systematic uncertainty for method or model bias. It may be that the CNN's shape bias is not fully corrected, but the paper doesn't investigate. Second, the unbiasedness of Eq. (1) rests on the assumption that the ML prediction bias measured on the 5 bias-correction source times is representative of the 18 unlabeled source times. The paper shows only an aggregate comparison over source times, so that assumption is not directly tested. This is the load-bearing assumption, and it is testable: plot the residual <O - O_pred> as a function of source time. Third, the efficiency claim is not supported by any cost comparison; the paper says ML saves computational expense but gives no estimate of training cost against the cost of extra inversions. Fourth, the boosted-ratio exponent alpha in Eq. (6) is tuned on the same data it evaluates, which weakens but does not invalidate that benchmark. The citation pattern is fine; the relevant prior work is cited and the relationship to it is stated.\n\nFor a Lattice proceedings writeup, this is a reasonable contribution. It is not a breakthrough, but it describes a workable method with a mostly honest validation. The missing systematic uncertainties and the source-time stationarity check are addressable in a final version or a follow-up paper. I would send it to peer review because it is clearly written, the core idea is plausible, and the numerical evidence is positive enough to warrant referee time. I wouldn't cite it myself yet, but I would not be surprised if the method becomes useful after the caveats are closed.\n\nRecommendation: engage with it, but ask for the source-time-resolved bias check and an explicit discussion of the excited-state deviations and the alpha tuning before considering it final.","headline":"A competent methods note on ML-inferred lattice correlators; the ground-state results hold together, but the central unbiasedness claim needs a source-time-resolved check and the paper would benefit from a fuller systematic uncertainty budget.","tokens_in":6018,"tokens_out":969,"would_cite":false,"duration_ms":12097,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["12.38.Gc","11.15.Ha"],"model":"deepseek-v4-flash","headline":"This paper claims that a bias-corrected machine-learning estimator can infer lattice correlation functions at new quark masses from a single ensemble, with spectral fit parameters matching truth-level data within quoted uncertainties.","keywords":["lattice QCD","machine learning","correlation functions","bias correction","per-configuration inference","ratio estimator","quark mass extrapolation","uncertainty estimation"],"falsifier":"Take a held-out set of configurations from the same ensemble where both $O_1$ and $O_2$ are measured at all 24 source times, apply the trained model and Eq. (1) using only the same split, and compare the corrected per-configuration predictions to the truth. If the mean residual on the held-out unlabeled source times differs from the bias-correction residual by more than the statistical uncertainty, the estimator's quoted covariance is missing a systematic source-time-dependent bias.","tokens_in":5043,"feed_emoji":"⚛️","tokens_out":7117,"duration_ms":62899,"temperature":0.7,"pith_summary":"Lattice QCD correlators are expensive to produce, and this paper asks whether supervised learning can infer new ones from existing ones at the analysis stage rather than by running new simulations. It proposes splitting data along the time-source index into training, bias-correction, and unlabeled subsets, so that the model makes per-configuration predictions and every configuration keeps its own statistical identity. The central estimator subtracts the average prediction bias measured on the bias-correction set from the raw model predictions on the unlabeled set. On a heavy-light meson correlator dataset, the bias-corrected predictions yield fitted ground- and excited-state amplitudes and energy splittings that agree with the truth-level fits. A simple ratio method serves as benchmark and matches or approaches the ML results, showing that some of the gain comes from using correlated observables rather than from the neural network alone.","feed_headline":"Bias-corrected ML matches truth-level fits for lattice correlators","feed_subtitle":"Per-configuration predictions give statistical uncertainties without bootstrapping or retraining; a ratio benchmark matches most of the…","key_machinery":"The load-bearing object is the bias-corrected per-configuration estimator of Eq. (1), together with the data split that makes it possible. Time sources are partitioned into a training set (both source and target observables measured, used to train the network), a bias-correction set (both measured, used to estimate the model's average prediction error), and an unlabeled set (only the source observable measured, where the model predicts the target). Because every configuration appears in all three sets, the correction term $\\langle O_i(\\tau)-O^{\\rm pred}_i(\\tau)\\rangle_{\\rm BC}$ is computed for the same configuration $i$, which is what converts a black-box prediction into a configurable observable with honest statistical errors.","core_discovery":"The paper's central claim is that Eq. (1), $O_i(\\tau)=\\langle O_i(\\tau)_{\\rm pred}\\rangle_{\\rm UD}+\\langle O_i(\\tau)-O^{\\rm pred}_i(\\tau)\\rangle_{\\rm BC}$, produces per-configuration estimators of a target-mass two-point function that are unbiased enough that spectral fits give parameters ($a_0$, $dE_0$, $a_1$, $dE_1$) agreeing with truth-level data within quoted errors. The method treats source time as the index that separates training, bias-correction, and unlabeled data while preserving the configuration axis, so the corrected $O_i$ can be used as ordinary per-measurement observables. The paper benchmarks MLP, CNN, transformer, and decision-tree regressors and finds all four yield fit parameters consistent with truth; it also shows the bias correction visibly reduces the correlated difference of predicted correlators.","pith_inferences":["If the bias correction remains representative across source times and masses, the same data set could be reused to infer correlators at interpolated quark masses without new simulations, reducing the cost of scanning parameter space.","A natural extension is to three-point functions or smeared operators, where the per-configuration split would work the same way as long as a cheaper correlated observable exists for every configuration.","The method's practical limit is the assumption that ML error does not vary with source time; if it does, the Eq. (1) correction is incomplete and the quoted covariance misses a systematic term.","One testable extension is to use the variance of predictions across network training seeds as a model-error estimate and compare it with the residual bias on a fully measured held-out set."],"forward_implications":["Spectral fit parameters from bias-corrected MLP, CNN, transformer, and decision-tree predictions all overlap the truth-level values in Table 3, so the estimator preserves the physics content of the correlator.","Per-configuration predictions mean means, covariances, and resampling can be computed directly, without bootstrapping or retraining the network for each resample.","The ratio method and its boosted variant match or come close to the ML fits while using far less machinery, suggesting correlations between observables carry much of the statistical power.","The bias-correction step is essential: Figure 1 shows the relative correlated difference drops substantially after applying Eq. (1).","The method transfers across heavy-quark masses: training on $am_h=0.548$ and strange valence mass and predicting at $am_h=0.164$ works for these ensembles."],"supporting_citations":[{"why":"It supplies the ML estimator formulation that Eq. (1) adapts to per-configuration lattice correlators.","marker":"[1]"},{"why":"It is one of the truncated-solver-method references whose variance-reduction idea motivates splitting by source time.","marker":"[3]"},{"why":"It provides another truncated-solver-method reference used as inspiration for the bias-correction structure.","marker":"[4]"},{"why":"It gives the class of variance-reduction techniques using lattice symmetries that the setup extends.","marker":"[5]"},{"why":"It gives the gauge ensemble parameters on which all correlators are computed.","marker":"[6]"},{"why":"It provides the dataset of meson two-point functions from the semileptonic decay program.","marker":"[7]"},{"why":"It introduces the ratio estimators used as the benchmark in Section 5.","marker":"[8]"},{"why":"It is the software package used to handle correlated errors in the spectral fits.","marker":"[9]"},{"why":"It is the software package used for least-squares fitting of the spectral decomposition.","marker":"[10]"},{"why":"It is the software package used for the correlator fitting machinery.","marker":"[11]"}],"fun_headline_variants":["Bias-corrected AI predicts lattice correlators, matching truth-level fits","ML per-config estimators give unbiased lattice fits without bootstrapping","ML-bias-corrected lattice correlators produce truth-level spectral fits","AI infers lattice correlators across masses with bias-corrected fidelity"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The method assumes that the average prediction bias measured on the five-source-time bias-correction set equals the bias on the eighteen-source-time unlabeled set for every configuration, so that subtracting the one fixes the other; if ML error varies with source time, the correction leaves a systematic bias that the quoted covariance does not include.","fun_headline_variants_meta":{"raw":{"variants":["Bias-corrected AI predicts lattice correlators, matching truth-level fits","ML per-config estimators give unbiased lattice fits without bootstrapping","ML-bias-corrected lattice correlators produce truth-level spectral fits","AI infers lattice correlators across masses with bias-corrected fidelity"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00069,"raw_usage":{"total_tokens":3063,"prompt_tokens":819,"completion_tokens":2244,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":435,"completion_tokens_details":{"reasoning_tokens":2166}},"tokens_in":435,"tokens_out":2244,"duration_ms":15896,"temperature":1.0,"reasoning_tokens":2166,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T23:01:42.542079+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a held-out set of configurations from the same ensemble where both $O_1$ and $O_2$ are measured at all 24 source times, apply the trained model and Eq. (1) using only the same split, and compare the corrected per-configuration predictions to the truth. If the mean residual on the held-out unlabeled source times differs from the bias-correction residual by more than the statistical uncertainty, the estimator's quoted covariance is missing a systematic source-time-dependent bias.","supporting_citations":[{"cited_title":"Calculation of fermion loops for $\\eta^\\prime$ and nucleon scalar and electromagnetic form factors","cited_arxiv_id":"1108.2473","evidence_quote":"It provides another truncated-solver-method reference used as inspiration for the bias-correction structure."},{"cited_title":"Peter Lepage,gvar(Version 13.1), 2024","cited_arxiv_id":null,"evidence_quote":"It is the software package used to handle correlated errors in the spectral fits."},{"cited_title":"Peter Lepage,lsqfit (Version 13.2.2), 2024","cited_arxiv_id":null,"evidence_quote":"It is the software package used for least-squares fitting of the spectral decomposition."},{"cited_title":"Peter Lepage,corrfitter (Version 8.2), 2024","cited_arxiv_id":null,"evidence_quote":"It is the software package used for the correlator fitting machinery."}],"review_version":1}