{"id":"47201b4f-234b-485e-b140-086646a0a181","arxiv_id":"2608.06450","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"An amplitude surrogate trained on a few thousand exact LHC amplitude points statistically outperforms the training data, with largest amplification in sparsely populated kinematic tails of Z+g and Z+4g production.","lead":"This paper shows that a neural network trained on a small set of exact LHC amplitude calculations can effectively replace a much larger Monte Carlo sample, especially in rare kinematic regions. It quantifies this statistical amplification for Z-plus-gluons processes and finds the surrogate's benefit is largest precisely where traditional simulation runs out of events.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Reported G factors divide by in-window training counts (3 or 7) instead of the total training sample; with the standard total-ntrain normalization the factors are ≤0.56, so the claimed order-of-magnitude amplification in Eqs.(43),(46) is an artifact of the denominator choice.","rationale":"The reader's weakest assumption (σ_syst calibration) is not the most load-bearing part: the amplification factors in §3 are extracted from bootstrap replicas of the weighted tail sample (Eq.(42) and Figure 3), so the plateau of MI is an empirical estimate of the surrogate's squared bias, not a propagation of the learned σ_syst from Eq.(38). Even a perfectly calibrated σ_syst would not change G as read off the bootstrap plateau. The genuinely load-bearing issue is the normalization of G. Eq.(5) borrows the generative-network definition G=n_equiv/n_train, but in Eqs.(43),(46) the paper plugs n_train=3 and 7 — the number of training points inside the tail window — instead of the total training sample size (1M). This is inconsistent with the generative-network comparison in Ref.[19], where n_train is the total training set size. With the total-training denominator, G becomes 0.06 (Z+g) and 0.56 (Z+4g), both below 1, directly contradicting the abstract's claim of order-of-magnitude amplification and of far outperforming generative networks. The in-window denominator is also statistically fragile (3 or 7 counts, no error bars, sensitive to bin choice). Because this is a definitional and normalization error, it is addressable in revision; the reader's CONDITIONAL verdict remains appropriate, but the primary condition should be a re-derivation of G under a justified normalization rather than only the σ_syst calibration check.","tokens_in":15557,"tokens_out":18976,"duration_ms":166339,"concrete_test":"Recompute G for Eqs.(43) and (46) using total training size n_train=10k, 100k, 1M in the denominator, as in the generative-amplification definition of Ref.[19]; additionally report bootstrap uncertainties on n_equiv and on the in-window training count. If G ≤ 1 for the 1M-trained surrogates under the total-ntrain normalization, the abstract's claim of order-of-magnitude amplification and of outperforming generative networks is unsupported and should be revised.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Eq.(5) defines the amplification factor as G = n_equiv/n_train, following the generative-network literature (Refs.[16,19]) where n_train is the total number of training samples. In §3.1 and §3.2 the paper instead evaluates Eq.(43) G=60,000/3=20,000 and Eq.(46) G=560,000/7=80,000, taking n_train as the number of training events that happen to fall inside the targeted tail window ([700,800] GeV for Z+g, [1200,1400] GeV for Z+4g). The surrogate is trained on the full 1M-event sample; the tail prediction is an extrapolation constrained by all 1M amplitude evaluations, so the in-window count (3 or 7) is not the number of constraints supporting that prediction. If the standard normalization with total training size n_train=1,000,000 is used, the same n_equiv values give G=0.06 (Z+g) and G=0.56 (Z+4g), i.e. no amplification. The large in-window G values are also noisy: with only 3 or 7 in-window points, a single training-event fluctuation changes G by 33% or 14%, yet no error bars are reported. The paper's headline claim that 'surrogate Monte Carlo far outperforms the density estimation in current generative networks' rests on these inflated factors; a fair comparison at equal total training cost would require the same n_train normalization used for the generative models in Ref.[19].","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript adapts the concept of \"generative amplification\" from generative networks to amplitude surrogates for LHC event generation. For a surrogate trained on n_train exact amplitude evaluations, the authors define the equivalent sample size n_equiv as the number of true Monte Carlo events whose statistical uncertainty matches the surrogate's systematic model uncertainty, and the amplification factor G = n_equiv / n_train. Two measures are introduced: an averaging measure based on the variance of a reweighted estimator of a fiducial rate (Section 2.1), and a differential measure based on the Kolmogorov-Smirnov statistic between weighted surrogate events and an independent truth sample (Section 2.2). Stratified sampling is used to populate kinematic tails (Section 2.3). The method is applied to Z+ng (n=1,...,4) production with an uncertainty-aware Lorentz-equivariant transformer (L-GATr-slim), reporting large amplification factors in high-pT tails, e.g., G = 20,000 (averaging) and G = 2,670 (differential) for Z+g in pT in [700,800] GeV, and G = 80,000 and G = 290 for Z+4g in pT in [1200,1400] GeV. The paper concludes that surrogate Monte Carlo \"far outperforms\" density estimation in current generative networks.","tokens_in":15887,"tokens_out":10116,"duration_ms":79281,"significance":"If the results are correct, the paper provides a useful framework for quantifying the statistical benefit of amplitude surrogates, which is directly relevant for precision LHC simulation at the HL-LHC. Strengths include the clean derivation of the statistical variance using the Kish effective sample size, the explicit reweighting formulation that avoids generating surrogate events, the use of an independent truth reference for the differential measure, and the bootstrap evaluation of the amplification metrics directly against truth samples. The hyperparameters and training details are fully documented. The main reported quantitative claims, however, rest on a nonstandard normalization of n_train that is inconsistent with Eq. (5) and with the generative-network literature, and on an unverified calibration of the learned uncertainty in the kinematic tails; both points need to be addressed before the headline amplification factors can be accepted.","major_comments":[{"comment":"The amplification factor G is evaluated using n_train as the number of training events inside the targeted pT window (3 for Z+g, 7 for Z+4g), whereas Eq. (5) and the generative-network convention (Ref. [19]) define n_train as the total training sample size (here 1M). The tail prediction is constrained by the full 1M-event training set, so the in-window count is not the number of constraints supporting the tail estimate. With the standard total-training normalization, the same n_equiv values yield G = 0.06 (Z+g averaging), 0.008 (Z+g KS), 0.56 (Z+4g averaging), and 0.002 (Z+4g KS), which do not support the claimed \"massive statistical amplification\" relative to the training sample. The headline comparison with generative networks in the Abstract and Section 4 is therefore not on equal footing. The authors should either justify the in-window normalization as a distinct local measure (and rename it), or report the total-training normalization and temper the comparative claims. Without this fix, the central quantitative claim of the paper is not supported.","section":"§3.1 Eq. (43); §3.2 Eq. (46); §B Eqs. (47)-(48)"},{"comment":"The numerical values of n_equiv for the averaging measure are defined through sigma_model, which is propagated from the learned sigma_syst. The only calibration evidence is a single pull distribution for Z+4g trained on 10^5 events (Figure 1); no calibration check is shown for the 1M-trained surrogates or restricted to the tail windows [700,800] GeV and [1200,1400] GeV used for the headline amplification numbers. If sigma_syst is mis-calibrated in these tails, n_equiv and G change correspondingly. Moreover, the text does not state clearly whether n_equiv is extracted from Eq. (20) or from the crossing of the bootstrap MI with the statistical scaling in Figures 3 and 6; these two procedures agree only if the learned uncertainty is well calibrated. The authors should provide tail-restricted pull distributions (or an equivalent calibration check) and specify the extraction procedure.","section":"§3, Figs. 1, 3, 6; Eq. (20)"},{"comment":"The Z+3g KS amplification factor is quoted as G = 770/110 = 290, but 770/110 = 7. This internal numerical inconsistency suggests an error in the numerator or denominator and should be corrected; it also raises concerns about the reliability of the other reported factors.","section":"§B.2, Eq. (48)"}],"minor_comments":[{"comment":"The approximation Var(w_bin) ~= bar{I} sum_i w_i^2 is stated to hold when the weight distribution is uncorrelated with the selection of V; this is a strong assumption in tails where the approximation breaks down, as the text itself notes. Please clarify the regime of validity or use the exact expression in Eq. (11) when reporting n_equiv.","section":"§2.1, Eq. (13)"},{"comment":"The caption of Figure 1 should state which process and training size the pull corresponds to, and whether it is evaluated over the full phase space or in the tail; the current caption only says \"Z+4g surrogate trained on 10^5 events,\" which is ambiguous.","section":"§3, Fig. 1"},{"comment":"The meaning of the \"n_train\" curve in Figures 4 and 7 should be defined explicitly (in-window count versus total training size), since the text and the figures use the same symbol for different quantities, contributing to the normalization ambiguity discussed above.","section":"§3.1, Figs. 4 and 7"},{"comment":"The claim that \"the amplitude surrogate amplifies much more than a generative network learning the full phase-space density [19]\" should cite the specific G values from Ref. [19] to make the comparison quantitative and fair, rather than relying on a qualitative statement.","section":"§4, Outlook"},{"comment":"The smoothness prior that \"there are no finer structures than intermediate mass peaks with GeV-scale widths\" is an important assumption for extrapolation into the high-pT tails; it should be revisited in the Outlook and ideally tested with a dedicated resolution scan in the tail region.","section":"§1, Introduction"}],"recommendation":"major_revision","confidential_remarks":"The paper fits the scope of SciPost Physics as a methods paper on machine learning for LHC simulation. The derivation in Section 2 is sound and the empirical validation against reference samples is encouraging, but the headline amplification factors are currently normalized in a way that contradicts Eq. (5) and the generative-network literature, and the tail calibration of the learned uncertainty is not demonstrated. These issues are fixable within the scope of a revision, so I recommend major revision rather than rejection. I would also ask the editor to ensure that a corrected version explicitly reports both normalizations (in-window and total training size) and provides tail-specific calibration checks."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. First, the paper cleanly adapts the amplification framework from generative networks to amplitude surrogates, and the explicit reweighting trick (Eq. 7) is a nice simplification: you never have to generate surrogate events, you reweight truth events by the surrogate/truth amplitude ratio. The math in Sec. 2 is straightforward, and the KS-based differential measure is sensible. Second, the headline G factors are inflated by an inconsistent denominator. Eq. (5) defines G = n_equiv/n_train, and for generative networks n_train is the total training sample size. In Secs. 3.1-3.2 the authors instead divide by the number of training points that happen to fall inside the tail window (3 for Z+g, 7 for Z+4g), despite the surrogate being trained on the full 1M samples. With the standard total-training normalization, G becomes 0.06 and 0.56, respectively—no amplification. The stress-test note has this right. The comparison to generative networks in the abstract and Outlook is therefore not on equal footing.\n\nWhat is genuinely good: the idea of measuring the equivalent sample size of a surrogate in a sparse tail is useful, and the paper demonstrates that the L-GATr surrogate predicts the tail shape accurately with only a handful of in-window training points. That is a meaningful empirical result, even if the amplification factor is not the giant number advertised. The treatment of systematic uncertainty propagation through the reweighting is careful, and the appendices extend the study sensibly to 2g and 3g.\n\nSoft spots: (1) the denominator issue above is major and changes the main quantitative claim; (2) the learned heteroscedastic uncertainty is validated only via one pull distribution for Z+4g (Fig. 1), not separately in the tail bins where n_equiv is extracted—if sigma_syst is off there, all n_equiv values shift; (3) no error bars on n_equiv or G anywhere, despite the in-window counts being 3 and 7, so the numbers are noisy; (4) the target windows look post hoc, chosen where the effect is largest.\n\nWho is it for: people using amplitude surrogates in LHC simulation chains, and those comparing surrogate vs. generative approaches. It deserves a serious referee: the methodology is worth publishing, but the authors need to fix the normalization, report uncertainties, and either drop the 'far outperforms' claim or make the comparison with n_train=1M for both. I would send it to review with that expectation rather than desk reject.","headline":"Useful methodology, inflated headline numbers: the reported G factors divide by in-window training counts instead of total training size, which changes the main quantitative claim.","tokens_in":16430,"tokens_out":2772,"would_cite":true,"duration_ms":24166,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"An amplitude surrogate trained on a handful of exact matrix elements can deliver the statistical power of hundreds of thousands of truth events in the kinematic tail.","keywords":["generative amplification","amplitude surrogate","surrogate Monte Carlo","Lorentz-equivariant transformer","uncertainty calibration","kinematic tails","Z+jets","event reweighting"],"falsifier":"Compute the systematic pull $t_{\\mathrm{syst}}(x) = (A_{\\mathrm{NN}}(x)-A_{\\mathrm{true}}(x))/\\sigma_{\\mathrm{syst}}(x)$ separately inside the tail windows $p_T^Z \\in [700,800]$ GeV and $[1200,1400]$ GeV for a large independent set of matrix-element evaluations; if its variance there differs strongly from one, or if the mean is biased, the equality defining $n_{\\mathrm{equiv}}$ in Eq. (20) no longer holds, and the reported amplification factors are not trustworthy.","tokens_in":15322,"feed_emoji":"⚛️","tokens_out":9840,"duration_ms":73556,"temperature":0.7,"pith_summary":"This paper establishes that a neural amplitude surrogate trained on a small set of exact scattering amplitudes behaves like a much larger Monte Carlo sample: in the high-transverse-momentum tail, where the truth simulation has almost no events, the surrogate's equivalent sample size $n_{\\mathrm{equiv}}$ exceeds its number of training points by orders of magnitude. The paper defines the amplification factor $G = n_{\\mathrm{equiv}} / n_{\\mathrm{train}}$ and measures it with two complementary metrics: an averaging metric based on rate estimates in a kinematic bin and a differential Kolmogorov-Smirnov metric on the per-event log-likelihood ratio. For $Z+g$ production in the window $p_T^Z \\in [700,800]\\,\\mathrm{GeV}$ it reports $G = 20{,}000$ (averaging) and $G = 2{,}670$ (differential); for $Z+4g$ in $[1200,1400]\\,\\mathrm{GeV}$ it reports $G = 80{,}000$ (averaging) and $G = 290$ (differential). The practical stake is that expensive matrix-element evaluations in sparsely populated LHC tails could be replaced by network evaluations without losing statistical power.","feed_headline":"Amplitude surrogates amplify scarce LHC tail samples 20,000-fold","feed_subtitle":"One small training set yields the equivalent of 560,000 truth events in the Z+4g tail.","key_machinery":"The central object is the amplification factor $G = n_{\\mathrm{equiv}} / n_{\\mathrm{train}}$, defined by matching the statistical uncertainty of a true dataset of size $n_{\\mathrm{equiv}}$ to the systematic model uncertainty of an infinitely large surrogate dataset; the equivalent size is read off where the statistical $1/\\sqrt{n_{\\mathrm{eff}}}$ scaling crosses the model-systematics plateau. The machinery has three load-bearing parts: the explicit density-ratio reweighting $w_i = |\\mathcal{M}|^2_{\\mathrm{surr}}/|\\mathcal{M}|^2_{\\mathrm{true}}$, which converts surrogate Monte Carlo into weighted truth events; a calibrated heteroscedastic uncertainty $\\sigma_{\\mathrm{syst}}(x)$ from the loss $\\mathcal{L} = (A_{\\mathrm{NN}}(x)-A_{\\mathrm{true}}(x))^2/(2\\sigma_{\\mathrm{syst}}^2(x)) + \\log\\sigma_{\\mathrm{syst}}(x)$, propagated into a per-weight model variance; and the Kolmogorov-Smirnov statistic on $\\log w$, whose known asymptotic distribution provides the differential amplification measure. These components let the paper avoid generating surrogate events and instead reweight a fixed test sample, making the amplification measurable bin by bin.","core_discovery":"The central claim is that for a smooth scattering amplitude, a surrogate trained on $n_{\\mathrm{train}}$ exact evaluations describes the amplitude with the statistical power of $n_{\\mathrm{equiv}}$ truth events, where $n_{\\mathrm{equiv}}$ is fixed by equating the statistical fluctuation of $n_{\\mathrm{equiv}}$ true samples with the systematic model uncertainty of an infinite surrogate sample, $n_{\\mathrm{equiv}} = \\bar I (1-\\bar I)/\\sigma_{\\mathrm{model}}^2$ for an averaging rate estimate. The paper implements this for gluon-associated $Z$ production by reweighting a truth test sample with the explicit density ratio $w_i = |\\mathcal{M}|^2_{\\mathrm{surr}}(x_i) / |\\mathcal{M}|^2_{\\mathrm{true}}(x_i)$, so no surrogate event generation is needed. A Lorentz-equivariant transformer with a heteroscedastic head supplies a per-event uncertainty $\\sigma_{\\mathrm{syst}}(x)$ that is propagated into the model variance; a second, differential measure uses the Kolmogorov-Smirnov distance between weighted surrogate and truth on the optimal statistic $\\log w$. In the targeted tail bins the paper finds $n_{\\mathrm{equiv}}$ far above $n_{\\mathrm{train}}$, with the largest ratio $n_{\\mathrm{equiv}} = 560{,}000$ from seven training points in the $Z+4g$ averaging metric, and concludes that amplitude surrogates amplify far more than generative networks that must learn the full phase-space density.","pith_inferences":["Inference: the same procedure applied to an amplitude with a narrow resonance inside the tail window would test the smoothness prior; if the learned surrogate cannot resolve the resonance, the model-systematics plateau should rise and $G$ should drop.","Inference: a tail-localized pull distribution for $t_{\\mathrm{syst}}$ would settle whether the reported factors hold, since the paper only shows a global calibration check for $Z+4g$.","Inference: because amplification is measured by reweighting truth events, the quoted $n_{\\mathrm{equiv}}$ is an upper bound for direct generation from the surrogate; generating events and re-running the KS test would give the practical amplification.","Inference: the approximate flatness of $n_{\\mathrm{equiv}}$ across $p_T$ suggests low-$p_T$ training alone captures the amplitude's functional form; a training set restricted to $p_T < 200$ GeV could probe how much tail amplification survives."],"forward_implications":["For tail-sensitive LHC analyses, a surrogate trained on $10^5$–$10^6$ exact amplitude points can replace rate estimates that would otherwise need hundreds of thousands to millions of matrix-element evaluations, because the equivalent sample size in the tail exceeds the training count by three to five orders of magnitude.","Once the amplification curve reaches the model-systematics plateau, additional tail training points buy no further statistical power; the limiting factor becomes the calibrated surrogate uncertainty rather than training statistics.","Because the method reweights existing truth samples by $w_i$, weighted-event analyses can inherit the amplified statistical power without changing the factorization of the simulation chain.","The differential factor is consistently below the averaging factor, so the reported numbers bracket the practical gain: rate-style observables gain more than shape-resolving ones."],"supporting_citations":[{"why":"Defines the averaging and differential amplification measures and the equivalent-sample-size construction that this paper adapts to amplitude surrogates.","marker":"[19]"},{"why":"Introduces the generative amplification effect for trained network samples, the phenomenon the paper transfers to surrogate Monte Carlo.","marker":"[16]"},{"why":"Supplies the Lorentz-equivariant geometric-algebra transformer architecture used to model log-squared matrix elements.","marker":"[29]"},{"why":"Provides the slim L-GATr variant and hyperparameters the paper uses for fast training and evaluation.","marker":"[54]"},{"why":"Establishes the heteroscedastic surrogate training with calibrated uncertainties that yields the $\\sigma_{\\mathrm{syst}}$ used here.","marker":"[31]"},{"why":"Shows for the same processes that learned surrogate uncertainties are calibrated, the check the paper relies on for its systematic error.","marker":"[37]"},{"why":"Generates the exact training and test amplitude datasets for $Z+n g$ through automated matrix-element evaluation.","marker":"[56]"},{"why":"Supports the regression of log-amplitudes and the heteroscedastic loss, and provides the $Z+n$ gluons benchmark process context.","marker":"[26]"}],"fun_headline_variants":["Seven training points beat 560k truth events","Surrogate Monte Carlo: 560k effective events from 7","560k effective events from just 7 truth samples","One tiny training set, 560k-event payoff"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole amplification framework rests on the assumption that the per-event uncertainty $\\sigma_{\\mathrm{syst}}(x)$ learned by the surrogate is a faithful measure of its actual error inside the high-$p_T$ tail bins where the amplification is quoted, since any miscalibration changes $n_{\\mathrm{equiv}}$ and hence $G$ directly.","fun_headline_variants_meta":{"raw":{"variants":["Seven training points beat 560k truth events","Surrogate Monte Carlo: 560k effective events from 7","560k effective events from just 7 truth samples","One tiny training set, 560k-event payoff"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001591,"raw_usage":{"total_tokens":6334,"prompt_tokens":930,"completion_tokens":5404,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":546,"completion_tokens_details":{"reasoning_tokens":5339}},"tokens_in":546,"tokens_out":5404,"duration_ms":32296,"temperature":1.0,"reasoning_tokens":5339,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T14:32:19.381553+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute the systematic pull $t_{\\mathrm{syst}}(x) = (A_{\\mathrm{NN}}(x)-A_{\\mathrm{true}}(x))/\\sigma_{\\mathrm{syst}}(x)$ separately inside the tail windows $p_T^Z \\in [700,800]$ GeV and $[1200,1400]$ GeV for a large independent set of matrix-element evaluations; if its variance there differs strongly from one, or if the mean is biased, the equality defining $n_{\\mathrm{equiv}}$ in Eq. (20) no longer holds, and the reported amplification factors are not trustworthy.","supporting_citations":[],"review_version":1}