REVIEW 2 major objections 4 minor 17 references
A multinomial negative-log-likelihood estimator recovers a mixture parameter from toy thrust distributions more reliably than a simple Pearson-type chi-square baseline, with interpolation reducing off-grid bias.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 05:13 UTC pith:IUTOND6Z
load-bearing objection A careful, reproducible toy closure study whose main estimator comparison is solid; the interpolated-RMSE reversal is plausible but not directly tested. the 2 major comments →
Statistical validation of template-based parameter extraction from toy thrust distributions for future FCC-ee studies
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
On its own terms, the paper establishes that for the task of recovering a mixture parameter from binned thrust-like distributions, the multinomial negative-log-likelihood is a better default than a simple Pearson-type chi-square applied to normalized histograms. With 10,000 pseudo-experiments, the correct-template fraction rises from 0.467±0.005 at 200 events to 0.934±0.002 at 2000 events for the multinomial likelihood, versus 0.422±0.005 and 0.879±0.003 for the chi-square baseline. Linear interpolation between neighboring templates cuts the maximum absolute mean off-grid bias from 0.0114 (chi-square) and 0.0067 (discrete likelihood) to 0.0030. At 2000 events the interpolated likelihood has
What carries the argument
The central object is the two-component mixture density P(τ|p) = (1−p)f2(τ) + p f3(τ) for the thrust complement τ=1−T, discretized into templates at p=0.10, 0.15, …, 0.50 with 5000 events per template. The main estimator is the multinomial negative log-likelihood, −2Σ_i d_i ln T_i(p), which uses binned pseudo-data counts directly rather than normalized contents. A second version linearly interpolates between neighboring templates and is minimized on a fine grid with spacing 0.001. The machinery separates estimator bias, finite-template fluctuations, grid discretization, and model mismatch from physical modeling.
Load-bearing premise
The key premise is that the 5000-event templates are representative enough for interpolation to transfer only the shape evolution, not spurious fluctuations; the paper attributes the interpolated estimator's high-statistics reversal to this finite-template noise but does not vary template size for the interpolated estimator.
What would settle it
Run the template-statistics test (varying template size from 1,000 to 20,000 events) with the interpolated likelihood estimator at N=2000 pseudo-data events. If the interpolated RMSE remains above the discrete-likelihood RMSE even with 20,000-event templates, the finite-template explanation for the reversal is contradicted, and the recommendation to use interpolation with fixed-size templates would need to change.
If this is right
- For template-based fits to event-shape observables, using multinomial counts is a more reliable default than a normalized chi-square that ignores bin correlations.
- Linear template interpolation can reduce discretization bias by several factors, but its benefit shrinks when the template sample size is small relative to the pseudo-data sample.
- Finer histogram binning does not automatically improve recovery; the 10-bin configuration outperformed 20- and 40-bin configurations in the tested setup.
- Increasing the number of events used to build templates improves recovery, from 0.677 correct-template fraction with 1000 template events to 0.895 with 20,000.
- A successful closure test under the generating model does not imply robustness: a 10% change in momentum smearing produces parameter-dependent biases that are invisible in nominal closure.
Where Pith is reading between the lines
- The RMSE reversal at 2000 events suggests a practical design rule: template statistics should be matched to the expected pseudo-data statistics, and an adaptive scheme could switch between interpolated and discrete estimators when pseudo-data uncertainty drops below template uncertainty.
- The same closure and mismatch protocol could be applied directly to other event-shape observables, such as the C-parameter, and to likelihood-free approaches where interpolated templates act as smooth emulators.
- The parameter-dependent mismatch biases hint that detector-response uncertainties should be introduced as constrained nuisance parameters and profiled jointly with the parameter of interest, rather than fixed at nominal values.
- A natural testable extension is to repeat the template-statistics scan with the interpolated estimator, which the paper does not do; if the reversal persists with 20,000-event templates, the finite-template explanation would need revision.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper develops a toy-model benchmark for template-based parameter extraction from thrust-like distributions, targeting future FCC-ee event-shape studies. Events are generated as a two-jet/three-jet mixture with a continuous control parameter p, and templates are built from 5000-event samples on a discrete p-grid. Three estimators are compared in 10,000 pseudo-experiments each: a discrete Pearson-type chi-square on normalized histograms, a discrete multinomial negative-log-likelihood (NLL), and a linearly interpolated NLL. The main empirical claim is that the multinomial NLL consistently outperforms the simple uncorrelated Pearson baseline in correct-template fraction, bias, and RMSE. Interpolation reduces off-grid discretization bias but shows an RMSE reversal at N=2000, which the authors attribute to finite template statistics. The paper also studies binning, template-statistics variations, and a 10% momentum-smearing model-mismatch test, concluding that nominal closure is insufficient for robustness and that no physical alpha_s claim is made.
Significance. If the results hold, the paper provides a transparent, reproducible validation protocol for comparing statistical estimators before applying them to realistic FCC-ee analyses. The strengths are the large number of pseudo-experiments, the clear separation of estimator bias from physical modeling, the honest limitation to a deliberately simple baseline, and the explicit scope restriction against overclaiming alpha_s. The study demonstrates a useful closure-testing framework and quantifies the expected benefit of a multinomial likelihood over a normalized chi-square in a controlled setting. The main weakness is the unvalidated causal explanation for the interpolated-RMSE reversal, which underpins part of the practical guidance in Section 4.
major comments (2)
- [§3.1, Figure 3, Table 1 and §3.4, Figure 7] The RMSE reversal for the interpolated NLL at N=2000 is attributed to finite template statistics: 'This deterioration is therefore interpreted as a finite-template effect' and Section 4 uses this to recommend propagating template uncertainties or using larger templates. However, Section 3.4 varies the number of events per template only for the discrete NLL (Figure 7); the interpolated estimator is never tested with larger templates. The causal claim is therefore not validated. If the reversal instead stems from the interpolation procedure itself rather than from template noise, the practical recommendation about when interpolation helps would be unsupported. Please add an interpolated-NLL curve to the template-statistics test, or explicitly reframe the explanation as a hypothesis requiring further study and soften the recommendation in Section 4.
- [Introduction/Abstract] The abstract states 'For an input parameter of 0.30, the correct-template fraction increases from 0.467±0.005 for 200 events to 0.934±0.002 for 2000 events' without specifying which estimator these values refer to. In Section 3.1 these numbers correspond to the discrete multinomial NLL. Please qualify this in the abstract to avoid ambiguity, since the whole point of the paper is to compare estimators.
minor comments (4)
- [References] Reference [15] lists a DOI '10.1103/243z-g9x8' that appears malformed or is a placeholder. Please verify and correct the DOI or provide a working identifier.
- [§2.5, Eq. (8)] The symbol NPE is used in Eq. (8) but is not defined at that point. Please define it explicitly as the number of pseudo-experiments.
- [§3.4, Figure 7] The horizontal axis label 'Events per template' is on a logarithmic scale; the caption should state this explicitly, as it is currently only apparent from the tick marks.
- [§2.4, Eq. (6)] The interpolated likelihood is minimized on a grid with spacing 0.001. It would be helpful to state whether this is a dense grid search or a continuous optimizer, as this affects the numerical precision of the reported RMSE values.
Circularity Check
No significant circularity: the closure-test design is self-contained and claims are limited to estimator comparison.
full rationale
The paper is a controlled toy closure study: pseudo-data are generated from a known input parameter, templates are constructed independently from the same model, and the estimators are scored by how well they recover that known input. This is the intended validation structure, not a circular derivation. The performance claims (correct-template fraction, bias, RMSE) are direct measurements of estimator behavior under known truth, with no parameter fitted to a subset of data and then renamed as a prediction. There are no self-citations and no uniqueness theorem imported from the authors' prior work; the only methodological citations are external closure-testing references. The multinomial NLL and Pearson-type scores in Eqs. (4)-(6) are explicitly defined estimators, and the paper repeatedly limits its claims, stating it 'does not determine the physical strong coupling constant.' The weakest point, attributing the interpolated NLL RMSE reversal at N=2000 to finite-template fluctuations without varying template size for the interpolated estimator, is an unvalidated mechanistic assumption and a correctness risk, not a circular reduction: no equation reduces the conclusion to its input by construction. Therefore no circularity is identified.
Axiom & Free-Parameter Ledger
free parameters (2)
- pseudocount lambda =
0.5
- momentum-smearing widths for 2-jet and 3-jet components =
not stated in the manuscript; in code
axioms (5)
- domain assumption The toy mixture model P(τ|p) = (1-p) f2(τ) + p f3(τ) with fixed component densities is a meaningful proxy for jet topology mixtures.
- standard math The binned template contents and pseudo-data counts follow a multinomial distribution given fixed event count.
- domain assumption The thrust computed by evaluating along the momentum directions of final-state particles is sufficiently accurate for 2- and 3-particle toy configurations.
- domain assumption Templates and pseudo-data are generated from the same underlying stochastic model (closure assumption) for the nominal tests.
- domain assumption Linear interpolation between neighboring templates yields a valid continuous probability model.
read the original abstract
A statistically reliable parameter-extraction procedure should be validated independently of the physical and detector models to which it will eventually be applied. A template-based inference framework is developed and tested using controlled toy thrust distributions as a methodological benchmark for future FCC-ee event-shape studies. The model contains two-jet-like and three-jet-like event components whose relative contribution is governed by a continuous control parameter. Three extraction strategies are compared: a discrete Pearson-type $\chi^2$ distance, a discrete multinomial negative-log-likelihood estimator, and a likelihood estimator based on linearly interpolated templates. Their performance is quantified with 10,000 pseudo-experiments for samples containing 200--2000 events. Off-grid recovery, histogram-binning variations, template-statistics variations, and a model-mismatch test based on modified momentum smearing are also studied. The multinomial likelihood consistently outperforms the simple uncorrelated Pearson-type baseline used here. For an input parameter of 0.30, the correct-template fraction increases from $0.467\pm0.005$ for 200 events to $0.934\pm0.002$ for 2000 events. Interpolation reduces off-grid discretization effects, with a maximum absolute mean bias of approximately 0.003 in the tested configuration. Model mismatch nevertheless produces measurable biases, demonstrating that nominal closure alone is insufficient to establish robustness. This study does not determine the physical strong coupling constant. It provides a controlled framework for comparing estimators, quantifying statistical resolution, and diagnosing model dependence before application to realistic parton-shower simulations, hadronization models, detector effects, and FCC-ee data.
Reference graph
Works this paper leans on
-
[1]
S. Brandt, C. Peyrou, R. Sosnowski, and A. Wroblewski, Phys. Lett. 12, 57 (1964). doi:10.1016/0031-9163(64)91176- X
-
[2]
E. Farhi, Phys. Rev. Lett. 39, 1587 (1977). doi:10.1103/PhysRevLett.39.1587
-
[3]
A. Gehrmann-De Ridder, T. Gehrmann, E.W.N. Glover, and G. Heinrich, JHEP 12, 094 (2007). doi:10.1088/1126- 6708/2007/12/094
doi:10.1088/1126- 2007
-
[4]
R. Abbate, M. Fickinger, A.H. Hoang, V. Mateu, and I.W. Stewart, Phys. Rev. D 83, 074021 (2011). doi:10.1103/PhysRevD.83.074021
-
[5]
D. d’Enterria et al., J. Phys. G: Nucl. Part. Phys. 51, 090501 (2024). doi:10.1088/1361-6471/ad1a78
-
[6]
S. Navas et al. (Particle Data Group), Phys. Rev. D 110, 030001 (2024). doi:10.1103/PhysRevD.110.030001
-
[7]
A. Abada et al. (FCC Collaboration), Eur. Phys. J. C 79, 474 (2019). doi:10.1140/epjc/s10052-019-6904-3
-
[8]
A. Abada et al. (FCC Collaboration), Eur. Phys. J. Spec. Top. 228, 261 (2019). doi:10.1140/epjst/e2019-900045-4
-
[9]
D. d’Enterria and P.Z. Skands, eds., High-Precisionα s Measurements from LHC to FCC-ee, Proceedings of the Workshop held at CERN, Geneva, 12–13 October 2015, CERN-PH-TH-2015-299, arXiv:1512.05194 [hep-ph] (2015)
Pith/arXiv arXiv 2015
-
[10]
P.F. Monni and G. Zanderighi, Eur. Phys. J. Plus 136, 1162 (2021). doi:10.1140/epjp/s13360-021-02105-4
-
[11]
L. Del Debbio, T. Giani, and M. Wilson, Eur. Phys. J. C 82, 330 (2022). doi:10.1140/epjc/s10052-022-10297-x
-
[12]
Deans, Closure testing the NNPDF3.0 methodology, Proceedings of QCD14, pp
C.S. Deans, Closure testing the NNPDF3.0 methodology, Proceedings of QCD14, pp. 15–18 (2014), arXiv:1409.4283 [hep-ph]
Pith/arXiv arXiv 2014
-
[13]
A. Kardos, G. Somogyi, and A. Verbytskyi, SciPost Phys. Proc. 10, 014 (2022). doi:10.21468/SciPostPhysProc.10.014
-
[14]
Physics case for low- √sQCD studies at FCC-ee,
D. d’Enterria, P.F. Monni, P. Skands, and A. Verbytskyi, “Physics case for low- √sQCD studies at FCC-ee,” CERN-TH-2025-064, arXiv:2503.23855 [hep-ex] (2025). 12
Pith/arXiv arXiv 2025
-
[15]
P. Mathew, R. Aggarwal, and M. Kaur, Phys. Rev. D 113, 116017 (2026). doi:10.1103/243z-g9x8
-
[16]
J. Baron, S. Marzani, and V. Theeuwes, JHEP 08, 105 (2018). doi:10.1007/JHEP08(2018)105
-
[17]
S. Marzani, D. Reichelt, S. Schumann, G. Soyez, and V. Theeuwes, JHEP 11, 179 (2019). doi:10.1007/JHEP11(2019)179. 13
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.