Pith. sign in

REVIEW 2 major objections 4 minor 17 references

A multinomial negative-log-likelihood estimator recovers a mixture parameter from toy thrust distributions more reliably than a simple Pearson-type chi-square baseline, with interpolation reducing off-grid bias.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 05:13 UTC pith:IUTOND6Z

load-bearing objection A careful, reproducible toy closure study whose main estimator comparison is solid; the interpolated-RMSE reversal is plausible but not directly tested. the 2 major comments →

arxiv 2607.22282 v1 pith:IUTOND6Z submitted 2026-07-24 hep-ph hep-ex

Statistical validation of template-based parameter extraction from toy thrust distributions for future FCC-ee studies

classification hep-ph hep-ex
keywords template fittingthrustclosure testmultinomial likelihoodstatistical validationmodel mismatchtoy Monte Carloevent shapes
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper argues that a template-based parameter extraction procedure should be validated statistically before it is connected to full physics and detector simulations. To test that logic, the author builds a minimal toy model in which pseudo-data events are either two-jet-like or three-jet-like, with a single parameter p controlling their mixture, and compares three ways to extract p from binned thrust distributions. The central result is that the multinomial negative-log-likelihood estimator beats the uncorrelated normalized chi-square baseline at every pseudo-data sample size, and that linear interpolation between fixed templates reduces discretization bias. The paper also shows that a controlled model mismatch — 10% broader momentum smearing — creates biases that disappear in nominal closure, so self-consistency tests alone are not enough to establish robustness. The goal is a reusable validation layer for future electron-positron event-shape analyses, not a measurement of the strong coupling.

Core claim

On its own terms, the paper establishes that for the task of recovering a mixture parameter from binned thrust-like distributions, the multinomial negative-log-likelihood is a better default than a simple Pearson-type chi-square applied to normalized histograms. With 10,000 pseudo-experiments, the correct-template fraction rises from 0.467±0.005 at 200 events to 0.934±0.002 at 2000 events for the multinomial likelihood, versus 0.422±0.005 and 0.879±0.003 for the chi-square baseline. Linear interpolation between neighboring templates cuts the maximum absolute mean off-grid bias from 0.0114 (chi-square) and 0.0067 (discrete likelihood) to 0.0030. At 2000 events the interpolated likelihood has

What carries the argument

The central object is the two-component mixture density P(τ|p) = (1−p)f2(τ) + p f3(τ) for the thrust complement τ=1−T, discretized into templates at p=0.10, 0.15, …, 0.50 with 5000 events per template. The main estimator is the multinomial negative log-likelihood, −2Σ_i d_i ln T_i(p), which uses binned pseudo-data counts directly rather than normalized contents. A second version linearly interpolates between neighboring templates and is minimized on a fine grid with spacing 0.001. The machinery separates estimator bias, finite-template fluctuations, grid discretization, and model mismatch from physical modeling.

Load-bearing premise

The key premise is that the 5000-event templates are representative enough for interpolation to transfer only the shape evolution, not spurious fluctuations; the paper attributes the interpolated estimator's high-statistics reversal to this finite-template noise but does not vary template size for the interpolated estimator.

What would settle it

Run the template-statistics test (varying template size from 1,000 to 20,000 events) with the interpolated likelihood estimator at N=2000 pseudo-data events. If the interpolated RMSE remains above the discrete-likelihood RMSE even with 20,000-event templates, the finite-template explanation for the reversal is contradicted, and the recommendation to use interpolation with fixed-size templates would need to change.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • For template-based fits to event-shape observables, using multinomial counts is a more reliable default than a normalized chi-square that ignores bin correlations.
  • Linear template interpolation can reduce discretization bias by several factors, but its benefit shrinks when the template sample size is small relative to the pseudo-data sample.
  • Finer histogram binning does not automatically improve recovery; the 10-bin configuration outperformed 20- and 40-bin configurations in the tested setup.
  • Increasing the number of events used to build templates improves recovery, from 0.677 correct-template fraction with 1000 template events to 0.895 with 20,000.
  • A successful closure test under the generating model does not imply robustness: a 10% change in momentum smearing produces parameter-dependent biases that are invisible in nominal closure.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The RMSE reversal at 2000 events suggests a practical design rule: template statistics should be matched to the expected pseudo-data statistics, and an adaptive scheme could switch between interpolated and discrete estimators when pseudo-data uncertainty drops below template uncertainty.
  • The same closure and mismatch protocol could be applied directly to other event-shape observables, such as the C-parameter, and to likelihood-free approaches where interpolated templates act as smooth emulators.
  • The parameter-dependent mismatch biases hint that detector-response uncertainties should be introduced as constrained nuisance parameters and profiled jointly with the parameter of interest, rather than fixed at nominal values.
  • A natural testable extension is to repeat the template-statistics scan with the interpolated estimator, which the paper does not do; if the reversal persists with 20,000-event templates, the finite-template explanation would need revision.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. The paper develops a toy-model benchmark for template-based parameter extraction from thrust-like distributions, targeting future FCC-ee event-shape studies. Events are generated as a two-jet/three-jet mixture with a continuous control parameter p, and templates are built from 5000-event samples on a discrete p-grid. Three estimators are compared in 10,000 pseudo-experiments each: a discrete Pearson-type chi-square on normalized histograms, a discrete multinomial negative-log-likelihood (NLL), and a linearly interpolated NLL. The main empirical claim is that the multinomial NLL consistently outperforms the simple uncorrelated Pearson baseline in correct-template fraction, bias, and RMSE. Interpolation reduces off-grid discretization bias but shows an RMSE reversal at N=2000, which the authors attribute to finite template statistics. The paper also studies binning, template-statistics variations, and a 10% momentum-smearing model-mismatch test, concluding that nominal closure is insufficient for robustness and that no physical alpha_s claim is made.

Significance. If the results hold, the paper provides a transparent, reproducible validation protocol for comparing statistical estimators before applying them to realistic FCC-ee analyses. The strengths are the large number of pseudo-experiments, the clear separation of estimator bias from physical modeling, the honest limitation to a deliberately simple baseline, and the explicit scope restriction against overclaiming alpha_s. The study demonstrates a useful closure-testing framework and quantifies the expected benefit of a multinomial likelihood over a normalized chi-square in a controlled setting. The main weakness is the unvalidated causal explanation for the interpolated-RMSE reversal, which underpins part of the practical guidance in Section 4.

major comments (2)
  1. [§3.1, Figure 3, Table 1 and §3.4, Figure 7] The RMSE reversal for the interpolated NLL at N=2000 is attributed to finite template statistics: 'This deterioration is therefore interpreted as a finite-template effect' and Section 4 uses this to recommend propagating template uncertainties or using larger templates. However, Section 3.4 varies the number of events per template only for the discrete NLL (Figure 7); the interpolated estimator is never tested with larger templates. The causal claim is therefore not validated. If the reversal instead stems from the interpolation procedure itself rather than from template noise, the practical recommendation about when interpolation helps would be unsupported. Please add an interpolated-NLL curve to the template-statistics test, or explicitly reframe the explanation as a hypothesis requiring further study and soften the recommendation in Section 4.
  2. [Introduction/Abstract] The abstract states 'For an input parameter of 0.30, the correct-template fraction increases from 0.467±0.005 for 200 events to 0.934±0.002 for 2000 events' without specifying which estimator these values refer to. In Section 3.1 these numbers correspond to the discrete multinomial NLL. Please qualify this in the abstract to avoid ambiguity, since the whole point of the paper is to compare estimators.
minor comments (4)
  1. [References] Reference [15] lists a DOI '10.1103/243z-g9x8' that appears malformed or is a placeholder. Please verify and correct the DOI or provide a working identifier.
  2. [§2.5, Eq. (8)] The symbol NPE is used in Eq. (8) but is not defined at that point. Please define it explicitly as the number of pseudo-experiments.
  3. [§3.4, Figure 7] The horizontal axis label 'Events per template' is on a logarithmic scale; the caption should state this explicitly, as it is currently only apparent from the tick marks.
  4. [§2.4, Eq. (6)] The interpolated likelihood is minimized on a grid with spacing 0.001. It would be helpful to state whether this is a dense grid search or a continuous optimizer, as this affects the numerical precision of the reported RMSE values.

Circularity Check

0 steps flagged

No significant circularity: the closure-test design is self-contained and claims are limited to estimator comparison.

full rationale

The paper is a controlled toy closure study: pseudo-data are generated from a known input parameter, templates are constructed independently from the same model, and the estimators are scored by how well they recover that known input. This is the intended validation structure, not a circular derivation. The performance claims (correct-template fraction, bias, RMSE) are direct measurements of estimator behavior under known truth, with no parameter fitted to a subset of data and then renamed as a prediction. There are no self-citations and no uniqueness theorem imported from the authors' prior work; the only methodological citations are external closure-testing references. The multinomial NLL and Pearson-type scores in Eqs. (4)-(6) are explicitly defined estimators, and the paper repeatedly limits its claims, stating it 'does not determine the physical strong coupling constant.' The weakest point, attributing the interpolated NLL RMSE reversal at N=2000 to finite-template fluctuations without varying template size for the interpolated estimator, is an unvalidated mechanistic assumption and a correctness risk, not a circular reduction: no equation reduces the conclusion to its input by construction. Therefore no circularity is identified.

Axiom & Free-Parameter Ledger

2 free parameters · 5 axioms · 0 invented entities

The paper's central claims rest on the toy mixture model, the multinomial assumption for binned counts, and the specific choices of pseudocount and smearing widths. No new physical entities are introduced, but the smearing widths are not reported numerically, which slightly weakens reproducibility.

free parameters (2)
  • pseudocount lambda = 0.5
    Pseudocount added to template bin contents (Eq. 3) to stabilize empty bins; chosen by hand; not varied in sensitivity tests.
  • momentum-smearing widths for 2-jet and 3-jet components = not stated in the manuscript; in code
    Define the component densities f2, f3 and set the toy model; chosen by hand, not fitted, but not reported in the text.
axioms (5)
  • domain assumption The toy mixture model P(τ|p) = (1-p) f2(τ) + p f3(τ) with fixed component densities is a meaningful proxy for jet topology mixtures.
    Eq. (1). The entire benchmark rests on this choice; the paper explicitly states it is a toy model and not physical alpha_s.
  • standard math The binned template contents and pseudo-data counts follow a multinomial distribution given fixed event count.
    Used in Eq. (5); assumes fixed event count and independent events.
  • domain assumption The thrust computed by evaluating along the momentum directions of final-state particles is sufficiently accurate for 2- and 3-particle toy configurations.
    Section 2.2 acknowledges this is not a general-purpose thrust algorithm.
  • domain assumption Templates and pseudo-data are generated from the same underlying stochastic model (closure assumption) for the nominal tests.
    Section 2.5; this is intentional closure, and the model-mismatch test relaxes it.
  • domain assumption Linear interpolation between neighboring templates yields a valid continuous probability model.
    Eq. (6); the paper tests this and finds limitations at high statistics.

pith-pipeline@v1.3.0-alltime-deepseek · 6461 in / 11592 out tokens · 89793 ms · 2026-08-01T05:13:01.856244+00:00 · methodology

0 comments
read the original abstract

A statistically reliable parameter-extraction procedure should be validated independently of the physical and detector models to which it will eventually be applied. A template-based inference framework is developed and tested using controlled toy thrust distributions as a methodological benchmark for future FCC-ee event-shape studies. The model contains two-jet-like and three-jet-like event components whose relative contribution is governed by a continuous control parameter. Three extraction strategies are compared: a discrete Pearson-type $\chi^2$ distance, a discrete multinomial negative-log-likelihood estimator, and a likelihood estimator based on linearly interpolated templates. Their performance is quantified with 10,000 pseudo-experiments for samples containing 200--2000 events. Off-grid recovery, histogram-binning variations, template-statistics variations, and a model-mismatch test based on modified momentum smearing are also studied. The multinomial likelihood consistently outperforms the simple uncorrelated Pearson-type baseline used here. For an input parameter of 0.30, the correct-template fraction increases from $0.467\pm0.005$ for 200 events to $0.934\pm0.002$ for 2000 events. Interpolation reduces off-grid discretization effects, with a maximum absolute mean bias of approximately 0.003 in the tested configuration. Model mismatch nevertheless produces measurable biases, demonstrating that nominal closure alone is insufficient to establish robustness. This study does not determine the physical strong coupling constant. It provides a controlled framework for comparing estimators, quantifying statistical resolution, and diagnosing model dependence before application to realistic parton-shower simulations, hadronization models, detector effects, and FCC-ee data.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

17 extracted references · 7 canonical work pages

  1. [1]

    Brandt, C

    S. Brandt, C. Peyrou, R. Sosnowski, and A. Wroblewski, Phys. Lett. 12, 57 (1964). doi:10.1016/0031-9163(64)91176- X

  2. [2]

    Farhi, Phys

    E. Farhi, Phys. Rev. Lett. 39, 1587 (1977). doi:10.1103/PhysRevLett.39.1587

  3. [3]

    Gehrmann-De Ridder, T

    A. Gehrmann-De Ridder, T. Gehrmann, E.W.N. Glover, and G. Heinrich, JHEP 12, 094 (2007). doi:10.1088/1126- 6708/2007/12/094

  4. [4]

    Abbate, M

    R. Abbate, M. Fickinger, A.H. Hoang, V. Mateu, and I.W. Stewart, Phys. Rev. D 83, 074021 (2011). doi:10.1103/PhysRevD.83.074021

  5. [5]

    d’Enterria et al., J

    D. d’Enterria et al., J. Phys. G: Nucl. Part. Phys. 51, 090501 (2024). doi:10.1088/1361-6471/ad1a78

  6. [6]

    Navas et al

    S. Navas et al. (Particle Data Group), Phys. Rev. D 110, 030001 (2024). doi:10.1103/PhysRevD.110.030001

  7. [7]

    Abada et al

    A. Abada et al. (FCC Collaboration), Eur. Phys. J. C 79, 474 (2019). doi:10.1140/epjc/s10052-019-6904-3

  8. [8]

    Abada et al

    A. Abada et al. (FCC Collaboration), Eur. Phys. J. Spec. Top. 228, 261 (2019). doi:10.1140/epjst/e2019-900045-4

  9. [9]

    d’Enterria and P.Z

    D. d’Enterria and P.Z. Skands, eds., High-Precisionα s Measurements from LHC to FCC-ee, Proceedings of the Workshop held at CERN, Geneva, 12–13 October 2015, CERN-PH-TH-2015-299, arXiv:1512.05194 [hep-ph] (2015)

  10. [10]

    Monni and G

    P.F. Monni and G. Zanderighi, Eur. Phys. J. Plus 136, 1162 (2021). doi:10.1140/epjp/s13360-021-02105-4

  11. [11]

    Del Debbio, T

    L. Del Debbio, T. Giani, and M. Wilson, Eur. Phys. J. C 82, 330 (2022). doi:10.1140/epjc/s10052-022-10297-x

  12. [12]

    Deans, Closure testing the NNPDF3.0 methodology, Proceedings of QCD14, pp

    C.S. Deans, Closure testing the NNPDF3.0 methodology, Proceedings of QCD14, pp. 15–18 (2014), arXiv:1409.4283 [hep-ph]

  13. [13]

    Kardos, G

    A. Kardos, G. Somogyi, and A. Verbytskyi, SciPost Phys. Proc. 10, 014 (2022). doi:10.21468/SciPostPhysProc.10.014

  14. [14]

    Physics case for low- √sQCD studies at FCC-ee,

    D. d’Enterria, P.F. Monni, P. Skands, and A. Verbytskyi, “Physics case for low- √sQCD studies at FCC-ee,” CERN-TH-2025-064, arXiv:2503.23855 [hep-ex] (2025). 12

  15. [15]

    Mathew, R

    P. Mathew, R. Aggarwal, and M. Kaur, Phys. Rev. D 113, 116017 (2026). doi:10.1103/243z-g9x8

  16. [16]

    Baron, S

    J. Baron, S. Marzani, and V. Theeuwes, JHEP 08, 105 (2018). doi:10.1007/JHEP08(2018)105

  17. [17]

    Marzani, D

    S. Marzani, D. Reichelt, S. Schumann, G. Soyez, and V. Theeuwes, JHEP 11, 179 (2019). doi:10.1007/JHEP11(2019)179. 13