Pith. sign in

REVIEW 3 major objections 3 minor 72 references

The HDSense score, built only from single-observable Fisher information matrices, ranks K-observable subsets so that the selected set gives near-maximal parameter sensitivity, validated on Lund string hadronization parameters.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-03 05:37 UTC pith:JMM2W6WM

load-bearing objection The empirical HDSense heuristic looks useful for MC tuning, but the paper's claimed theoretical derivation collapses at Eqs. (30) and (34), so it should be peer-reviewed as a heuristic with strong empirical support, not as a derived bound. the 3 major comments →

arxiv 2602.01509 v2 pith:JMM2W6WM submitted 2026-02-02 hep-ph hep-exstat.ME

HDSense: An efficient method for ranking observable sensitivity

classification hep-ph hep-exstat.ME
keywords HDSenseobservable selectionFisher informationsensitivity rankinghadronizationLund string modelcorrelation penaltyprofile likelihood
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper tries to establish that a score computed only from single-observable histograms—the HDSense score, S_HD(X)=Info(X)[1−β P_overlap(X)]—ranks subsets of observables so that the chosen set gives nearly maximal parameter sensitivity. The score sums the traces of per-observable Fisher information matrices and then imposes a penalty for overlap between their parameter sensitivities, absorbing the unknown correlation structure into a single hyperparameter β. This matters because full-likelihood comparison of all observable subsets is computationally prohibitive, especially for hadronization models where correlations are difficult to model. In a simulated e+e−→Z→jets application to five Lund string hadronization parameters, the paper shows the selected subsets are near-optimal when scored against a machine-learned approximation to the full likelihood. The intended payoff is a cheap, widely applicable tool for deciding which measurements deserve precision.

Core claim

On its own terms, the central claim is that the HDSense score—defined as the total trace of single-observable Fisher information matrices times a correlation-redundancy factor—identifies observable subsets whose parameter constraints are nearly as good as those obtained from the full joint likelihood. The paper derives the score by profiling over unknown correlations (the copula) in a Gaussian limit, which yields an approximate lower bound on the trace of the profiled Fisher information matrix, with the unknown inverse correlation matrix compressed into the single parameter β. Validation against machine-learned full-likelihood approximations in the Lund string example shows that selections m

What carries the argument

The load-bearing object is the HDSense score S_HD(X)=Info(X)[1−β P_overlap(X)]. Info(X) is the summed trace of single-observable Fisher information matrices, each computed from a binned histogram via a chain rule and event reweighting; P_overlap(X) is a normalized pairwise sum of the Frobenius inner product between Fisher matrices, which measures how much two observables constrain the same parameter directions. β is a heuristic penalty strength, set to 0.5 divided by the maximum overlap over candidate sets. The attached derivation positions S_HD as an approximate lower bound on the trace of the profiled Fisher information matrix, justifying the ansatz that one scalar correlation penalty suff

Load-bearing premise

The method presumes that a single scalar overlap penalty with a heuristic strength β captures the full correlation structure between observables well enough to rank subsets; the paper's derivation of this approximation uses an inequality that is not generally valid for multivariate Fisher matrices, so the near-optimality is only established empirically in the tested configuration.

What would settle it

Take two observables with known Gaussian joint likelihood and Fisher matrices of rank two that are positively correlated in the directions that matter, so that |cos Φ_ij| > sqrt(cos Φ^F_ij) — the violation of the paper's bound. Compute the exact A-optimal subset (minimizing the trace of the inverse profiled Fisher matrix) and compare with the subset HDSense selects; if HDSense does not select the exact optimal subset in this controlled setting, the general claim of near-optimal ranking is refuted.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • If HDSense is right, experimentalists can decide which observables to measure with high precision—and which to leave out—using only per-observable histograms, without building a full joint likelihood.
  • In the Lund string application, the selected subsets concentrate on multiplicity observables (hadron, charged, baryon, strangeness), suggesting that these measurements carry the most information about flavor-related hadronization parameters.
  • The framework combines measurements from experiments with different statistics and acceptances by adding Fisher information contributions, so lower-statistics observables can still be chosen if they probe otherwise unconstrained parameter directions.
  • Detector efficiencies can be included by reweighting the histogram bin occupancies; the score's structure is unchanged, and rankings remain largely stable.
  • The method is presented as generic: the same construction applies to any parameter estimation problem with many observables whose correlations are unknown, such as effective field theory fits, parton distributions, or astrophysical models.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The theoretical guarantee is weaker than stated: the inequality used to bound the overlap term (|cos Φ_ij| ≤ sqrt(cos Φ^F_ij)) does not hold in general when the single-observable Fisher matrices have rank greater than one, so the 'approximate lower bound' derivation is not a proof; the method's reliability may rest on the particular structure of observables in practice.
  • One testable extension: apply HDSense to synthetic correlated multivariate Gaussians with multi-parameter means and known covariance, then compare its selected subset against the exact A-optimal subset; the paper's toy example only covers perfectly correlated copies of identical observables, which is the most favorable case.
  • The heuristic β = 0.5/max P_overlap can produce negative scores for large K (as the paper notes in an appendix), signaling that the score no longer represents a Fisher information trace; users who want a principled stopping criterion for K may need a different adaptive rule.
  • The greedy remove-one procedure yields a ranking as a byproduct; whether that ranking always reproduces the exhaustive-search selection is not systematically checked in the paper, so an empirical comparison for larger observable pools would be a natural next step.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 3 minor

Summary. This paper introduces HDSense, a score S_HD(X) = Info(X)[1 - β P_overlap(X)] for ranking subsets of observables by their parameter sensitivity. Info(X) sums the traces of single-observable Fisher information matrices, while P_overlap(X) penalizes pairs whose Fisher matrices are aligned in the Frobenius sense. The overlap penalty strength is set heuristically to β = 0.5 / max P_overlap. The authors compute the ingredient Fisher matrices from binned histograms using Pythia event reweighting, and validate the score on a toy Gaussian model and on 15 observables relevant to five Lund string hadronization parameters at the Z pole. Validation is performed against an XGBoost-based approximate full likelihood, and the paper also shows how the framework combines experiments and detector effects. The paper claims that S_HD is derived by profiling over unknown correlations in the Fisher information framework and that it identifies near-optimal observable subsets.

Significance. If the empirical validation is representative, HDSense is a practically useful and computationally cheap tool: it needs only one-dimensional histograms, has a public implementation, and the independent ML-based validation indicates that for K=3,5,7 the selected subsets lie near the optimal region, with bootstrap checks supporting stability. The multi-experiment and detector-extension sections are also valuable. However, the advertised theoretical derivation in Section 3.3 contains an algebraic error and an invalid bounding step; as written, S_HD is not derived as a lower bound on the profiled Fisher information. The mathematical issues are localized, so the paper is salvageable, but the ``derived by profiling'' claim must be either repaired or carefully restated as a heuristic motivation.

major comments (3)
  1. [§3.3, Eq. (30)] The relation between Φ_ij and the Frobenius angle Φ^F_ij is algebraically inverted. From Eqs. (7) and (29), cos²Φ_ij = cosΦ^F_ij / (ξ_i ξ_j). The printed Eq. (30) instead has cos(Φ_ij) = ±√(ξ_i ξ_j cosΦ^F_ij), which for ξ_i ξ_j > 1 would give |cos Φ_ij| ≥ √(cosΦ^F_ij), the reverse of the inequality asserted in Eq. (33). With the corrected identity, the last step of Eq. (33) does follow because ξ_i ≥ 1 for positive-semidefinite I_i. This is a localized but load-bearing algebraic error.
  2. [§3.3, Eqs. (34), (40), (41)] Even after correcting Eq. (30), the lower bound in Eq. (34) does not follow. The partial-correlation bound in Eq. (33) bounds the off-diagonal contribution by a term proportional to √(cosΦ^F_ij) √((ρ^{-1})_ii (ρ^{-1})_jj), not by cosΦ^F_ij (ρ^{-1})_ii (ρ^{-1})_jj. Since √(cosΦ^F_ij) ≥ cosΦ^F_ij for cosΦ^F_ij ∈ [0,1], replacing the former by the latter weakens the bound in the wrong direction. The same problem propagates through Eq. (40) and Eq. (41). The revision should either carry √(cosΦ^F_ij) through the derivation, which changes the score, or explicitly present the overlap penalty as an approximation rather than a derived lower bound on the profiled score.
  3. [§3.3, Eqs. (38)-(40)] The derivation also relies on the unproved dominance assumption in Eq. (38), and the β defined in Eq. (39) depends on the unknown diagonal entries (ρ^{-1})_ii. Consequently, the heuristic β = 0.5 / max P_overlap in Eq. (8) is not a consequence of the profiling calculation. The paper is transparent about this in places, but the Abstract and Section 1 describe the score as derived by profiling. The authors should state precise sufficient conditions under which Eq. (38) holds, or recharacterize S_HD as an empirically motivated score whose form is motivated, but not rigorously derived, by the profiling argument.
minor comments (3)
  1. [Abstract] Typo: ``rank a set observables`` should read ``rank a set of observables``.
  2. [Fig. 3] The axes are labelled Δ(log Tr Î_full^{-1}) and Δ(log Det Î_full^{-1}), but the exact definition of Δ and the normalization used in the text would be easier to follow if repeated in the caption or defined as an equation in the main text.
  3. [Appendix C, after Eq. (68)] The text ``O(n^{2-3})`` is imprecise; the determinant computation is O(n^3) for generic matrices (or O(n^ω) for fast matrix multiplication). Please state the intended complexity.

Circularity Check

1 steps flagged

No significant circularity in the ranking itself; the only self-definitional step is the Section 3.3 'derivation', where β is defined so the bound reproduces S_HD by construction.

specific steps
  1. self definitional [Section 3.3, Eqs. (39)-(40)]
    "Under this assumption, we can take Tr[I(i)](ρ−1)ii as a common factor on the sum, replace it with its lower bound ... and define the effective hyperparameter β ≡ ... M, (39) yielding Spr. ≳ ... ≡ SHD. (40) ... Our ignorance about the correlation structure has been absorbed into the hyperparameter β."

    S_HD was already defined in Eq. (4) with a free parameter β. In Eqs. (39)-(40) the unknown correlation quantities (diagonal entries of ρ^{-1} and M) are absorbed into that same β, so the right-hand side reproduces S_HD by definition rather than independently deriving it. The subsequent choice β = 0.5/max P_overlap (Eq. 8) is then a heuristic inserted after the fact. This makes the advertised 'profiling derivation' partly self-definitional, but it does not force the empirical rankings, which are validated against an external XGBoost full-likelihood approximation.

full rationale

The central empirical claim of HDSense is not circular. β is chosen heuristically (Eq. 8) before comparison with the gold standard, and the validation metric is the Fisher information of an XGBoost classifier score computed from the joint binned observables, not from the single-observable traces that enter S_HD. The toy study and the Lund-string study compare against exhaustive or all-combination searches, so the ranking has independent content. The one circularity-adjacent element is internal to the theoretical justification: Section 3.3 closes by defining β so that the profiled-score bound becomes S_HD, absorbing all unknown correlation information into a free hyperparameter, which is then chosen heuristically. Thus the first-principles derivation does not independently fix the overlap-penalty form. In addition, the inequality |cos Φ_ij| ≤ sqrt(cos Φ_F_ij) used in Eq. (33) is not established as written for ξ_i ξ_j > 1, but that is a mathematical correctness issue rather than a circularity. There are no load-bearing self-citations: Refs. [50,51] are technical reweighting tools, and the validation uses independent XGBoost and Pythia machinery. The paper itself repeatedly flags the heuristic nature of β and the Gaussian assumptions, which further supports a low circularity score.

Axiom & Free-Parameter Ledger

4 free parameters · 5 axioms · 0 invented entities

The method rests on a small number of modeling choices: the Gaussian/copula profiling assumptions, the multinomial binning model, the linear-gradient reweighting approximation, and the heuristic beta penalty. None of these are derived from external benchmarks or data; they are assumptions of the framework. There are no new physical entities.

free parameters (4)
  • beta_0 (overlap penalty strength) = 0.5
    Set by hand in Eq. (8) as a compromise between no penalty and maximal penalty. The paper notes any beta_0 in [0,1] could be used; rankings are sensitive to its value (Fig. 2).
  • beta (adaptive overlap penalty) = 0.5 / max_X P_overlap(X)
    The effective penalty used in S_HD. It varies with K and is not derived from first principles; its heuristic form is central to the ranking.
  • Binning choice B = round(10^-4 N') = ≈1% per-bin statistical uncertainty
    Chosen ad hoc to balance bin statistics and resolution. Fisher information matrices depend on binning, though the method may be robust to this choice.
  • Gradient fit sampling radius and N_theta' = 5%, 150 points
    The linear-gradient extraction in Eq. (49) assumes a 5% radius around the reference point is both reweightable and linear; this is a free modeling choice affecting all Fisher matrices.
axioms (5)
  • standard math Sklar's theorem: any joint distribution can be decomposed into marginals and a copula.
    Invoked in Eq. (20) as the starting point for profiling over unknown correlations.
  • domain assumption Binned event counts follow a multinomial distribution with fixed total N.
    Used in Section 2.2 to derive the Fisher information matrix for binned histograms, Eq. (10)-(12).
  • domain assumption Gaussian approximation for observables and covariance matrix independent of theta.
    Assumed in Section 3.2 to set cross-terms I_theta-eta to zero and to express Fisher matrices via means; the paper acknowledges this is an idealization.
  • ad hoc to paper Moderate-correlation dominance assumption of Eq. (38).
    Needed to absorb the unknown correlation factors into a single beta. The paper says the assumption 'holds when correlations are moderate' but provides no external justification.
  • ad hoc to paper Heuristic beta = 0.5/max P_overlap is a valid summary of unknown correlations.
    The final form of S_HD depends on this choice; Appendix B shows that other fixed values produce negative or non-monotonic scores.

pith-pipeline@v1.3.0-alltime-deepseek · 27770 in / 12797 out tokens · 133355 ms · 2026-08-03T05:37:23.540000+00:00 · methodology

0 comments
read the original abstract

Identifying which observables most effectively constrain model parameters can be computationally prohibitive when considering full likelihoods of many correlated observables. This is especially important for, e.g., hadronization models, where high precision is required to interpret the results of collider experiments. We introduce the High-Dimensional Sensitivity (HDSense) score, a computationally efficient metric for ranking observable sets using only one-dimensional histograms. Derived by profiling over unknown correlations in the Fisher information framework, the score balances total information content against redundancy between observables. We apply HDSense to rank a set observables in terms of their constraining power with respect to five parameters of the Lund string model of hadronization implemented in Pythia using simulated leptonic collider events at the $Z$ pole. Validation against machine-learning--based full-likelihood approximations demonstrates that HDSense successfully identifies near-optimal observable subsets. The framework naturally handles data from multiple experiments with different acceptances and incorporates detector effects. While demonstrated on hadronization models, the methodology applies broadly to generic parameter estimation problems where correlations are unknown or difficult to model.

Figures

Figures reproduced from arXiv: 2602.01509 by Beno\^it Assi, Christian Bierlich, Jure Zupan, Manuel Szewc, Michael K. Wilkinson, Phil Ilten, Rikab Gambhir, Stephen Mrenna, Tony Menzo.

Figure 1
Figure 1. Figure 1: (Left) Covariance matrix Σ corresponding to the toy example outlined in eq. (42), with K = 20. The covariance is in block-diagonal form. (Right) The dependence of the means µi of each observable on each of the 5 parameters θi . consider a set of independent observables, copied several times, inducing perfect correla￾tions between the copies, thus resulting in a block-diagonal likelihood. HDSense should be … view at source ↗
Figure 2
Figure 2. Figure 2: Validation of HDSense selection prescription against approximate full likelihood for K = 3 (top), K = 5 (middle), and K = 7 (bottom) ob￾servables. Left: Dependence of the full Fisher information trace on β for the subset selected by HDSense at each β value. The smallest and largest possible traces, obtained by exploring all possible K-observable subsets, are shown for reference. Right: Trace versus determi… view at source ↗
Figure 3
Figure 3. Figure 3: Comparison of the selected subsets for all K. For K ≤ 7, the HD￾Sense score hovers close to the optimal region, while it remains far from the ‘worst’ possible score for all K. The quality of the approximation degrades for large K where correlations become more important. Here, the ∆ metric is defined as the ratio between the difference in log space of the worst combination and the combination selected by H… view at source ↗
Figure 4
Figure 4. Figure 4: HDSense ranking for all 15 observables, as assessed across boot￾strapped samples (see text). Top: HDSense score (blue) and heuristic β (red) versus number of selected observables K. The maximum HDSense score is high￾lighted with an additional blue circle and a dashed vertical line is drawn at the corresponding K-value. Bottom: Heatmap showing the fractional score reduc￾tion ∆SHD/SHD when each observable is… view at source ↗
Figure 5
Figure 5. Figure 5: HDSense ranking using only the 9 event-level observables. Layout and interpretation as in fig. 4. that the selected subset grows monotonically: the K-observable subset is always contained within the (K + 1)-observable subset. This monotonicity is not guaranteed for fixed β, as demonstrated in section B. Third, the HDSense score exhibits non-monotonic behavior, reaching a maximum around K = 10 in fig. 4. Th… view at source ↗
Figure 6
Figure 6. Figure 6: HDSense ranking for event-level observables when combining two experiments with different statistics and observable coverage. Experiment I: 106 events, no particle ID (cannot measure nbaryon, nstr). Experiment II: 105 events, full particle ID. Top: HDSense score (blue) and heuristic β (red) versus K. Bottom: Observable importance heatmap, with layout as in fig. 4. Despite 10× lower statistics, nbaryon and … view at source ↗
Figure 7
Figure 7. Figure 7: HDSense ranking with estimated detector efficiencies for charged particles as functions of |p| and particle id. Layout as in fig. 4. While overall scores decrease compared to perfect efficiency, the relative ranking of observables remains largely stable. their known sensitivity to flavor parameters. The method also naturally handles multi￾ple experiments with different detector capabilities, balancing stat… view at source ↗
Figure 8
Figure 8. Figure 8: HDSense ranking for all 15 observables with β = 0 (no overlap penalty). Layout as in fig. 4. The score increases monotonically, providing no indication of when additional observables cease providing meaningful constraints due to correlations with already-selected measurements. B Observable rankings with fixed overlap penalty The main text uses an adaptive hyperparameter β(K) that varies with the number of … view at source ↗
Figure 9
Figure 9. Figure 9: HDSense ranking for all 15 observables with fixed β = 0.5. Layout as in fig. 4. The score peaks at K = 2, becoming negative (dark grey shading) from K = 10 onwards. Negative scores indicate the approximation no longer represents a valid Fisher information matrix. Selected subsets (shown with white dashed lines for K = 10 onwards) do not grow monotonically. onwards. If SHD is interpreted as an proxy for A-o… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

72 extracted references · 3 canonical work pages

  1. [1]

    Neyman and E

    J. Neyman and E. S. Pearson,IX. On the problem of the most efficient tests of statis- tical hypotheses, Philosophical Transactions of the Royal Society of London, Series A: Containing Papers of a Mathematical or Physical Character231(694-706), 289 (1933), doi:10.1098/rsta.1933.0009, e-print:https://royalsocietypublishing.org/rsta/article- pdf/231/694-706/...

  2. [2]

    Cranmer, J

    K. Cranmer, J. Pavez and G. Louppe,Approximating Likelihood Ratios with Cali- brated Discriminative Classifiers(2015), e-print:1506.02169

  3. [3]

    Coccaro, M

    A. Coccaro, M. Pierini, L. Silvestrini and R. Torre,The DNNLikelihood: enhanc- ing likelihood distribution with Deep Learning, Eur. Phys. J. C80(7), 664 (2020), doi:10.1140/epjc/s10052-020-8230-1, e-print:1911.03305

  4. [4]

    Rizvi, M

    S. Rizvi, M. Pettee and B. Nachman,Learning likelihood ratios with neural network classifiers, JHEP02, 136 (2024), doi:10.1007/JHEP02(2024)136, e-print:2305.10500

  5. [5]

    Brehmer, K

    J. Brehmer, K. Cranmer, G. Louppe and J. Pavez,Constraining Effective Field Theories with Machine Learning, Phys. Rev. Lett.121(11), 111801 (2018), doi:10.1103/PhysRevLett.121.111801, e-print:1805.00013. 26 SciPost Physics Submission

  6. [6]

    Brehmer, K

    J. Brehmer, K. Cranmer, G. Louppe and J. Pavez,A Guide to Constraining Ef- fective Field Theories with Machine Learning, Phys. Rev. D98(5), 052004 (2018), doi:10.1103/PhysRevD.98.052004, e-print:1805.00020

  7. [7]

    Cranmer, J

    K. Cranmer, J. Brehmer and G. Louppe,The frontier of simulation- based inference, Proceedings of the National Academy of Sci- ences117(48), 30055 (2020), doi:10.1073/pnas.1912789117, e- print:https://www.pnas.org/doi/pdf/10.1073/pnas.1912789117

  8. [8]

    Deistler, J

    M. Deistler, J. Boelts, P. Steinbach, G. Moss, T. Moreau, M. Gloeckler, P. L. C. Ro- drigues, J. Linhart, J. K. Lappalainen, B. K. Miller, P. J. Gon¸ calves, J.-M. Lueckmann et al.,Simulation-based inference: A practical guide(2025), e-print:2508.12939

  9. [9]

    Aadet al.,An implementation of neural simulation-based inference for parameter estimation in ATLAS, Rept

    G. Aadet al.,An implementation of neural simulation-based inference for parameter estimation in ATLAS, Rept. Prog. Phys.88(6), 067801 (2025), doi:10.1088/1361- 6633/add370, e-print:2412.01600

  10. [10]

    Andersson, G

    B. Andersson, G. Gustafson, G. Ingelman and T. Sjostrand,Parton Fragmentation and String Dynamics, Phys. Rept.97, 31 (1983), doi:10.1016/0370-1573(83)90080-7

  11. [11]

    Andersson,The Lund model, Camb

    B. Andersson,The Lund model, Camb. Monogr. Part. Phys. Nucl. Phys. Cosmol.7, 1 (1997)

  12. [12]

    Andersson,The Lund Model, vol

    B. Andersson,The Lund Model, vol. 7, Cambridge University Press, ISBN 978-1-009- 40129-6, 978-1-009-40125-8, 978-1-009-40128-9, 978-0-521-01734-3, 978-0-521-42094- 5, 978-0-511-88149-7, doi:10.1017/9781009401296 (1998)

  13. [13]

    R. D. Field and S. Wolfram,A QCD Model for e+ e- Annihilation, Nucl. Phys. B 213, 65 (1983), doi:10.1016/0550-3213(83)90175-X

  14. [14]

    T. D. Gottschalk,An Improved Description of Hadronization in the QCD Cluster Model fore +e− Annihilation, Nucl. Phys. B239, 349 (1984), doi:10.1016/0550- 3213(84)90253-0

  15. [15]

    Webber,A QCD Model for Jet Fragmentation Including Soft Gluon Interference, Nucl

    B. Webber,A QCD Model for Jet Fragmentation Including Soft Gluon Interference, Nucl. Phys. B238, 492 (1984), doi:10.1016/0550-3213(84)90333-X

  16. [16]

    Hayrapetyanet al.,Combination of Measurements of the Top Quark Mass from Data Collected by the ATLAS and CMS Experiments at s=7 and 8 TeV, Phys

    A. Hayrapetyanet al.,Combination of Measurements of the Top Quark Mass from Data Collected by the ATLAS and CMS Experiments at s=7 and 8 TeV, Phys. Rev. Lett.132(26), 261902 (2024), doi:10.1103/PhysRevLett.132.261902, e- print:2402.08713

  17. [17]

    Kluth,Tests of Quantum Chromo Dynamics at e+ e- Colliders, Rept

    S. Kluth,Tests of Quantum Chromo Dynamics at e+ e- Colliders, Rept. Prog. Phys. 69, 1771 (2006), doi:10.1088/0034-4885/69/6/R04, e-print:hep-ex/0603011

  18. [18]

    Bethke,Experimental tests of asymptotic freedom, Prog

    S. Bethke,Experimental tests of asymptotic freedom, Prog. Part. Nucl. Phys.58, 351 (2007), doi:10.1016/j.ppnp.2006.06.001, e-print:hep-ex/0606035

  19. [19]

    Heisteret al.,Studies of QCD at e+ e- centre-of-mass energies between 91-GeV and 209-GeV, Eur

    A. Heisteret al.,Studies of QCD at e+ e- centre-of-mass energies between 91-GeV and 209-GeV, Eur. Phys. J. C35, 457 (2004), doi:10.1140/epjc/s2004-01891-4

  20. [20]

    Abbiendiet al.,Determination ofalpha s using OPAL hadronic event shapes at√s= 91- 209 GeV and resummed NNLO calculations, Eur

    G. Abbiendiet al.,Determination ofalpha s using OPAL hadronic event shapes at√s= 91- 209 GeV and resummed NNLO calculations, Eur. Phys. J. C71, 1733 (2011), doi:10.1140/epjc/s10052-011-1733-z, e-print:1101.1470. 27 SciPost Physics Submission

  21. [21]

    Abbate, M

    R. Abbate, M. Fickinger, A. H. Hoang, V. Mateu and I. W. Stewart,Thrust atN 3LL with Power Corrections and a Precision Global Fit forα s(mZ), Phys. Rev. D83, 074021 (2011), doi:10.1103/PhysRevD.83.074021, e-print:1006.3080

  22. [22]

    d’Enterriaet al.,The strong coupling constant: state of the art and the decade ahead, J

    D. d’Enterriaet al.,The strong coupling constant: state of the art and the decade ahead, J. Phys. G51(9), 090501 (2024), doi:10.1088/1361-6471/ad1a78, e- print:2203.08271

  23. [23]

    M. A. Benitez, A. H. Hoang, V. Mateu, I. W. Stewart and G. Vita,On determiningα s(mZ) from dijets in e +e− thrust, JHEP07, 249 (2025), doi:10.1007/JHEP07(2025)249, e-print:2412.15164

  24. [24]

    A. H. Hoang, D. W. Kolodrubetz, V. Mateu and I. W. Stewart,Precise determina- tion ofα s from theC-parameter distribution, Phys. Rev. D91(9), 094018 (2015), doi:10.1103/PhysRevD.91.094018, e-print:1501.04111

  25. [25]

    Huston, K

    J. Huston, K. Rabbertz and G. Zanderighi,Quantum Chromodynamics(2023), e- print:2312.14015

  26. [26]

    de Blaset al.,Physics Briefing Book: Input for the 2026 update of the Eu- ropean Strategy for Particle Physics(2025), doi:10.17181/CERN.35CH.2O2P, e- print:2511.03883

    J. de Blaset al.,Physics Briefing Book: Input for the 2026 update of the Eu- ropean Strategy for Particle Physics(2025), doi:10.17181/CERN.35CH.2O2P, e- print:2511.03883

  27. [27]

    Skands, S

    P. Skands, S. Carrazza and J. Rojo,Tuning PYTHIA 8.1: the Monash 2013 Tune, Eur. Phys. J. C74(8), 3024 (2014), doi:10.1140/epjc/s10052-014-3024-y, e-print:1404.5630

  28. [28]

    Buckley, H

    A. Buckley, H. Hoeth, H. Lacker, H. Schulz and J. E. von Seggern,System- atic event generator tuning for the LHC, Eur. Phys. J. C65, 331 (2010), doi:10.1140/epjc/s10052-009-1196-7, e-print:0907.2973

  29. [29]

    Buckley,ATLAS Pythia 8 tunes to 7 TeV data, In6th International Workshop on Multiple Partonic Interactions at the LHC, p

    A. Buckley,ATLAS Pythia 8 tunes to 7 TeV data, In6th International Workshop on Multiple Partonic Interactions at the LHC, p. 29 (2014)

  30. [30]

    Ilten, T

    P. Ilten, T. Menzo, A. Youssef and J. Zupan,Modeling hadronization using machine learning, SciPost Phys.14(3), 027 (2023), doi:10.21468/SciPostPhys.14.3.027, e- print:2203.04983

  31. [31]

    Ghosh, X

    A. Ghosh, X. Ju, B. Nachman and A. Siodmok,Towards a deep learning model for hadronization, Phys. Rev. D106(9), 096020 (2022), doi:10.1103/PhysRevD.106.096020, e-print:2203.12660

  32. [32]

    J. Chan, X. Ju, A. Kania, B. Nachman, V. Sangli and A. Siodmok,Fitting a deep gen- erative hadronization model, JHEP09, 084 (2023), doi:10.1007/JHEP09(2023)084, e-print:2305.17169

  33. [33]

    Bierlich, P

    C. Bierlich, P. Ilten, T. Menzo, S. Mrenna, M. Szewc, M. K. Wilkinson, A. Youssef and J. Zupan,Towards a data-driven model of hadronization using normalizing flows, Sci- Post Phys.17(2), 045 (2024), doi:10.21468/SciPostPhys.17.2.045, e-print:2311.09296

  34. [34]

    J. Chan, X. Ju, A. Kania, B. Nachman, V. Sangli and A. Siodmok,Integrating particle flavor into deep learning models for hadronization, Phys. Rev. D111(11), 116015 (2025), doi:10.1103/hgbg-k7js, e-print:2312.08453

  35. [35]

    M. K. Wilkinson,Simulating Hadronization with Machine Learning, EPJ Web Conf. 295, 09026 (2024), doi:10.1051/epjconf/202429509026. 28 SciPost Physics Submission

  36. [36]

    Bierlich, P

    C. Bierlich, P. Ilten, T. Menzo, S. Mrenna, M. Szewc, M. K. Wilkinson, A. Youssef and J. Zupan,Describing hadronization via histories and observ- ables for Monte-Carlo event reweighting, SciPost Phys.18(2), 054 (2025), doi:10.21468/SciPostPhys.18.2.054, e-print:2410.06342

  37. [37]

    Heller, P

    N. Heller, P. Ilten, T. Menzo, S. Mrenna, B. Nachman, A. Siodmok, M. Szewc and A. Youssef,Rejection Sampling with Autodifferentiation - Case study: Fitting a Hadronization Model(2024), e-print:2411.02194

  38. [38]

    B. Assi, C. Bierlich, P. Ilten, T. Menzo, S. Mrenna, M. Szewc, M. K. Wilkinson, A. Youssef and J. Zupan,Characterizing the hadronization of parton showers using the HOMER method, SciPost Phys.19, 125 (2025), doi:10.21468/SciPostPhys.19.5.125, e-print:2503.05667

  39. [39]

    Butteret al.,Iterative HOMER with uncertainties(2025), e-print:2509.03592

    A. Butteret al.,Iterative HOMER with uncertainties(2025), e-print:2509.03592

  40. [40]

    B. Assi, S. H¨ oche, K. Lee and J. Thaler,QCD Theory Meets Information Theory, Phys. Rev. Lett.135(13), 131901 (2025), doi:10.1103/gf42-qzd9, e-print:2501.17219

  41. [41]

    Bierlichet al.,A comprehensive guide to the physics and usage of PYTHIA 8.3, SciPost Phys

    C. Bierlichet al.,A comprehensive guide to the physics and usage of PYTHIA 8.3, SciPost Phys. Codeb.2022, 8 (2022), doi:10.21468/SciPostPhysCodeb.8, e- print:2203.11601

  42. [42]

    R. A. Fisher,Theory of Statistical Estimation, Proc. Cambridge Phil. Soc.22, 700 (1925)

  43. [43]

    J. J. Ethier, G. Magni, F. Maltoni, L. Mantani, E. R. Nocera, J. Rojo, E. Slade, E. Vryonidou and C. Zhang,Combined SMEFT interpretation of Higgs, diboson, and top quark data from the LHC, JHEP11, 089 (2021), doi:10.1007/JHEP11(2021)089, e-print:2105.00006

  44. [44]

    Aghanimet al.,Planck 2018 results

    N. Aghanimet al.,Planck 2018 results. VI. Cosmological parameters, Astron. Astrophys.641, A6 (2020), doi:10.1051/0004-6361/201833910, [Erratum: As- tron.Astrophys. 652, C4 (2021)], e-print:1807.06209

  45. [45]

    R. D. Ballet al.,The path to proton structure at 1% accuracy, Eur. Phys. J. C82(5), 428 (2022), doi:10.1140/epjc/s10052-022-10328-7, e-print:2109.02653

  46. [46]

    Houet al.,New CTEQ global analysis of quantum chromodynamics with high-precision data from the LHC, Phys

    T.-J. Houet al.,New CTEQ global analysis of quantum chromodynamics with high-precision data from the LHC, Phys. Rev. D103(1), 014013 (2021), doi:10.1103/PhysRevD.103.014013, e-print:1912.10053

  47. [47]

    Ellis, M

    J. Ellis, M. Madigan, K. Mimasu, V. Sanz and T. You,Top, Higgs, Diboson and Electroweak Fit to the Standard Model Effective Field Theory, JHEP04, 279 (2021), doi:10.1007/JHEP04(2021)279, e-print:2012.02779

  48. [48]

    Cram´ er,Mathematical Methods of Statistics, Princeton University Press (1946)

    H. Cram´ er,Mathematical Methods of Statistics, Princeton University Press (1946)

  49. [49]

    C. R. Rao,Information and the accuracy attainable in the estimation of statistical parameters, Bull. Calcutta Math. Soc.37, 81 (1945)

  50. [50]

    Bierlich, P

    C. Bierlich, P. Ilten, T. Menzo, S. Mrenna, M. Szewc, M. K. Wilkinson, A. Youssef and J. Zupan,Reweighting Monte Carlo predictions and auto- mated fragmentation variations in Pythia 8, SciPost Phys.16(5), 134 (2024), doi:10.21468/SciPostPhys.16.5.134, e-print:2308.13459. 29 SciPost Physics Submission

  51. [51]

    B. Assi, C. Bierlich, P. Ilten, T. Menzo, S. Mrenna, M. Szewc, M. K. Wilkinson, A. Youssef and J. Zupan,Post-hoc reweighting of hadron production in the Lund string model, SciPost Phys.19(4), 104 (2025), doi:10.21468/SciPostPhys.19.4.104, e-print:2505.00142

  52. [52]

    Sklar,Fonctions de r´ epartition ` a n dimensions et leurs marges, Publications de l’Institut de Statistique de l’Universit´ e de Paris8, 229 (1959)

    A. Sklar,Fonctions de r´ epartition ` a n dimensions et leurs marges, Publications de l’Institut de Statistique de l’Universit´ e de Paris8, 229 (1959)

  53. [53]

    Chaloner and I

    K. Chaloner and I. Verdinelli,Bayesian experimental design: A review, Statist. Sci. 10(3), 273 (1995), doi:10.1214/ss/1177009939

  54. [54]

    X. Huan, J. Jagalur and Y. Marzouk,Optimal experimental design: Formulations and computations, Acta Numerica33, 715 (2024), doi:10.1017/S0962492924000023, OSTI ID: 2440515. [55]A measurement of energy correlations and a determination ofα s(m2 z)from e+e− annihilations at √s= 91gev, Physics Letters B252(1), 159 (1990), doi:https://doi.org/10.1016/0370-2693...

  55. [56]

    Pedregosa, G

    F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blon- del, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passoset al.,Scikit- learn: Machine learning in Python, Journal of Machine Learning Research12, 2825 (2011)

  56. [57]

    Chen and C

    T. Chen and C. Guestrin,XGBoost, InProceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. ACM, doi:10.1145/2939672.2939785 (2016)

  57. [58]

    Banfi, G

    A. Banfi, G. P. Salam and G. Zanderighi,Resummed event shapes at hadron - hadron colliders, JHEP08, 062 (2004), doi:10.1088/1126-6708/2004/08/062, e-print:hep- ph/0407287

  58. [59]

    Schyns,NEWTAG -\pi , K, p Tagging for Delphi RICHes(1996)

    E. Schyns,NEWTAG -\pi , K, p Tagging for Delphi RICHes(1996)

  59. [60]

    Brandt, C

    S. Brandt, C. Peyrou, R. Sosnowski and A. Wroblewski,The Principal axis of jets. An Attempt to analyze high-energy collisions as two-body processes, Phys. Lett.12, 57 (1964), doi:10.1016/0031-9163(64)91176-X

  60. [61]

    Farhi,A QCD Test for Jets, Phys

    E. Farhi,A QCD Test for Jets, Phys. Rev. Lett.39, 1587 (1977), doi:10.1103/PhysRevLett.39.1587

  61. [62]

    Catani, G

    S. Catani, G. Turnock and B. R. Webber,Jet broadening measures ine +e− annihi- lation, Phys. Lett. B295, 269 (1992), doi:10.1016/0370-2693(92)91565-Q

  62. [63]

    Y. L. Dokshitzer, A. Lucenti, G. Marchesini and G. P. Salam,On the QCD analysis of jet broadening, JHEP01, 011 (1998), doi:10.1088/1126-6708/1998/01/011, e- print:hep-ph/9801324

  63. [64]

    Parisi,Super Inclusive Cross-Sections, Phys

    G. Parisi,Super Inclusive Cross-Sections, Phys. Lett. B74, 65 (1978), doi:10.1016/0370-2693(78)90061-8

  64. [65]

    J. F. Donoghue, F. E. Low and S.-Y. Pi,Tensor Analysis of Hadronic Jets in Quantum Chromodynamics, Phys. Rev. D20, 2759 (1979), doi:10.1103/PhysRevD.20.2759

  65. [66]

    Catani, Y

    S. Catani, Y. L. Dokshitzer, M. Olsson, G. Turnock and B. R. Webber,New clustering algorithm for multi - jet cross-sections in e+ e- annihilation, Phys. Lett. B269, 432 (1991), doi:10.1016/0370-2693(91)90196-W. 30 SciPost Physics Submission

  66. [67]

    Cacciari and G

    M. Cacciari and G. P. Salam,Dispelling theN 3 myth for thek t jet-finder, Phys. Lett. B641, 57 (2006), doi:10.1016/j.physletb.2006.08.037, e-print:hep-ph/0512210

  67. [68]

    Cacciari, G

    M. Cacciari, G. P. Salam and G. Soyez,FastJet User Manual, Eur. Phys. J. C72, 1896 (2012), doi:10.1140/epjc/s10052-012-1896-2, e-print:1111.6097

  68. [69]

    Thaler and K

    J. Thaler and K. Van Tilburg,Identifying Boosted Objects with N-subjettiness, JHEP 03, 015 (2011), doi:10.1007/JHEP03(2011)015, e-print:1011.2268

  69. [70]

    de Florian and M

    D. de Florian and M. Grazzini,The Back-to-back region in e+ e- energy-energy correlation, Nucl. Phys. B704, 387 (2005), doi:10.1016/j.nuclphysb.2004.10.051, e-print:hep-ph/0407241

  70. [71]

    Kiefer,Optimum Experimental Designs, Journal of the Royal Statistical Society: Series B (Methodological)21(2), 272 (1959), doi:10.1111/j.2517-6161.1959.tb00338.x

    J. Kiefer,Optimum Experimental Designs, Journal of the Royal Statistical Society: Series B (Methodological)21(2), 272 (1959), doi:10.1111/j.2517-6161.1959.tb00338.x

  71. [72]

    Hedayat,Study of optimality criteria in design of experiments, Tech

    A. Hedayat,Study of optimality criteria in design of experiments, Tech. Rep. AFOSR- TR-80-0514, Department of Mathematics, University of Illinois at Chicago, Invited paper presented at the International Symposium on Statistics and Related Topics, Carleton University, Ottawa, Canada, May 5–8, 1980 (1980)

  72. [73]

    symmetric design

    K. R. Shah,Optimality criteria for incomplete block designs, Annals of Mathematical Statistics31(3), 791 (1960). A Observable definitions In this appendix, we collect the definitions of the 15 hadronization-sensitive observables used in our numerical studies (see also eq. (48) in the main text). Multiplicities:n had,n ch,n baryon,n str.The hadron multipli...