Pith. sign in

REVIEW 3 major objections 5 minor 121 references

A Step Toward Interpretability: Smearing the Likelihood

T0 review · 3 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read Smearing the likelihood over an energy metric isolates the physical energy scales a jet classifier exploits, with the minimal smearing scale following an extreme-value-theory power law.

desk verdict A genuinely new interpretability framework with honest limitations; the scale-mixing issue in the metric is the main thing standing between the proposal and the headline claim. read the letter →

arxiv 2501.07643 v2 pith:PGH63SL7 submitted 2025-01-13 hep-ph cs.LGhep-exstat.ML

classification hep-phcs.LGhep-exstat.ML
keywords machinelearninginterpretabilitysmearedlikelihoodSpectralEnergyMover'sDistanceextremevaluetheoryscalinglawsquarkversusgluondiscriminationjetsubstructureIRCsafemetric
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper proposes that interpretability of a machine-learning tagger in particle physics should mean identifying the physical energy scales the tagger exploits, and offers a concrete way to do it: smear the likelihood by averaging over all events within an energy distance $\epsilon$ of each other. On a finite dataset this smearing makes otherwise discrete data continuous, so ratios like the quark-to-gluon likelihood are well defined everywhere. The paper shows that the smallest smearing radius that can be probed with $n$ events is controlled by extreme value theory through $n\,\Sigma(d_n)=1$, which produces the approximate power-law scaling seen in practice, and verifies this on quark versus gluon jet discrimination. The central case-study result is that discrimination power improves smoothly as $\epsilon$ is lowered from 30 to 10 GeV, which the paper reads as evidence that the true likelihood is sensitive to emissions at every scale. The approach matters because it gives a physics-grounded, architecture-independent notion of what a jet-classifier has learned, rather than an opaque list of feature importances.

What carries the argument

The machinery is the smeared likelihood built from a metric ball. The metric is the p=2 Spectral Energy Mover's Distance (SEMD), an infrared- and collinear-safe, closed-form distance on collider events that measures how differently two jets distribute their energy in angle and has units of energy. Smearing replaces each discrete event with the integral of the probability distribution over all events within distance $\epsilon$, which regularizes the finite dataset into a continuous function. The analytic engine is extreme value theory: for $n$ independent draws, the characteristic minimal pairwise distance $d_n$ is defined by $n\,\Sigma(d_n)=1$, and expanding the cumulative distance distribution near $d\to 0$ gives the scaling exponent; for quark versus gluon jets the double-logarithmic Sudakov form yields a Gumbel distribution for the minimum and an effective exponent $\gamma = -\sqrt{\pi/(8\alpha_s(C_A+C_F)\log n)}$, while for massive jets the metric's linear term vanishes so $d_n\propto n^{-1/2+O(\alpha_s)}$.

What would settle it

Measure the mean minimal metric distance between quark and gluon jets for dataset sizes spanning, say, $10^2$ to $10^6$ events and compare with the predicted near-power-law curve $\langle\epsilon_{\min}\rangle \approx 19.4\,n^{-0.12}$ GeV; a statistically significant break from this curve or a plateau at large $n$ would falsify the extreme-value mechanism, as would a discontinuous jump in the ROC curve at any fixed $\epsilon$.

Watch

Extended reading notes

Core claim

The central claim is that the question 'what is the machine using?' can be answered in physical terms by the set of energy scales at which its decisions change. Concretely, the paper defines the smeared likelihood $L(\vec{x}|\epsilon)$ as the ratio of background to signal probability contained in a metric ball of radius $\epsilon$ around a phase-space point $\vec{x}$, using the p=2 Spectral Energy Mover's Distance as the metric. Varying $\epsilon$ from large to small then acts like a Wilsonian renormalization-group step: physics above $\epsilon$ is kept, physics below $\epsilon$ is averaged out, so plateaus or jumps in discrimination power locate dynamically important scales. Because the dataset is finite, $\epsilon$ cannot go to zero; the paper shows that the minimal usable $\epsilon$ for $n$ events obeys $n\,\Sigma(d_n)=1$, where $\Sigma$ is the cumulative distribution of pairwise event distances, and that in the quark-versus-gluon example this gives an approximate power law $\langle\epsilon_{\min}\rangle \approx 19.4\,n^{-0.12}\,\mathrm{GeV}$. Empirically, the smeared likelihood's ROC curves improve smoothly as $\epsilon$ decreases from 30 to 10 GeV, indicating that quark versus gluon discrimination is sensitive to emissions at all scales rather than to one special scale.

Load-bearing premise

The load-bearing premise is that smearing over a metric ball of radius $\epsilon$ removes only physics below that energy scale, a clean cutoff that the paper's own footnote qualifies by noting a 10 GeV SEMD distance can arise from a wide-angle emission with relative transverse momentum near 1 GeV, so the ball can mix different scales.

Editorial extensions

If this is right

  • For quark versus gluon discrimination on 20,000+20,000 jets, the minimum meaningful smearing scale is about 10 GeV, below which some smeared likelihoods evaluate to 0 or infinity and the method loses meaning.
  • The measured scaling $\langle\epsilon_{\min}\rangle \approx 19.4\,n^{-0.12}$ GeV implies that reaching the roughly 1 GeV hadronization scale would require about $10^{10}$ events, far beyond current training-set sizes.
  • Extreme value theory predicts that the scaling exponent is only approximately constant, $\gamma = -\sqrt{\pi/(8\alpha_s(C_A+C_F)\log n)}$, so apparent power laws in finite datasets are approximations rather than exact scaling.
  • For massive jets (W, Z, or Higgs decays), the same reasoning predicts $d_n \propto n^{-1/2+O(\alpha_s)}$, because the squared SEMD is proportional to the squared jet mass and the linear term in the small-distance cumulative distribution vanishes.
  • A smooth, monotonic improvement of the smeared-likelihood ROC curves as $\epsilon$ decreases signals approximate scale invariance; a sudden jump at a particular $\epsilon$ would pinpoint a special physical scale.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A testable extension is to smear a trained network's output rather than the true likelihood; comparing which $\epsilon$ values most change a given architecture's score would rank architectures by the scales they rely on and could expose dependence on non-perturbative or spurious correlations.
  • The relation $n\,\Sigma(d_n)=1$ offers a geometric account of neural scaling laws: if the small-distance cumulative distribution has a power-law tail, the minimal extrapolation distance shrinks as a power of dataset size. Testing whether the exponent tracks the small-distance behavior of $\Sigma$ across different tasks would separate this mechanism from other explanations.
  • Because the SEMD is one choice among many IRC-safe metrics, repeating the smearing analysis with a different metric would test whether the extracted scales are physical or partly a metric artifact; agreement would support the interpretability claim, disagreement would qualify it.
  • The paper's hyper-focused generation suggestion, showering one event many times with different random seeds to populate a small metric ball, could directly measure the local dimensionality or curvature of the dataspace manifold, giving geometric meaning to the minimal smearing scale.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes a definition of interpretability for particle-physics machine learning based on smearing the likelihood over metric balls in event space, using the infrared-and-collinear-safe p=2 Spectral Energy Mover's Distance. The central claims are that varying the smearing radius epsilon resolves the physical energy scales exploited by a classifier, that the minimal smearing scale on a finite dataset follows approximate power-law scaling derived from extreme value theory via n Sigma(d_n)=1, and that a quark-versus-gluon jet case study demonstrates sensitivity of the likelihood to emissions at all scales. The manuscript includes a large-scale simulation study of 20000 quark and 20000 gluon jets, a power-law fit of the mean minimal cross-class distance, an extreme-value-theory analysis of the Sudakov-type distance distribution, a proposed massive-jet scaling argument, and smeared-likelihood ROC curves for epsilon = 10, 15, 20, 25, 30 GeV.

Significance. If the claims were fully established, the smeared-likelihood framework would be a valuable, architecture-independent tool for connecting machine-learning behavior to physical energy scales, and the connection to extreme value theory would give a principled origin for empirical scaling laws. The paper is commendably explicit about its limitations and uses reproducible public tools: the SPECTER code is specified, the event-generation chain is standard, and roughly 8e8 pairwise distances were computed. The quark-versus-gluon study is a genuine, nontrivial proof of concept. However, two load-bearing parts of the argument are not currently supported: the interpretation of SEMD radius epsilon as a clean energy-scale cutoff is contradicted by the metric's own mass-like character, and the massive-jet scaling derivation in Sec. 3.2.3 rests on an integral that the paper itself states does not exist. The headline interpretability conclusion therefore overreaches the evidence presented.

major comments (3)
  1. [Sec. 2 and Sec. 3.3 (Eq. (2.5), Fig. 3, footnote 4)] The central interpretability claim requires that averaging over events within metric distance epsilon removes only physics below the scale epsilon, i.e., that epsilon acts as a Wilsonian cutoff. This is not established for the p=2 SEMD and is in fact undermined by footnote 4: because the SEMD is closely related to the squared jet mass, a 10 GeV metric distance can arise from a wide-angle non-perturbative emission with relative transverse momentum near 1 GeV. Thus a metric ball of radius epsilon mixes emissions with widely different physical scales, and the monotonic ROC improvement shown in Sec. 3.3 does not by itself demonstrate that the true likelihood is 'sensitive to emissions at all scales.' The paper should either use a metric with better scale separation, quantify the mixing (e.g., by studying the composition of metric balls as a function of internal kT), or validate the diagnostic on a problem with a known single physical scale, before drawing the headline conclusion.
  2. [Sec. 3.2.3 (Eqs. (3.18)-(3.24), footnote 6)] The derivation of d_n proportional to n^{-1/2+O(alpha_s)} relies on the claimed quadratic small-distance behavior of the cumulative distance distribution, Eq. (3.22). Footnote 6 states that the leading-order integral involving the derivative of a delta function actually does not exist. Consequently, the relation n Sigma(d_n) proportional to n d_n^2 is not established, and Eq. (3.24) cannot be presented as a result of this manuscript. This is load-bearing because the abstract claims that approximate scaling laws are explicitly demonstrated as a consequence of extreme value theory; the massive-jet example is one of only two concrete theoretical demonstrations. The derivation should be either completed with a proper regularization or removed/reframed as a conjecture for future work.
  3. [Sec. 3.2.2 (Eqs. (3.6)-(3.12), Fig. 1)] There is a mismatch between the theoretical model and the empirical construction. The theory considers the minimum of n i.i.d. draws from the pair-distance distribution, but in Fig. 1 the empirical quantity is, for each event of one class, the minimum over the n events of the other class; the draws are correlated and the total number of cross-class pairs is n^2, not n. In addition, Eq. (3.12) shows that the effective scaling exponent gamma depends on log n, so the predicted relation is not a power law; comparing it to the constant exponent in Eq. (3.5) is an approximation whose accuracy over the plotted range should be quantified. The paper should specify exactly which n enters the extreme-value-theory calculation, state the simplifications connecting the pair model to the empirical minima, and quantify the expected curvature of log <epsilon_min> versus log n.
minor comments (5)
  1. [Eq. (1.4)] The relation n Sigma(d_n)=1 is introduced without derivation; it should be stated explicitly that this is a quantile condition defining a characteristic minimum distance, not the mean minimum distance plotted in Fig. 1.
  2. [Sec. 3.2.1 (Eq. (3.5))] The power-law fit quotes no uncertainties on the exponent or normalization and no goodness-of-fit; given that the theoretical prediction has curvature, a fit range and uncertainty estimate are needed for a meaningful comparison.
  3. [Sec. 3.1] The exclusive kT reclustering used to reduce jets above 100 particles changes the jet state and therefore the SEMD; the paper states this restriction was weak in terms of multiplicity but does not comment on the systematic shift in the distance distribution from the reclustering itself.
  4. [Appendix A (Eq. (A.2))] The displayed formula for d^2(Pi, Pi') appears to have an unbalanced factor or a missing parenthesis in the term involving the minimum of z(1-z) and z'(1-z'); please check and correct the expression.
  5. [Sec. 3.2.2 (Eq. (3.6))] The Sudakov-type distance distribution is central to the extreme-value-theory argument; the paper should state the validity conditions of Eq. (3.6) (double logarithmic accuracy, collinear limit, fixed coupling) in the main text rather than only citing references.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the EVT scaling relation is applied to an independently sourced cumulative distance distribution, and the empirical fit is a comparison, not an input.

full rationale

I find no load-bearing circularity in this paper. The central relation n Sigma(d_n) = 1 is the standard extreme-value-theory definition of the characteristic minimum distance; it is not fitted from data. The cumulative distribution Sigma(d) is taken from published QCD/Sudakov results, including an independent external source (Ref. [104]), rather than from the paper's own simulated dataset or fitted parameters. The empirical power-law fit of Eq. (3.5) is compared with the EVT-based estimate, but the fitted exponent is never inserted back into the theory to produce the claimed scaling law. The choice of the p=2 SEMD is supported by prior, externally reproducible metric constructions and the SPECTER code; citing those works is not circular because the cited metric results do not assume the paper's interpretability claim. The limitations flagged in footnote 4 (metric scale mixing) and footnote 6 (non-existent integral in the massive-jet analysis) are genuine correctness and interpretability concerns, but they are not instances of a derivation reducing to its own inputs. The paper is therefore self-contained in the sense required for a circularity analysis, and no specific circular step can be exhibited.

Assumptions & free parameters 2 free parameters · 5 assumptions · 0 invented entities

The central claims rest on the metric-scale correspondence, the assumed cumulative distance distribution, and the validity of the small-distance expansion. The free parameters are the two numbers in the fitted scaling law. No new physical entities are introduced.

free parameters (2)
  • power-law normalization d0 = 19.4 GeV
    Fit to mean minimum quark-gluon distance versus dataset size in Eq. (3.5).
  • scaling exponent gamma = -0.12
    Fit exponent in Eq. (3.5); not independently predicted by the EVT analysis, which gives a slowly varying gamma around -0.3 for n=20000.
assumptions (5)
  • domain assumption Pairwise event distances are iid draws from a well-defined cumulative distribution Sigma(d).
    Required to apply Fisher-Tippett-Gnedenko extreme value theory in Sec. 3.2.2; stated in the text.
  • ad hoc to paper The SEMD metric has units of energy and is IRC safe, so metric balls correspond to physical energy scales.
    Imposed in Sec. 1 properties 5 and 6; the paper's footnote 4 notes a 10 GeV distance can come from a 1 GeV wide-angle emission, so the scale interpretation is not exact.
  • domain assumption The quark-gluon cumulative distance distribution is the double-log Sudakov form of Eq. (3.6).
    Taken from prior QCD results (refs. 52 and 104), not derived in this paper.
  • ad hoc to paper The small-distance expansion of Sigma(d) is smooth with vanishing linear term and a convergent quadratic term.
    Used for the massive-jet scaling in Sec. 3.2.3; footnote 6 states the quadratic integral does not exist, so this axiom fails in the presented form.
  • domain assumption The leading-order flavor definition of quark and gluon jets from the matrix element is sufficient for the case study.
    The paper admits in Sec. 3.1 that this definition is not theoretically well-defined, but uses it for ubiquity.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A Step Toward Interpretability: Smearing the Likelihood." pith.science (2026). https://pith.science/paper/PGH63SL7

@misc{pith2026250107643,
  author       = {Pith},
  title        = {Pith review of: A Step Toward Interpretability: Smearing the Likelihood},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/PGH63SL7}},
  note         = {Machine review of arXiv:2501.07643}
}
read the original abstract

The problem of interpretability of machine learning architecture in particle physics has no agreed-upon definition, much less any proposed solution. We present a first modest step toward these goals by proposing a definition and corresponding practical method for isolation and identification of relevant physical energy scales exploited by the machine. This is accomplished by smearing or averaging over all input events that lie within a prescribed metric energy distance of one another and correspondingly renders any quantity measured on a finite, discrete dataset continuous over the dataspace. Within this approach, we are able to explicitly demonstrate that (approximate) scaling laws are a consequence of extreme value theory applied to analysis of the distribution of the irreducible minimal distance over which a machine must extrapolate given a finite dataset. As an example, we study quark versus gluon jet identification, construct the smeared likelihood, and show that discrimination power steadily increases as resolution decreases, indicating that the true likelihood for the problem is sensitive to emissions at all scales.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

121 extracted references · 20 canonical work pages

  1. [1]

    A. J. Larkoski, I. Moult, and B. Nachman, Jet Substructure at the Large Hadron Collider: A Review of Recent Advances in Theory and Machine Learning , Phys. Rept. 841 (2020) 1–63, [arXiv:1709.04464]

  2. [2]

    Kogler et al., Jet Substructure at the Large Hadron Collider: Experimental Review , Rev

    R. Kogler et al., Jet Substructure at the Large Hadron Collider: Experimental Review , Rev. Mod. Phys. 91 (2019), no. 4 045003, [ arXiv:1803.06991]

  3. [3]

    Guest, K

    D. Guest, K. Cranmer, and D. Whiteson, Deep Learning and its Application to LHC Physics , Ann. Rev. Nucl. Part. Sci. 68 (2018) 161–181, [ arXiv:1806.11484]

  4. [4]

    Albertsson et al., Machine Learning in High Energy Physics Community White Paper , J

    K. Albertsson et al., Machine Learning in High Energy Physics Community White Paper , J. Phys. Conf. Ser. 1085 (2018), no. 2 022008, [ arXiv:1807.02876]

  5. [5]

    Radovic, M

    A. Radovic, M. Williams, D. Rousseau, M. Kagan, D. Bonacorsi, A. Himmel, A. Aurisano, K. Terao, and T. Wongjirad, Machine learning at the energy and intensity frontiers of particle physics, Nature 560 (2018), no. 7716 41–48

  6. [6]

    Carleo, I

    G. Carleo, I. Cirac, K. Cranmer, L. Daudet, M. Schuld, N. Tishby, L. Vogt-Maranto, and L. Zdeborov´ a,Machine learning and the physical sciences , Rev. Mod. Phys. 91 (2019), no. 4 045002, [arXiv:1903.10563]

  7. [7]

    Bourilkov, Machine and Deep Learning Applications in Particle Physics , Int

    D. Bourilkov, Machine and Deep Learning Applications in Particle Physics , Int. J. Mod. Phys. A 34 (2020), no. 35 1930019, [ arXiv:1912.08245]

  8. [8]

    M. D. Schwartz, Modern Machine Learning and Particle Physics , arXiv:2103.12226

Show all 121 references
  1. [9]

    Karagiorgi, G

    G. Karagiorgi, G. Kasieczka, S. Kravitz, B. Nachman, and D. Shih, Machine Learning in the Search for New Fundamental Physics , arXiv:2112.03769

  2. [10]

    Boehnlein et al., Colloquium: Machine learning in nuclear physics , Rev

    A. Boehnlein et al., Colloquium: Machine learning in nuclear physics , Rev. Mod. Phys. 94 (2022), no. 3 031003, [ arXiv:2112.02309]

  3. [11]

    Shanahan et al., Snowmass 2021 Computational Frontier CompF03 Topical Group Report: Machine Learning, arXiv:2209.07559

    P. Shanahan et al., Snowmass 2021 Computational Frontier CompF03 Topical Group Report: Machine Learning, arXiv:2209.07559. – 19 –

  4. [12]

    Plehn, A

    T. Plehn, A. Butter, B. Dillon, T. Heimel, C. Krause, and R. Winterhalder, Modern Machine Learning for LHC Physicists , arXiv:2211.01421

  5. [13]

    Nachman et al., Jets and Jet Substructure at Future Colliders , Front

    B. Nachman et al., Jets and Jet Substructure at Future Colliders , Front. in Phys. 10 (2022) 897719, [arXiv:2203.07462]

  6. [14]

    DeZoort, P

    G. DeZoort, P. W. Battaglia, C. Biscarat, and J.-R. Vlimant, Graph neural networks at the Large Hadron Collider , Nature Rev. Phys. 5 (2023), no. 5 281–303

  7. [15]

    K. Zhou, L. Wang, L.-G. Pang, and S. Shi, Exploring QCD matter in extreme conditions with Machine Learning, Prog. Part. Nucl. Phys. 135 (2024) 104084, [ arXiv:2303.15136]

  8. [16]

    Belis, P

    V. Belis, P. Odagiu, and T. K. Aarrestad, Machine learning for anomaly detection in particle physics, Rev. Phys. 12 (2024) 100091, [ arXiv:2312.14190]

  9. [17]

    Mondal and L

    S. Mondal and L. Mastrolorenzo, Machine Learning in High Energy Physics: A review of heavy-flavor jet tagging at the LHC , arXiv:2404.01071

  10. [18]

    Feickert and B

    M. Feickert and B. Nachman, A Living Review of Machine Learning for Particle Physics , arXiv:2102.02770

  11. [19]

    A. J. Larkoski, QCD masterclass lectures on jet physics and machine learning , Eur. Phys. J. C 84 (2024), no. 10 1117, [ arXiv:2407.04897]

  12. [20]

    Halverson, TASI Lectures on Physics for Machine Learning , arXiv:2408.00082

    J. Halverson, TASI Lectures on Physics for Machine Learning , arXiv:2408.00082

  13. [21]

    Rudin, Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead, Nature machine intelligence 1 (2019), no

    C. Rudin, Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead, Nature machine intelligence 1 (2019), no. 5 206–215

  14. [22]

    Molnar, Interpretable machine learning

    C. Molnar, Interpretable machine learning . Lulu. com, 2020

  15. [23]

    Grojean, A

    C. Grojean, A. Paul, Z. Qian, and I. Str¨ umke, Lessons on interpretable machine learning from particle physics, Nature Rev. Phys. 4 (2022), no. 5 284–286, [ arXiv:2203.08021]

  16. [24]

    Interpretable machine learning for particle physics

    “Interpretable machine learning for particle physics.” https://indico.cern.ch/event/1407421/contributions/6055393/attachments/ 2923026/5130688/jthaler_2024_09_10_InterpretableML_PHYSTAT.pdf. Accessed: 2024-12-24

  17. [25]

    10,000 einsteins: Ai and the future of theoretical physics

    “10,000 einsteins: Ai and the future of theoretical physics.” https://scholar.harvard.edu/ sites/scholar.harvard.edu/files/iaifi-workshop-schwartz.pdf. Accessed: 2024-12-24

  18. [26]

    L. S. Shapley, Notes on the n-person game-ii: The value of an n-person game , 1951

  19. [27]

    A. E. Roth, The Shapley value: essays in honor of Lloyd S. Shapley . Cambridge University Press, 1988

  20. [28]

    Chang, T

    S. Chang, T. Cohen, and B. Ostdiek, What is the Machine Learning? , Phys. Rev. D 97 (2018), no. 5 056009, [ arXiv:1709.10106]

  21. [29]

    Roxlo and M

    T. Roxlo and M. Reece, Opening the black box of neural nets: case studies in stop/top discrimination, arXiv:1804.09278

  22. [30]

    Faucett, J

    T. Faucett, J. Thaler, and D. Whiteson, Mapping Machine-Learned Physics into a Human-Readable Space, Phys. Rev. D 103 (2021), no. 3 036020, [ arXiv:2010.11998]

  23. [31]

    Grojean, A

    C. Grojean, A. Paul, and Z. Qian, Resurrecting bbh with kinematic shapes , JHEP 04 (2021) 139, [arXiv:2011.13945]. – 20 –

  24. [32]

    R. Das, G. Kasieczka, and D. Shih, Feature selection with distance correlation, Phys. Rev. D 109 (2024), no. 5 054009, [ arXiv:2212.00046]

  25. [33]

    Bhattacherjee, C

    B. Bhattacherjee, C. Bose, A. Chakraborty, and R. Sengupta, Boosted top tagging and its interpretation using Shapley values , Eur. Phys. J. Plus 139 (2024), no. 12 1131, [arXiv:2212.11606]

  26. [34]

    J. M. Munoz, I. Batatia, C. Ortner, and F. Romeo, Retrieval of Boost Invariant Symbolic Observables via Feature Importance, arXiv:2306.13496

  27. [35]

    Chowdhury, A

    S. Chowdhury, A. Chakraborty, and S. Dutta, Boosted Top Tagging through Flavour-violating interactions at the LHC , arXiv:2310.10763

  28. [36]

    Englert, Improved Precision in V h(→ b¯b) via Boosted Decision Trees, arXiv:2407.21239

    P. Englert, Improved Precision in V h(→ b¯b) via Boosted Decision Trees, arXiv:2407.21239

  29. [37]

    C. Bose, A. Chakraborty, S. Chowdhury, and S. Dutta, Interplay of traditional methods and machine learning algorithms for tagging boosted objects , Eur. Phys. J. ST 233 (2024), no. 15-16 2531–2558, [ arXiv:2408.01138]

  30. [38]

    Neyman and E

    J. Neyman and E. S. Pearson, On the Problem of the Most Efficient Tests of Statistical Hypotheses, Phil. Trans. Roy. Soc. Lond. A 231 (1933), no. 694-706 289–337

  31. [39]

    K. G. Wilson, The renormalization group and critical phenomena , Rev. Mod. Phys. 55 (1983) 583–600

  32. [40]

    Kinoshita, Mass singularities of Feynman amplitudes , J

    T. Kinoshita, Mass singularities of Feynman amplitudes , J. Math. Phys. 3 (1962) 650–677

  33. [41]

    T. D. Lee and M. Nauenberg, Degenerate Systems and Mass Singularities , Phys. Rev. 133 (1964) B1549–B1562

  34. [42]

    R. K. Ellis, W. J. Stirling, and B. R. Webber, QCD and collider physics , vol. 8. Cambridge University Press, 2, 2011

  35. [43]

    P. T. Komiske, E. M. Metodiev, and J. Thaler, Metric Space of Collider Events , Phys. Rev. Lett. 123 (2019), no. 4 041801, [ arXiv:1902.02346]

  36. [44]

    Mullin, S

    A. Mullin, S. Nicholls, H. Pacey, M. Parker, M. White, and S. Williams, Does SUSY have friends? A new approach for LHC event analysis , JHEP 02 (2021) 160, [ arXiv:1912.10625]

  37. [45]

    Crispim Rom˜ ao, N

    M. Crispim Rom˜ ao, N. F. Castro, J. G. Milhano, R. Pedro, and T. Vale, Use of a generalized energy Mover’s distance in the search for rare phenomena at colliders , Eur. Phys. J. C 81 (2021), no. 2 192, [ arXiv:2004.09360]

  38. [46]

    T. Cai, J. Cheng, N. Craig, and K. Craig, Linearized optimal transport for collider events , Phys. Rev. D 102 (2020), no. 11 116019, [ arXiv:2008.08604]

  39. [47]

    A. J. Larkoski and T. Melia, Covariantizing phase space , Phys. Rev. D 102 (2020), no. 9 094014, [arXiv:2008.06508]

  40. [48]

    S. Tsan, R. Kansal, A. Aportela, D. Diaz, J. Duarte, S. Krishna, F. Mokhtar, J.-R. Vlimant, and M. Pierini, Particle Graph Autoencoders and Differentiable, Learned Energy Mover’s Distance, in 35th Conference on Neural Information Processing Systems , 11, 2021. arXiv:2111.12849

  41. [49]

    T. Cai, J. Cheng, K. Craig, and N. Craig, Which metric on the space of collider events? , Phys. Rev. D 105 (2022), no. 7 076003, [ arXiv:2111.03670]. – 21 –

  42. [50]

    Kitouni, N

    O. Kitouni, N. Nolte, and M. Williams, Finding NEEMo: Geometric Fitting using Neural Estimation of the Energy Mover’s Distance , arXiv:2209.15624

  43. [51]

    Alipour-Fard, P

    S. Alipour-Fard, P. T. Komiske, E. M. Metodiev, and J. Thaler, Pileup and Infrared Radiation Annihilation (PIRANHA): a paradigm for continuous jet grooming , JHEP 09 (2023) 157, [arXiv:2305.00989]

  44. [52]

    A. J. Larkoski and J. Thaler, A spectral metric for collider geometry , JHEP 08 (2023) 107, [arXiv:2305.03751]

  45. [53]

    Davis, T

    A. Davis, T. Menzo, A. Youssef, and J. Zupan, Earth mover’s distance as a measure of CP violation, JHEP 06 (2023) 098, [ arXiv:2301.13211]

  46. [54]

    D. Ba, A. S. Dogra, R. Gambhir, A. Tasissa, and J. Thaler, SHAPER: can you hear the shape of a jet? , JHEP 06 (2023) 195, [ arXiv:2302.12266]

  47. [55]

    Craig, J

    N. Craig, J. N. Howard, and H. Li, Exploring Optimal Transport for Event-Level Anomaly Detection at the Large Hadron Collider , arXiv:2401.15542

  48. [56]

    T. Cai, J. Cheng, N. Craig, G. Koszegi, and A. J. Larkoski, The phase space distance between collider events , JHEP 09 (2024) 054, [ arXiv:2405.16698]

  49. [57]

    Gambhir, A

    R. Gambhir, A. J. Larkoski, and J. Thaler, SPECTER: efficient evaluation of the spectral EMD, JHEP 12 (2025) 219, [ arXiv:2410.05379]

  50. [58]

    Datta and A

    K. Datta and A. Larkoski, How Much Information is in a Jet? , JHEP 06 (2017) 073, [arXiv:1704.08249]

  51. [59]

    A. J. Larkoski and E. M. Metodiev, A Theory of Quark vs. Gluon Discrimination , JHEP 10 (2019) 014, [ arXiv:1906.01639]

  52. [60]

    Kasieczka, S

    G. Kasieczka, S. Marzani, G. Soyez, and G. Stagnitto, Towards Machine Learning Analytics for Jet Substructure , JHEP 09 (2020) 195, [ arXiv:2007.04319]

  53. [61]

    Rosenblat, Remarks on some nonparametric estimates of a density function , Ann

    M. Rosenblat, Remarks on some nonparametric estimates of a density function , Ann. Math. Stat 27 (1956) 832–837

  54. [62]

    Parzen, On estimation of a probability density function and mode , The annals of mathematical statistics 33 (1962), no

    E. Parzen, On estimation of a probability density function and mode , The annals of mathematical statistics 33 (1962), no. 3 1065–1076

  55. [63]

    Ahmad and G

    S. Ahmad and G. Tesauro, Scaling and generalization in neural networks: a case study , Advances in neural information processing systems 1 (1988)

  56. [64]

    Cohn and G

    D. Cohn and G. Tesauro, Can neural networks do better than the vapnik-chervonenkis bounds?, Advances in Neural Information Processing Systems 3 (1990)

  57. [65]

    Hestness, S

    J. Hestness, S. Narang, N. Ardalani, G. Diamos, H. Jun, H. Kianinejad, M. M. A. Patwary, Y. Yang, and Y. Zhou, Deep learning scaling is predictable, empirically , arXiv preprint arXiv:1712.00409 (2017)

  58. [66]

    Kaplan, S

    J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei, Scaling laws for neural language models , arXiv preprint arXiv:2001.08361 (2020)

  59. [67]

    J. S. Rosenfeld, A. Rosenfeld, Y. Belinkov, and N. Shavit, A constructive prediction of the generalization error across scales , arXiv preprint arXiv:1909.12673 (2019). – 22 –

  60. [68]

    Henighan, J

    T. Henighan, J. Kaplan, M. Katz, M. Chen, C. Hesse, J. Jackson, H. Jun, T. B. Brown, P. Dhariwal, S. Gray, et al., Scaling laws for autoregressive generative modeling , arXiv preprint arXiv:2010.14701 (2020)

  61. [69]

    J. S. Rosenfeld, J. Frankle, M. Carbin, and N. Shavit, On the predictability of pruning across scales, in International Conference on Machine Learning , pp. 9075–9083, PMLR, 2021

  62. [70]

    Fr´ echet,Sur la loi de probabilit´ e de l’´ ecart maximum, Ann

    M. Fr´ echet,Sur la loi de probabilit´ e de l’´ ecart maximum, Ann. de la Soc. Polonaise de Math. (1927)

  63. [71]

    R. A. Fisher and L. H. C. Tippett, Limiting forms of the frequency distribution of the largest or smallest member of a sample , in Mathematical proceedings of the Cambridge philosophical society, vol. 24, pp. 180–190, Cambridge University Press, 1928

  64. [72]

    Von Mises, La distribution de la plus grande de n valuers , Rev

    R. Von Mises, La distribution de la plus grande de n valuers , Rev. math. Union interbalcanique 1 (1936) 141–160

  65. [73]

    Gnedenko, Sur la distribution limite du terme maximum d’une serie aleatoire , Annals of mathematics 44 (1943), no

    B. Gnedenko, Sur la distribution limite du terme maximum d’une serie aleatoire , Annals of mathematics 44 (1943), no. 3 423–453

  66. [74]

    Geuskens, N

    J. Geuskens, N. Gite, M. Kr¨ amer, V. Mikuni, A. M¨ uck, B. Nachman, and H. Reyes-Gonz´ alez, The Fundamental Limit of Jet Tagging , 11, 2024. arXiv:2411.02628

  67. [75]

    Fukushima, Visual feature extraction by a multilayered network of analog threshold elements, IEEE Transactions on Systems Science and Cybernetics 5 (1969), no

    K. Fukushima, Visual feature extraction by a multilayered network of analog threshold elements, IEEE Transactions on Systems Science and Cybernetics 5 (1969), no. 4 322–333

  68. [76]

    L. M. Jones, Tests for Determining the Parton Ancestor of a Hadron Jet , Phys. Rev. D 39 (1989) 2550

  69. [77]

    Fodor, How to See the Differences Between Quark and Gluon Jets , Phys

    Z. Fodor, How to See the Differences Between Quark and Gluon Jets , Phys. Rev. D 41 (1990) 1726

  70. [78]

    Lonnblad, C

    L. Lonnblad, C. Peterson, and T. Rognvaldsson, Finding Gluon Jets With a Neural Trigger , Phys. Rev. Lett. 65 (1990) 1321–1324

  71. [79]

    Lonnblad, C

    L. Lonnblad, C. Peterson, and T. Rognvaldsson, Using neural networks to identify jets , Nucl. Phys. B 349 (1991) 675–702

  72. [80]

    Csabai, F

    I. Csabai, F. Czako, and Z. Fodor, Quark and gluon jet separation using neural networks , Phys. Rev. D 44 (1991) 1905–1908

  73. [81]

    Jones, TOWARDS A SYSTEMATIC JET CLASSIFICATION , Phys

    L. Jones, TOWARDS A SYSTEMATIC JET CLASSIFICATION , Phys. Rev. D 42 (1990) 811–814

  74. [82]

    Pumplin, How to tell quark jets from gluon jets , Phys

    J. Pumplin, How to tell quark jets from gluon jets , Phys. Rev. D 44 (1991) 2025–2032

  75. [83]

    OP ALCollaboration, P. D. Acton et al., A Study of differences between quark and gluon jets using vertex tagging of quark jets , Z. Phys. C 58 (1993) 387–404

  76. [84]

    Private communication

    B. Nachman, “Private communication.”

  77. [85]

    C. Frye, A. J. Larkoski, J. Thaler, and K. Zhou, Casimir Meets Poisson: Improved Quark/Gluon Discrimination with Counting Observables , JHEP 09 (2017) 083, [arXiv:1704.06266]

  78. [86]

    Bright-Thonney, I

    S. Bright-Thonney, I. Moult, B. Nachman, and S. Prestel, Systematic quark/gluon identification with ratios of likelihoods , JHEP 12 (2022) 021, [ arXiv:2207.12411]. – 23 –

  79. [87]

    Alwall, R

    J. Alwall, R. Frederix, S. Frixione, V. Hirschi, F. Maltoni, O. Mattelaer, H. S. Shao, T. Stelzer, P. Torrielli, and M. Zaro, The automated computation of tree-level and next-to-leading order differential cross sections, and their matching to parton shower simulations, JHEP 07...

  80. [88]

    Bierlich et al., A comprehensive guide to the physics and usage of PYTHIA 8.3 , SciPost Phys

    C. Bierlich et al., A comprehensive guide to the physics and usage of PYTHIA 8.3 , SciPost Phys. Codeb. 2022 (2022) 8, [ arXiv:2203.11601]

  81. [89]

    Cacciari, G

    M. Cacciari, G. P. Salam, and G. Soyez, The anti- kt jet clustering algorithm , JHEP 04 (2008) 063, [arXiv:0802.1189]

  82. [90]

    Cacciari, G

    M. Cacciari, G. P. Salam, and G. Soyez, FastJet User Manual , Eur. Phys. J. C 72 (2012) 1896, [arXiv:1111.6097]

  83. [91]

    Catani, Y

    S. Catani, Y. L. Dokshitzer, M. H. Seymour, and B. R. Webber, Longitudinally invariant Kt clustering algorithms for hadron hadron collisions , Nucl. Phys. B 406 (1993) 187–224

  84. [92]

    S. D. Ellis and D. E. Soper, Successive combination jet algorithm for hadron collisions , Phys. Rev. D 48 (1993) 3160–3166, [ hep-ph/9305266]

  85. [93]

    P. Gras, S. H¨ oche, D. Kar, A. Larkoski, L. L¨ onnblad, S. Pl¨ atzer, A. Si´ odmok, P. Skands, G. Soyez, and J. Thaler, Systematics of quark/gluon tagging , JHEP 07 (2017) 091, [arXiv:1704.03878]

  86. [94]

    Banfi, G

    A. Banfi, G. P. Salam, and G. Zanderighi, Infrared safe definition of jet flavor , Eur. Phys. J. C 47 (2006) 113–124, [ hep-ph/0601139]

  87. [95]

    Caletti, A

    S. Caletti, A. J. Larkoski, S. Marzani, and D. Reichelt, Practical jet flavour through NNLO , Eur. Phys. J. C 82 (2022), no. 7 632, [ arXiv:2205.01109]

  88. [96]

    Caletti, A

    S. Caletti, A. J. Larkoski, S. Marzani, and D. Reichelt, A fragmentation approach to jet flavor , JHEP 10 (2022) 158, [ arXiv:2205.01117]

  89. [97]

    Czakon, A

    M. Czakon, A. Mitov, and R. Poncelet, Infrared-safe flavoured anti-kT jets, JHEP 04 (2023) 138, [arXiv:2205.11879]

  90. [98]

    Gauld, A

    R. Gauld, A. Huss, and G. Stagnitto, Flavor Identification of Reconstructed Hadronic Jets , Phys. Rev. Lett. 130 (2023), no. 16 161901, [ arXiv:2208.11138]

  91. [99]

    Caola, R

    F. Caola, R. Grabarczyk, M. L. Hutt, G. P. Salam, L. Scyboz, and J. Thaler, Flavored jets with exact anti-kt kinematics and tests of infrared and collinear safety , Phys. Rev. D 108 (2023), no. 9 094010, [ arXiv:2306.07314]

  92. [100]

    G. P. Salam and D. Wicke, Hadron masses and power corrections to event shapes , JHEP 05 (2001) 061, [ hep-ph/0102343]

  93. [101]

    Mateu, I

    V. Mateu, I. W. Stewart, and J. Thaler, Power Corrections to Event Shapes with Mass-Dependent Operators, Phys. Rev. D 87 (2013), no. 1 014025, [ arXiv:1209.3781]

  94. [102]

    H. Qu, C. Li, and S. Qian, Particle Transformer for Jet Tagging , arXiv:2202.03772

  95. [103]

    Amram, L

    O. Amram, L. Anzalone, J. Birk, D. A. Faroughy, A. Hallin, G. Kasieczka, M. Kr¨ amer, I. Pang, H. Reyes-Gonzalez, and D. Shih, Aspen Open Jets: Unlocking LHC Data for Foundation Models in Particle Physics , arXiv:2412.10504

  96. [104]

    P. T. Komiske, S. Kryhin, and J. Thaler, Disentangling quarks and gluons in CMS open data , Phys. Rev. D 106 (2022), no. 9 094021, [ arXiv:2205.04459]. – 24 –

  97. [105]

    E. J. Gumbel, Les valeurs extrˆ emes des distributions statistiques, in Annales de l’institut Henri Poincar´ e, vol. 5, pp. 115–158, 1935

  98. [106]

    E. J. Gumbel, The return period of flood flows , The annals of mathematical statistics 12 (1941), no. 2 163–190

  99. [107]

    Batson and Y

    J. Batson and Y. Kahn, Scaling Laws in Jet Classification , arXiv:2312.02264

  100. [108]

    Catani, L

    S. Catani, L. Trentadue, G. Turnock, and B. R. Webber, Resummation of large logarithms in e+ e- event shape distributions , Nucl. Phys. B 407 (1993) 3–42

  101. [109]

    M. H. Seymour, Searches for new particles using cone and cluster jet algorithms: A Comparative study, Z. Phys. C 62 (1994) 127–138

  102. [110]

    J. M. Butterworth, B. E. Cox, and J. R. Forshaw, W Wscattering at the CERN LHC , Phys. Rev. D 65 (2002) 096014, [ hep-ph/0201098]

  103. [111]

    J. M. Butterworth, J. R. Ellis, and A. R. Raklev, Reconstructing sparticle mass spectra using hadronic decays, JHEP 05 (2007) 033, [ hep-ph/0702150]

  104. [112]

    J. M. Butterworth, A. R. Davison, M. Rubin, and G. P. Salam, Jet substructure as a new Higgs search channel at the LHC , Phys. Rev. Lett. 100 (2008) 242001, [ arXiv:0802.2470]

  105. [113]

    Bisla, A

    D. Bisla, A. N. Saridena, and A. Choromanska, A theoretical-empirical approach to estimating sample complexity of dnns , CoRR abs/2105.01867 (2021) [arXiv:2105.01867]

  106. [114]

    Bahri, E

    Y. Bahri, E. Dyer, J. Kaplan, J. Lee, and U. Sharma, Explaining neural scaling laws , Proceedings of the National Academy of Sciences 121 (2024), no. 27 e2311878121

  107. [115]

    E. M. Metodiev and J. Thaler, Jet Topics: Disentangling Quarks and Gluons at Colliders , Phys. Rev. Lett. 120 (2018), no. 24 241602, [ arXiv:1802.00008]

  108. [116]

    D. A. Roberts, S. Yaida, and B. Hanin, The Principles of Deep Learning Theory , arXiv:2106.10165

  109. [117]

    Y. L. Dokshitzer, Calculation of the Structure Functions for Deep Inelastic Scattering and e+ e- Annihilation by Perturbation Theory in Quantum Chromodynamics. , Sov. Phys. JETP 46 (1977) 641–653

  110. [118]

    V. N. Gribov and L. N. Lipatov, Deep inelastic e p scattering in perturbation theory , Sov. J. Nucl. Phys. 15 (1972) 438–450

  111. [119]

    V. N. Gribov and L. N. Lipatov, e+ e- pair annihilation and deep inelastic e p scattering in perturbation theory, Sov. J. Nucl. Phys. 15 (1972) 675–684

  112. [120]

    L. N. Lipatov, The parton model and perturbation theory , Yad. Fiz. 20 (1974) 181–198

  113. [121]

    Altarelli and G

    G. Altarelli and G. Parisi, Asymptotic Freedom in Parton Language , Nucl. Phys. B 126 (1977) 298–318. – 25 –

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.