Pith. sign in

REVIEW 3 major objections 5 minor 7 references

Machine Learning Based Top Quark and W Jet Tagging to Hadronic Four-Top Final States Induced by SM as well as BSM Processes

T0 review · 3 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read The paper claims that, in simulated hadronic four-top events, machine-learning taggers match cut-based real tagging efficiencies while significantly lowering fake rates for light jets, at a modest cost in signal significance but with…

desk verdict A short proceedings paper whose headline ML fake-rate advantage is undermined by a circular truth-label mass window; the paper is honest but the central claim is not established. read the letter →

arxiv 2501.07589 v1 pith:5NXBJ2LA submitted 2025-01-08 hep-ph hep-ex

classification hep-phhep-ex
keywords topquarktaggingWbosonjetsubstructuremachinelearningboostedjetsfour-topproductionsubjettinessimbalancedclassification
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks whether machine-learning classifiers can replace the usual cut-based taggers for identifying hadronically decaying top quarks and W bosons in boosted final states, specifically in simulated SM and BSM four-top events. It finds that gradient-boosted trees and a multilayer perceptron, trained on subjettiness ratios and jet mass with undersampling, reach the same real tagging efficiency as the cut-based method, about 80%, while mistagging significantly fewer light jets. The cut-based tagger has a light-jet fake rate of roughly 65–70%, whereas the ML mistag rate is suppressed. The ML taggers also give a sharper signal peak in the reconstructed top-pair mass, at the price of a slightly lower fitted signal significance than the cut-based tagger (5.6 versus 6.1). The practical interest is that four-top searches are background-limited, so a tagger with a lower fake rate could directly improve the purity and sensitivity of such analyses.

What carries the argument

The central objects are the subjettiness ratios tau_21 = tau_2/tau_1 and tau_32 = tau_3/tau_2, together with the large-radius jet mass m_J, which describe how many prongs a boosted jet has and what mass it carries. These variables are used as input features for a gradient-boosting classifier and a multilayer perceptron, with random undersampling and cluster-centroid undersampling to balance the heavily top-jet-dominated training sets. The cut-based tagger uses explicit windows on the same variables: W-jets require 0.10 < tau_21 < 0.60, 0.50 < tau_32 < 0.85 and m_J in [60, 100] GeV, while top-jets require 0.30 < tau_21 < 0.70, 0.30 < tau_32 < 0.80 and m_J in [138, 208] GeV. The ML classifiers learn to separate the jet classes from the same feature space, and the comparison isolates what the learned decision boundary adds beyond the hand-made cuts.

What would settle it

Retrain both taggers with truth labels defined only by angular matching to the partonic top or W, without the m_J window, and re-measure real and fake efficiencies as a function of jet mass; if the ML fake-rate suppression persists, the central claim is robust, and if it disappears, the reported gain was an artifact of the label’s mass requirement.

Watch

Extended reading notes

Core claim

The central result is a direct performance comparison on MadGraph5 simulated events with a parameterized detector simulation. For both top-tagging and W-tagging, the ML classifiers give the same real efficiencies as the cut-based algorithm—high, about 80%, and mostly flat across jet mass—while the fake efficiencies are significantly lower. In the four-top exercise, the fitted signal significance is slightly lower for the ML tagger (5.6 versus 6.1), but the mass resolution of the signal peak is better (sigma about 80 GeV versus 106 GeV). The authors’ stated conclusion is that ML tagging is a viable lower-mistag alternative to cut-based tagging in hadronic four-top final states.

Load-bearing premise

The conclusions depend on the truth-label definition in Section 3, which requires a jet’s reconstructed mass to fall in the same window used by the cut-based tagger; because that mass is also an ML input feature, the reported fake-rate advantage may partly reflect the model applying the label’s own cut.

Editorial extensions

If this is right

  • Replacing the cut-based top and W tagger with the ML tagger in hadronic four-top searches would keep the true-tagging rate near 80% while cutting the light-jet mistag rate, directly reducing the dominant QCD background.
  • The sharper signal peak in the reconstructed top-pair mass (sigma about 80 GeV versus 106 GeV) means mass-window analyses could get better signal-to-background separation even though the ML significance is slightly lower.
  • Because the taggers are trained on combined pp and Z-prime datasets, the same ML tagger can be applied to both SM four-top background and BSM signal simulations without separate tuning.
  • The ML approach is implemented with off-the-shelf classifiers and undersampling, so it can be reproduced and adapted to other boosted-object tagging tasks without custom network architectures.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The reported fake-rate suppression may partly be the model learning the label’s own reconstructed-mass window, since the truth labels require m_J in the same range used by the cut-based tagger; if the labels were changed to parton-level matching only, the ML advantage could be smaller.
  • A testable extension is to decorrelate the ML tagger from m_J, for example by removing the mass feature or adding an adversarial loss; if fake suppression persists without the mass feature, the separation is genuinely substructure-based.
  • The feature set is minimal, so energy-flow polynomials or graph-based jet representations could plausibly push the fake rate lower still, although that remains to be demonstrated.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper studies ML-based tagging of boosted top quarks and W bosons in simulated hadronic four-top final states, comparing gradient boosting and multilayer perceptron classifiers against a cut-based tagger that uses subjettiness and reconstructed-mass windows. Truth labels are assigned by matching jets to parton-level top/W objects and requiring the reconstructed jet mass to lie inside fixed windows. The authors report that the ML taggers achieve real-tagging efficiencies similar to the cut-based tagger while substantially reducing fake efficiencies, and that a signal-plus-background fit to the dijet invariant mass yields comparable significances for the two methods.

Significance. If the central comparison were established, the result would be a useful benchmark for boosted-object tagging in four-top final states, showing that simple ML classifiers can outperform hand-built cut-based taggers in fake-rate suppression without sacrificing true-tagging efficiency. The paper is transparent about its sample definitions, enumerates its classifier variants and undersampling strategies, and provides ROC curves and efficiency plots, which is helpful for reproducibility. However, the analysis as presented does not yet support the central quantitative claim, because the truth labels, the cut-based tagger, and the ML classifier features share the same reconstructed-mass windows; the reported fake suppression may be an artifact of that overlap. In addition, no statistical uncertainties are given for any efficiency, AUC, or fit result, so the significance of the differences cannot be assessed.

major comments (3)
  1. [Section 3 together with Section 4.2] The truth labels for t-jets and W-jets are defined using reconstructed jet mass windows (138–208 GeV and 60–100 GeV) that are exactly the mass cuts used by the cut-based tagger in Section 4.2, and the reconstructed mass mJ is also an input feature to the ML classifiers. This makes the comparison circular: a classifier can reduce the fake efficiency defined by Eq. (2) simply by learning to reject jets whose reconstructed mass lies outside the label window, including truth-matched top or W jets that are mislabeled as light. Please repeat the training and evaluation with truth labels that do not require the reconstructed mass to be inside the cut-based mass window (e.g., using only the ∆R matching to the parton-level top/W), or explicitly demonstrate that the reported fake-rate suppression is unchanged when the label mass window is varied.
  2. [Section 5.2 and Figure 4] The claim that the ML method has "significantly less fake efficiencies" is not quantified with statistical uncertainties. The efficiencies plotted in Figure 4 and the numbers quoted in the conclusion are single values without confidence intervals, and the underlying event counts are not sufficient to judge whether the differences are significant. Please provide bootstrap or binomial uncertainties for every reported efficiency, together with the counts used in Eqs. (1) and (2), so the suppression can be assessed quantitatively.
  3. [Section 5.1] The signal significance comparison (Nsig/sqrt(Nbkg) = 6.1 versus 5.6) is based on a fit whose background ansatz (Bifurcated Gaussian plus Gaussian signal) is asserted without any goodness-of-fit measure, and no fit uncertainties are reported. Because the authors use this comparison to state that the cut-based method gives slightly higher significance, please report the fit range, the chi-squared per degree of freedom or an equivalent test, and the uncertainties on the fitted signal and background integrals, or soften the claim accordingly.
minor comments (5)
  1. [Section 2, Table 1] The labeling of the derived datasets is confusing: the text says samples 3 and 4 are unified into zp-sets and samples 0–2 into pp-sets, but the lower part of Table 1 shows ID 0 as "data_zp" and ID 1 as "data_pp". Please correct the table or the text so the reader can identify which files correspond to which training sample.
  2. [Section 3] The matching criterion ∆R(J,t) < 0.1 is not defined precisely: please state whether the reference is the parton-level top quark before hadronization or a particle-level object, and whether a jet can be matched to more than one truth object within the same angular radius.
  3. [Section 4.1.1] The undersampling techniques are only listed, not described, and no post-undersampling sizes are given. Please provide the final training-set sizes for each undersampling method so that the AUC values in Figure 1 can be compared on equal footing.
  4. [Section 5.1, Figure 3] The mass of the BSM resonance y0 is never stated in the text or the figure caption. Please specify it, and also define the event selection that produces the mJJ distribution (e.g., whether both jets are required to be tagged and which pT thresholds apply).
  5. [Conclusion] The conclusion states that "ML-based method has lower efficiencies", which appears to contradict Section 5.2 where the ML real efficiencies are described as the same as the cut-based ones. Please clarify whether the lower-efficiency statement refers to the W-tagging case or to a different operating point.

Circularity Check

1 steps flagged · score 6.0 of 10

Truth-label mass window is shared with the cut-based tagger and is an ML input, making part of the reported fake-rate suppression a construct of the label definition.

  1. self definitional [Section 3 (jet labels), Section 4.2 (cut-based mass window), Section 5.2 (Eqs. (1)-(2), Fig. 4)]
    "The true type jets labels are then based on the following criteria 1. truth t-jets: ∆R(J, t)< 0.1∧ 138 GeV≤ mJ≤ 208 GeV; 2. truth W-jets; ∆R(J, W )< 0.1∧ 60 GeV≤ mJ≤ 100 GeV; 3. truth light jets: otherwise. ... Variables defined and used for each jet in the classification are as follows ... m label ... εfake = N(tagged & not− matched) / (N(tagged & not− matched) + N(not− tagged & not− matched)) ... ML based algorithms give the same real efficiencies as cut-based, but significantly less fake efficiencies."

    The 'matched' class entering the fake-efficiency denominator is defined by the same reconstructed-mass window that the cut-based tagger uses and that is an input feature (m, Table 3) to the ML classifiers. Truth top/W jets whose reconstructed mass falls outside the label window are labeled 'light' and hence count as 'not matched'. An ML classifier can lower εfake simply by reproducing the label's own mass cut, moving such off-window true jets from 'tagged & not-matched' to 'not-tagged & not-matched' without learning any independent substructure separation. The claimed 'significantly less fake efficiencies' is therefore at least partly constructed by the truth-label definition rather than by genuine physics discrimination.

full rationale

The paper contains no load-bearing self-citations or imported uniqueness theorems; its ML classifiers and data generation are independent. However, the central performance comparison is partially circular. Section 3 defines truth t/W jets using reconstructed jet mass windows (138-208 GeV and 60-100 GeV) that are identical to the cut-based tagger's mass cuts (Section 4.2) and are available as input features to the ML classifiers. Because Eq. (2) calls any jet not satisfying this label 'not matched', the fake efficiency is defined in terms of the very mass window the algorithms are being compared on. An ML model can therefore reduce the reported fake rate by learning to reject off-window true jets, which the label already counts as light. This does not fully erase the comparison—subjettiness inputs and ROC analysis provide independent content—but the headline claim of 'significantly less fake efficiencies' is not established as an independent physics separation until the label's mass-window requirement is removed or varied. Score 6 reflects this partial, construction-level circularity in the central evaluation claim.

Assumptions & free parameters 9 free parameters · 5 assumptions · 2 invented entities

The central numbers rest on hand-chosen mass windows, arbitrary matching radii, a simplified detector simulation and an undefined BSM signal process y0. The mass windows and tau cuts are free parameters; the matching and detector assumptions are axioms; and the Z' and y0 particles are invented benchmark entities without independent evidence.

free parameters (9)
  • Truth t-jet mass window lower bound = 138 GeV
    Chosen by hand in Section 3 to define truth t-jets; this window directly sets the labels used for every efficiency and significance number.
  • Truth t-jet mass window upper bound = 208 GeV
    Chosen by hand in Section 3 as the upper edge of the truth t-jet definition.
  • Truth W-jet mass window lower bound = 60 GeV
    Chosen by hand in Section 3 to define truth W-jets.
  • Truth W-jet mass window upper bound = 100 GeV
    Chosen by hand in Section 3 as the upper edge of the truth W-jet definition.
  • Cut-based W-tag tau21 window = (0.10, 0.60)
    Hand-picked thresholds in Section 4.2 that define the baseline W tagger.
  • Cut-based W-tag tau32 window = (0.50, 0.85)
    Hand-picked thresholds in Section 4.2 for the baseline W tagger.
  • Cut-based top-tag tau21 window = (0.30, 0.70)
    Hand-picked thresholds in Section 4.2 that define the baseline top tagger.
  • Cut-based top-tag tau32 window = (0.30, 0.80)
    Hand-picked thresholds in Section 4.2 for the baseline top tagger.
  • ML classifier hyperparameters = not reported
    The gradient boosting and multilayer perceptron models are not specified with learning rates, tree counts, layer sizes or regularization, so the effective model complexity is unknown.
assumptions (5)
  • domain assumption Jet origin can be separated using subjettiness ratios tau21, tau32 and jet invariant mass mJ.
    The entire feature set and the cut-based selection rely on this; no comparison to alternative observables is provided.
  • domain assumption Truth matching with DeltaR below 0.1 to generator-level t or W correctly identifies the jet origin.
    Used in Section 3; the matching radius and the additional jet mass window are arbitrary choices.
  • domain assumption The parameterized detector simulation is an adequate stand-in for the LHC detector.
    All results are simulation-based; there is no validation against full simulation or collision data.
  • ad hoc to paper The mJJ distribution is described by a Bifurcated Gaussian background plus a Gaussian signal.
    Adopted in Section 5.1 to compute signal significance; no goodness-of-fit test or alternative fit is shown.
  • domain assumption A random 80/20 split of the same generated samples provides a valid generalization test.
    The training and test sets share the same generator and selection conditions; no k-fold or independent sample is used.
invented entities (2)
  • Z' boson with mass 1000 and 1250 GeV
    purpose: Generate the zp-sets (datasets 3 and 4) as benchmark BSM events with a resonance decaying to top quarks.
    No Lagrangian, width, couplings or search constraints are given; the particle appears only as a MadGraph sample label.
  • BSM resonance y0 with mass not specified
    purpose: Signal process tt y0 to tt tt used in Figure 3 to demonstrate the tagging comparison, scaled to 10 percent of the SM tttt background.
    The text never defines y0, its mass, production mechanism or decay, so the significance and mass-resolution numbers rest on this unspecified model.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Machine Learning Based Top Quark and W Jet Tagging to Hadronic Four-Top Final States Induced by SM as well as BSM Processes." pith.science (2026). https://pith.science/paper/5NXBJ2LA

@misc{pith2026250107589,
  author       = {Pith},
  title        = {Pith review of: Machine Learning Based Top Quark and W Jet Tagging to Hadronic Four-Top Final States Induced by SM as well as BSM Processes},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/5NXBJ2LA}},
  note         = {Machine review of arXiv:2501.07589}
}
read the original abstract

We study the application of selected ML techniques to the recognition of a substructure of hadronic final states (jets) and their tagging based on their possible origin in current HEP experiments using simulated events and a parameterized detector simulation. The results are then compared with the cut-based method.

Figures

Figures reproduced from arXiv: 2501.07589 by the authors.

Figure 1
Figure 1. Shapes of the τ21, τ32 subjettiness variables (top) and the large-R jet mass (bottom left) in the five samples used in training and testing of the tagging algo￾rithms. Performance of top-tagging for different classifiers shown via ROC curve (bottom right). The ratios between t-jets (W-jets) and light-jets are summarized in the following tables Variables defined and used for each jet in the classification are as foll… view at source ↗
Figure 2
Figure 2. GBC • Multi-layer Perceptron classifier (MLP) - based on neural networks. 4.1.1 Undersampling • very distorted ratio between t-jets and light-jets (in the direction of t-jets) • we settled for the undersampling applied to the training sets, which uses various tech￾niques to remove data from the major class • tested undersampling techniques: Random undersampling, Cluster centroids, Near miss, Repeated edited nearest … view at source ↗
Figure 3
Figure 3. Invariant mass of two t-tagged jets (top ML based, bottom cut-based algo￾rithm) for the process of SM t¯tt¯t (blue area) representing background process with the stacked signal process t¯t y0 → t¯tt¯t (red area) scaled to its 10%. The light red and blue areas show tagged and matched jets to highlight the tagging efficiencies. The background fit is given by black line using Bifurcated Gaussian and green line is the G… view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Efficiencies using cut-based and ML, t¯t y0 → t¯tt¯t. 6 Conclusion The real efficiencies of cut-based method in both t-jets and W-jets tagging are high about 80%, mostly flat, but unfortunatelly also having high mistagging rates about 65-70%. While ML￾based method has …

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

7 extracted references · 3 canonical work pages

  1. [1]

    , " * write output.state after.block =

    ENTRY address archive author booktitle chapter doi edition editor eid eprint howpublished institution isbn journal key month note number organization pages publisher school series title type url volume year label INTEGERS output.state before.all mid.sentence after.sentence after.block FUNCTION init.state.consts #0 'before.all := #1 'mid.sentence := #2 'af...

  2. [2]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...

  3. [3]

    50 Years of Quantum Chromodynamics

    Franz Gross et al. 50 Years of Quantum Chromodynamics . arXiv:2212.11107, December 2022. https://arxiv.org/abs/2212.11107

  4. [4]

    k-NN Approach to Unbalanced Data Distributions: A Case Study Involving Information Extraction

    Inderjeet Mani and Jianping Zhang. k-NN Approach to Unbalanced Data Distributions: A Case Study Involving Information Extraction . In Proceedings of the Workshop on Learning from Imbalanced Datasets , volume 126, 2003. https://www.site.uottawa.ca/ nat/Workshop2003/jzhang.pdf

  5. [5]

    Gradient Boosting Classifier

    Hemashree Kilari. Gradient Boosting Classifier . Medium.com, 2023. https://medium.com/@hemashreekilari9/understanding-gradient-boosting-632939b98764

  6. [6]

    M \"u ller and Sarah Guido

    Andreas G. M \"u ller and Sarah Guido. Introduction to Machine Learning with Python . O'Reilly Media, Beijing, Boston, Farnham, Sebastopol, Tokyo, 2016. ISBN: 978-1-449-36975-8

  7. [7]

    Cluster-Based Under-Sampling Approaches for Imbalanced Data Distributions

    Jane Yen and Yue-Shi Lee. Cluster-Based Under-Sampling Approaches for Imbalanced Data Distributions . Expert Systems with Applications , 36(3):5718--5727, April 2009. https://doi.org/10.1016/j.eswa.2008.06.108

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.