REVIEW 3 major objections 5 minor 7 references
Machine Learning Based Top Quark and W Jet Tagging to Hadronic Four-Top Final States Induced by SM as well as BSM Processes
T0 review · 3 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read The paper claims that, in simulated hadronic four-top events, machine-learning taggers match cut-based real tagging efficiencies while significantly lowering fake rates for light jets, at a modest cost in signal significance but with…
desk verdict A short proceedings paper whose headline ML fake-rate advantage is undermined by a circular truth-label mass window; the paper is honest but the central claim is not established. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central objects are the subjettiness ratios tau_21 = tau_2/tau_1 and tau_32 = tau_3/tau_2, together with the large-radius jet mass m_J, which describe how many prongs a boosted jet has and what mass it carries. These variables are used as input features for a gradient-boosting classifier and a multilayer perceptron, with random undersampling and cluster-centroid undersampling to balance the heavily top-jet-dominated training sets. The cut-based tagger uses explicit windows on the same variables: W-jets require 0.10 < tau_21 < 0.60, 0.50 < tau_32 < 0.85 and m_J in [60, 100] GeV, while top-jets require 0.30 < tau_21 < 0.70, 0.30 < tau_32 < 0.80 and m_J in [138, 208] GeV. The ML classifiers learn to separate the jet classes from the same feature space, and the comparison isolates what the learned decision boundary adds beyond the hand-made cuts.
What would settle it
Retrain both taggers with truth labels defined only by angular matching to the partonic top or W, without the m_J window, and re-measure real and fake efficiencies as a function of jet mass; if the ML fake-rate suppression persists, the central claim is robust, and if it disappears, the reported gain was an artifact of the label’s mass requirement.
Extended reading notes
Core claim
The central result is a direct performance comparison on MadGraph5 simulated events with a parameterized detector simulation. For both top-tagging and W-tagging, the ML classifiers give the same real efficiencies as the cut-based algorithm—high, about 80%, and mostly flat across jet mass—while the fake efficiencies are significantly lower. In the four-top exercise, the fitted signal significance is slightly lower for the ML tagger (5.6 versus 6.1), but the mass resolution of the signal peak is better (sigma about 80 GeV versus 106 GeV). The authors’ stated conclusion is that ML tagging is a viable lower-mistag alternative to cut-based tagging in hadronic four-top final states.
Load-bearing premise
The conclusions depend on the truth-label definition in Section 3, which requires a jet’s reconstructed mass to fall in the same window used by the cut-based tagger; because that mass is also an ML input feature, the reported fake-rate advantage may partly reflect the model applying the label’s own cut.
Editorial extensions
If this is right
- Replacing the cut-based top and W tagger with the ML tagger in hadronic four-top searches would keep the true-tagging rate near 80% while cutting the light-jet mistag rate, directly reducing the dominant QCD background.
- The sharper signal peak in the reconstructed top-pair mass (sigma about 80 GeV versus 106 GeV) means mass-window analyses could get better signal-to-background separation even though the ML significance is slightly lower.
- Because the taggers are trained on combined pp and Z-prime datasets, the same ML tagger can be applied to both SM four-top background and BSM signal simulations without separate tuning.
- The ML approach is implemented with off-the-shelf classifiers and undersampling, so it can be reproduced and adapted to other boosted-object tagging tasks without custom network architectures.
Reading between the lines
- The reported fake-rate suppression may partly be the model learning the label’s own reconstructed-mass window, since the truth labels require m_J in the same range used by the cut-based tagger; if the labels were changed to parton-level matching only, the ML advantage could be smaller.
- A testable extension is to decorrelate the ML tagger from m_J, for example by removing the mass feature or adding an adversarial loss; if fake suppression persists without the mass feature, the separation is genuinely substructure-based.
- The feature set is minimal, so energy-flow polynomials or graph-based jet representations could plausibly push the fake rate lower still, although that remains to be demonstrated.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies ML-based tagging of boosted top quarks and W bosons in simulated hadronic four-top final states, comparing gradient boosting and multilayer perceptron classifiers against a cut-based tagger that uses subjettiness and reconstructed-mass windows. Truth labels are assigned by matching jets to parton-level top/W objects and requiring the reconstructed jet mass to lie inside fixed windows. The authors report that the ML taggers achieve real-tagging efficiencies similar to the cut-based tagger while substantially reducing fake efficiencies, and that a signal-plus-background fit to the dijet invariant mass yields comparable significances for the two methods.
Significance. If the central comparison were established, the result would be a useful benchmark for boosted-object tagging in four-top final states, showing that simple ML classifiers can outperform hand-built cut-based taggers in fake-rate suppression without sacrificing true-tagging efficiency. The paper is transparent about its sample definitions, enumerates its classifier variants and undersampling strategies, and provides ROC curves and efficiency plots, which is helpful for reproducibility. However, the analysis as presented does not yet support the central quantitative claim, because the truth labels, the cut-based tagger, and the ML classifier features share the same reconstructed-mass windows; the reported fake suppression may be an artifact of that overlap. In addition, no statistical uncertainties are given for any efficiency, AUC, or fit result, so the significance of the differences cannot be assessed.
major comments (3)
- [Section 3 together with Section 4.2] The truth labels for t-jets and W-jets are defined using reconstructed jet mass windows (138–208 GeV and 60–100 GeV) that are exactly the mass cuts used by the cut-based tagger in Section 4.2, and the reconstructed mass mJ is also an input feature to the ML classifiers. This makes the comparison circular: a classifier can reduce the fake efficiency defined by Eq. (2) simply by learning to reject jets whose reconstructed mass lies outside the label window, including truth-matched top or W jets that are mislabeled as light. Please repeat the training and evaluation with truth labels that do not require the reconstructed mass to be inside the cut-based mass window (e.g., using only the ∆R matching to the parton-level top/W), or explicitly demonstrate that the reported fake-rate suppression is unchanged when the label mass window is varied.
- [Section 5.2 and Figure 4] The claim that the ML method has "significantly less fake efficiencies" is not quantified with statistical uncertainties. The efficiencies plotted in Figure 4 and the numbers quoted in the conclusion are single values without confidence intervals, and the underlying event counts are not sufficient to judge whether the differences are significant. Please provide bootstrap or binomial uncertainties for every reported efficiency, together with the counts used in Eqs. (1) and (2), so the suppression can be assessed quantitatively.
- [Section 5.1] The signal significance comparison (Nsig/sqrt(Nbkg) = 6.1 versus 5.6) is based on a fit whose background ansatz (Bifurcated Gaussian plus Gaussian signal) is asserted without any goodness-of-fit measure, and no fit uncertainties are reported. Because the authors use this comparison to state that the cut-based method gives slightly higher significance, please report the fit range, the chi-squared per degree of freedom or an equivalent test, and the uncertainties on the fitted signal and background integrals, or soften the claim accordingly.
minor comments (5)
- [Section 2, Table 1] The labeling of the derived datasets is confusing: the text says samples 3 and 4 are unified into zp-sets and samples 0–2 into pp-sets, but the lower part of Table 1 shows ID 0 as "data_zp" and ID 1 as "data_pp". Please correct the table or the text so the reader can identify which files correspond to which training sample.
- [Section 3] The matching criterion ∆R(J,t) < 0.1 is not defined precisely: please state whether the reference is the parton-level top quark before hadronization or a particle-level object, and whether a jet can be matched to more than one truth object within the same angular radius.
- [Section 4.1.1] The undersampling techniques are only listed, not described, and no post-undersampling sizes are given. Please provide the final training-set sizes for each undersampling method so that the AUC values in Figure 1 can be compared on equal footing.
- [Section 5.1, Figure 3] The mass of the BSM resonance y0 is never stated in the text or the figure caption. Please specify it, and also define the event selection that produces the mJJ distribution (e.g., whether both jets are required to be tagged and which pT thresholds apply).
- [Conclusion] The conclusion states that "ML-based method has lower efficiencies", which appears to contradict Section 5.2 where the ML real efficiencies are described as the same as the cut-based ones. Please clarify whether the lower-efficiency statement refers to the W-tagging case or to a different operating point.
Circularity Check
Truth-label mass window is shared with the cut-based tagger and is an ML input, making part of the reported fake-rate suppression a construct of the label definition.
-
self definitional
[Section 3 (jet labels), Section 4.2 (cut-based mass window), Section 5.2 (Eqs. (1)-(2), Fig. 4)]
"The true type jets labels are then based on the following criteria 1. truth t-jets: ∆R(J, t)< 0.1∧ 138 GeV≤ mJ≤ 208 GeV; 2. truth W-jets; ∆R(J, W )< 0.1∧ 60 GeV≤ mJ≤ 100 GeV; 3. truth light jets: otherwise. ... Variables defined and used for each jet in the classification are as follows ... m label ... εfake = N(tagged & not− matched) / (N(tagged & not− matched) + N(not− tagged & not− matched)) ... ML based algorithms give the same real efficiencies as cut-based, but significantly less fake efficiencies."
The 'matched' class entering the fake-efficiency denominator is defined by the same reconstructed-mass window that the cut-based tagger uses and that is an input feature (m, Table 3) to the ML classifiers. Truth top/W jets whose reconstructed mass falls outside the label window are labeled 'light' and hence count as 'not matched'. An ML classifier can lower εfake simply by reproducing the label's own mass cut, moving such off-window true jets from 'tagged & not-matched' to 'not-tagged & not-matched' without learning any independent substructure separation. The claimed 'significantly less fake efficiencies' is therefore at least partly constructed by the truth-label definition rather than by genuine physics discrimination.
full rationale
The paper contains no load-bearing self-citations or imported uniqueness theorems; its ML classifiers and data generation are independent. However, the central performance comparison is partially circular. Section 3 defines truth t/W jets using reconstructed jet mass windows (138-208 GeV and 60-100 GeV) that are identical to the cut-based tagger's mass cuts (Section 4.2) and are available as input features to the ML classifiers. Because Eq. (2) calls any jet not satisfying this label 'not matched', the fake efficiency is defined in terms of the very mass window the algorithms are being compared on. An ML model can therefore reduce the reported fake rate by learning to reject off-window true jets, which the label already counts as light. This does not fully erase the comparison—subjettiness inputs and ROC analysis provide independent content—but the headline claim of 'significantly less fake efficiencies' is not established as an independent physics separation until the label's mass-window requirement is removed or varied. Score 6 reflects this partial, construction-level circularity in the central evaluation claim.
Assumptions & free parameters
free parameters (9)
- Truth t-jet mass window lower bound =
138 GeV
- Truth t-jet mass window upper bound =
208 GeV
- Truth W-jet mass window lower bound =
60 GeV
- Truth W-jet mass window upper bound =
100 GeV
- Cut-based W-tag tau21 window =
(0.10, 0.60)
- Cut-based W-tag tau32 window =
(0.50, 0.85)
- Cut-based top-tag tau21 window =
(0.30, 0.70)
- Cut-based top-tag tau32 window =
(0.30, 0.80)
- ML classifier hyperparameters =
not reported
assumptions (5)
- domain assumption Jet origin can be separated using subjettiness ratios tau21, tau32 and jet invariant mass mJ.
- domain assumption Truth matching with DeltaR below 0.1 to generator-level t or W correctly identifies the jet origin.
- domain assumption The parameterized detector simulation is an adequate stand-in for the LHC detector.
- ad hoc to paper The mJJ distribution is described by a Bifurcated Gaussian background plus a Gaussian signal.
- domain assumption A random 80/20 split of the same generated samples provides a valid generalization test.
invented entities (2)
-
Z' boson with mass 1000 and 1250 GeV
-
BSM resonance y0 with mass not specified
Cite this review
Pith. "Pith review of Machine Learning Based Top Quark and W Jet Tagging to Hadronic Four-Top Final States Induced by SM as well as BSM Processes." pith.science (2026). https://pith.science/paper/5NXBJ2LA
@misc{pith2026250107589,
author = {Pith},
title = {Pith review of: Machine Learning Based Top Quark and W Jet Tagging to Hadronic Four-Top Final States Induced by SM as well as BSM Processes},
year = {2026},
howpublished = {\url{https://pith.science/paper/5NXBJ2LA}},
note = {Machine review of arXiv:2501.07589}
}
read the original abstract
We study the application of selected ML techniques to the recognition of a substructure of hadronic final states (jets) and their tagging based on their possible origin in current HEP experiments using simulated events and a parameterized detector simulation. The results are then compared with the cut-based method.
Figures
Figures from the paper (1 more)
Reference graph
Works this paper leans on
-
[1]
, " * write output.state after.block =
ENTRY address archive author booktitle chapter doi edition editor eid eprint howpublished institution isbn journal key month note number organization pages publisher school series title type url volume year label INTEGERS output.state before.all mid.sentence after.sentence after.block FUNCTION init.state.consts #0 'before.all := #1 'mid.sentence := #2 'af...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
50 Years of Quantum Chromodynamics
Franz Gross et al. 50 Years of Quantum Chromodynamics . arXiv:2212.11107, December 2022. https://arxiv.org/abs/2212.11107
arXiv 2022
-
[4]
k-NN Approach to Unbalanced Data Distributions: A Case Study Involving Information Extraction
Inderjeet Mani and Jianping Zhang. k-NN Approach to Unbalanced Data Distributions: A Case Study Involving Information Extraction . In Proceedings of the Workshop on Learning from Imbalanced Datasets , volume 126, 2003. https://www.site.uottawa.ca/ nat/Workshop2003/jzhang.pdf
work page 2003
-
[5]
Hemashree Kilari. Gradient Boosting Classifier . Medium.com, 2023. https://medium.com/@hemashreekilari9/understanding-gradient-boosting-632939b98764
work page 2023
-
[6]
Andreas G. M \"u ller and Sarah Guido. Introduction to Machine Learning with Python . O'Reilly Media, Beijing, Boston, Farnham, Sebastopol, Tokyo, 2016. ISBN: 978-1-449-36975-8
work page 2016
-
[7]
Cluster-Based Under-Sampling Approaches for Imbalanced Data Distributions
Jane Yen and Yue-Shi Lee. Cluster-Based Under-Sampling Approaches for Imbalanced Data Distributions . Expert Systems with Applications , 36(3):5718--5727, April 2009. https://doi.org/10.1016/j.eswa.2008.06.108
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.