{"id":"485d8a19-f5be-4dee-990b-729fa2bf4746","arxiv_id":"2501.07589","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":9,"one_line_summary":"Using gradient boosting and neural networks on simulated jets, the authors find lower mistagging rates than a cut-based tagger for hadronic four-top final states, at similar real efficiency.","lead":"This paper tests whether machine learning classifiers can tell hadronic jets from top quarks and W bosons apart better than a simple cut-based method, using simulated LHC events. It reports that the ML taggers cut the fake rate sharply while keeping real tagging efficiency about the same, which could sharpen searches for four-top-quark final states.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Reported fake-rate advantage may be an artifact of truth labels that require reconstructed jet mass in the same window as the cut-based tagger; with mass-window-free labels the ML advantage could disappear.","rationale":"The paper's central claim is specifically about fake efficiencies: ML tagging matches cut-based real efficiency while mistagging fewer light jets. The most load-bearing condition for that claim is that the truth labels do not already encode the same reconstructed-mass information that the classifiers see as a feature and that the cut-based baseline uses for selection. Section 3 violates this condition by requiring the reconstructed jet mass to lie in the top/W window for a jet to be labeled a truth top or W. As a result, the relative ordering of ML and cut-based fake rates is confounded with the label definition. This is a concrete, testable concern rather than a mere disagreement with the community consensus. The reader's weakest assumption identifies the same issue, and the conditional verdict already reflects the resulting uncertainty, so I do not recommend changing the verdict. A matching-only relabeling and retraining test would settle whether the reported fake-rate suppression is a genuine ML advantage or an artifact of the label mass window.","tokens_in":4287,"tokens_out":2884,"duration_ms":31269,"concrete_test":"Retrain the GBC and MLP classifiers on the same pp and zp samples with truth labels defined by matching only: t-jets ∆R(J,t)<0.1, W-jets ∆R(J,W)<0.1, and all other jets light, with no mJ restriction. At the operating point matching the cut-based real efficiency as in Figure 4, recompute εfake from Eq. (2). If the ML fake-rate advantage over cut-based shrinks or disappears, the central claim is an artifact of the Section 3 mass-window labels; if the advantage persists, the claim is robust.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3 defines truth t-jets as ∆R(J,t)<0.1 ∧ 138 GeV ≤ mJ ≤ 208 GeV and truth W-jets as ∆R(J,W)<0.1 ∧ 60 GeV ≤ mJ ≤ 100 GeV, with everything else labeled light. The reconstructed jet mass mJ is also an input feature to both ML classifiers, and the cut-based tagger in Section 4.2 uses exactly these same mass windows plus subjettiness cuts. Thus the ML classifiers can lower the reported fake efficiency simply by learning to reject objects whose reconstructed mass falls outside the label's own mass window, including truth-matched top or W jets that Section 3 mislabels as light. The fake efficiency in Eq. (2) then counts those off-window truth jets in the denominator, so any model that reproduces the label mass cut will appear to suppress fakes. This is especially relevant for the pp samples, where the light-jet class is contaminated by unmatched top and W jets. Until the label mass cut is removed or varied, the claimed 'significantly less fake efficiencies' does not demonstrate an independent physics separation.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies ML-based tagging of boosted top quarks and W bosons in simulated hadronic four-top final states, comparing gradient boosting and multilayer perceptron classifiers against a cut-based tagger that uses subjettiness and reconstructed-mass windows. Truth labels are assigned by matching jets to parton-level top/W objects and requiring the reconstructed jet mass to lie inside fixed windows. The authors report that the ML taggers achieve real-tagging efficiencies similar to the cut-based tagger while substantially reducing fake efficiencies, and that a signal-plus-background fit to the dijet invariant mass yields comparable significances for the two methods.","tokens_in":4509,"tokens_out":4445,"duration_ms":44283,"significance":"If the central comparison were established, the result would be a useful benchmark for boosted-object tagging in four-top final states, showing that simple ML classifiers can outperform hand-built cut-based taggers in fake-rate suppression without sacrificing true-tagging efficiency. The paper is transparent about its sample definitions, enumerates its classifier variants and undersampling strategies, and provides ROC curves and efficiency plots, which is helpful for reproducibility. However, the analysis as presented does not yet support the central quantitative claim, because the truth labels, the cut-based tagger, and the ML classifier features share the same reconstructed-mass windows; the reported fake suppression may be an artifact of that overlap. In addition, no statistical uncertainties are given for any efficiency, AUC, or fit result, so the significance of the differences cannot be assessed.","major_comments":[{"comment":"The truth labels for t-jets and W-jets are defined using reconstructed jet mass windows (138–208 GeV and 60–100 GeV) that are exactly the mass cuts used by the cut-based tagger in Section 4.2, and the reconstructed mass mJ is also an input feature to the ML classifiers. This makes the comparison circular: a classifier can reduce the fake efficiency defined by Eq. (2) simply by learning to reject jets whose reconstructed mass lies outside the label window, including truth-matched top or W jets that are mislabeled as light. Please repeat the training and evaluation with truth labels that do not require the reconstructed mass to be inside the cut-based mass window (e.g., using only the ∆R matching to the parton-level top/W), or explicitly demonstrate that the reported fake-rate suppression is unchanged when the label mass window is varied.","section":"Section 3 together with Section 4.2"},{"comment":"The claim that the ML method has \"significantly less fake efficiencies\" is not quantified with statistical uncertainties. The efficiencies plotted in Figure 4 and the numbers quoted in the conclusion are single values without confidence intervals, and the underlying event counts are not sufficient to judge whether the differences are significant. Please provide bootstrap or binomial uncertainties for every reported efficiency, together with the counts used in Eqs. (1) and (2), so the suppression can be assessed quantitatively.","section":"Section 5.2 and Figure 4"},{"comment":"The signal significance comparison (Nsig/sqrt(Nbkg) = 6.1 versus 5.6) is based on a fit whose background ansatz (Bifurcated Gaussian plus Gaussian signal) is asserted without any goodness-of-fit measure, and no fit uncertainties are reported. Because the authors use this comparison to state that the cut-based method gives slightly higher significance, please report the fit range, the chi-squared per degree of freedom or an equivalent test, and the uncertainties on the fitted signal and background integrals, or soften the claim accordingly.","section":"Section 5.1"}],"minor_comments":[{"comment":"The labeling of the derived datasets is confusing: the text says samples 3 and 4 are unified into zp-sets and samples 0–2 into pp-sets, but the lower part of Table 1 shows ID 0 as \"data_zp\" and ID 1 as \"data_pp\". Please correct the table or the text so the reader can identify which files correspond to which training sample.","section":"Section 2, Table 1"},{"comment":"The matching criterion ∆R(J,t) < 0.1 is not defined precisely: please state whether the reference is the parton-level top quark before hadronization or a particle-level object, and whether a jet can be matched to more than one truth object within the same angular radius.","section":"Section 3"},{"comment":"The undersampling techniques are only listed, not described, and no post-undersampling sizes are given. Please provide the final training-set sizes for each undersampling method so that the AUC values in Figure 1 can be compared on equal footing.","section":"Section 4.1.1"},{"comment":"The mass of the BSM resonance y0 is never stated in the text or the figure caption. Please specify it, and also define the event selection that produces the mJJ distribution (e.g., whether both jets are required to be tagged and which pT thresholds apply).","section":"Section 5.1, Figure 3"},{"comment":"The conclusion states that \"ML-based method has lower efficiencies\", which appears to contradict Section 5.2 where the ML real efficiencies are described as the same as the cut-based ones. Please clarify whether the lower-efficiency statement refers to the W-tagging case or to a different operating point.","section":"Conclusion"}],"recommendation":"major_revision","confidential_remarks":"The main blocker is the label/mass-window circularity. If the authors can retrain with truth labels that do not impose the cut-based reconstructed-mass window and still observe fake-rate suppression, the paper would be a solid proceedings contribution. As it stands, the central quantitative claim is not yet supported. The paper is quite thin for a full journal article; as a proceedings contribution it may be acceptable after the revision, provided the major comments are addressed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nA quick read of Kvita et al. 2501.07589. It's a six-page proceedings piece comparing two ML jet taggers (gradient boosting, MLP) against a cut-based tagger for boosted top and W jets in four-top events. The headline claim—ML gives the same real efficiency with significantly less fake efficiency—is not established, because the truth labels are defined with the same reconstructed-mass windows that are also features and cut-based selections.\n\nWhat's new and good: as far as the paper's own references go, this specific comparison of GBC/MLP to cut-based on MadGraph four-top samples with tau21/tau32/mJ features isn't published elsewhere. The paper is clearly written and honest about the efficiency trade-off. It also shows distributions and ROC curves, and does a signal-peak fit exercise. For a proceedings note, that's fine.\n\nThe soft spot is not minor: in Sec. 3, a truth t-jet requires ΔR(J,t)<0.1 and 138≤mJ≤208 GeV; a truth W-jet requires ΔR(J,W)<0.1 and 60≤mJ≤100 GeV. Everything else is labeled light. Those mass windows are exactly what the cut-based tagger uses (Sec. 4.2) and mJ is an input to the ML classifiers. So the fake efficiency of Eq. (2) counts truth-matched jets that fall outside the mass window as 'not matched,' and any classifier that learns the mass cut will appear to suppress fakes. The stress-test note holds up.\n\nOther issues: no statistical uncertainties on any efficiency or AUC; the BSM signal y0 is never defined; no hyperparameters or code given. The ROC AUCs are low (0.64–0.67), so the ML advantage is modest even setting the circularity aside. No comparison to modern taggers.\n\nIs the paper thoughtless? No. It just draws a stronger conclusion than the label definition supports. The fix is straightforward: define truth labels by parton-level matching only, not by reconstructed mass, and rerun. Or at least vary the window and show the ML advantage persists.\n\nWho's it for? People working on four-top searches or boosted jet tagging might read it as a cautionary tale about label circularity. Its current claims shouldn't be cited as evidence that ML beats cut-based tagging. If this came to me as a referee for a journal, I'd send it back for major revision, focusing on the label definition and uncertainties. For a proceedings contribution, it's acceptable as a talk but the conclusion as written overreaches.\n\nRecommendation: engage with it if you're interested in the method, but don't take the fake-rate comparison at face value. It deserves peer review only after revision.\n\nBest,\n\n[Your name]","headline":"A short proceedings paper whose headline ML fake-rate advantage is undermined by a circular truth-label mass window; the paper is honest but the central claim is not established.","tokens_in":5157,"tokens_out":4630,"would_cite":false,"duration_ms":42278,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that, in simulated hadronic four-top events, machine-learning taggers match cut-based real tagging efficiencies while significantly lowering fake rates for light jets, at a modest cost in signal significance but with…","keywords":["top quark tagging","W boson tagging","jet substructure","machine learning","boosted jets","four-top production","subjettiness","imbalanced classification"],"falsifier":"Retrain both taggers with truth labels defined only by angular matching to the partonic top or W, without the m_J window, and re-measure real and fake efficiencies as a function of jet mass; if the ML fake-rate suppression persists, the central claim is robust, and if it disappears, the reported gain was an artifact of the label’s mass requirement.","tokens_in":4019,"feed_emoji":"🤖","tokens_out":5314,"duration_ms":51159,"temperature":0.7,"pith_summary":"This paper asks whether machine-learning classifiers can replace the usual cut-based taggers for identifying hadronically decaying top quarks and W bosons in boosted final states, specifically in simulated SM and BSM four-top events. It finds that gradient-boosted trees and a multilayer perceptron, trained on subjettiness ratios and jet mass with undersampling, reach the same real tagging efficiency as the cut-based method, about 80%, while mistagging significantly fewer light jets. The cut-based tagger has a light-jet fake rate of roughly 65–70%, whereas the ML mistag rate is suppressed. The ML taggers also give a sharper signal peak in the reconstructed top-pair mass, at the price of a slightly lower fitted signal significance than the cut-based tagger (5.6 versus 6.1). The practical interest is that four-top searches are background-limited, so a tagger with a lower fake rate could directly improve the purity and sensitivity of such analyses.","feed_headline":"ML taggers match cut-based real-tag rates with fewer fakes","feed_subtitle":"In simulated four-top events, the ML tagger keeps about 80% true efficiency while suppressing light-jet mistagging.","key_machinery":"The central objects are the subjettiness ratios tau_21 = tau_2/tau_1 and tau_32 = tau_3/tau_2, together with the large-radius jet mass m_J, which describe how many prongs a boosted jet has and what mass it carries. These variables are used as input features for a gradient-boosting classifier and a multilayer perceptron, with random undersampling and cluster-centroid undersampling to balance the heavily top-jet-dominated training sets. The cut-based tagger uses explicit windows on the same variables: W-jets require 0.10 < tau_21 < 0.60, 0.50 < tau_32 < 0.85 and m_J in [60, 100] GeV, while top-jets require 0.30 < tau_21 < 0.70, 0.30 < tau_32 < 0.80 and m_J in [138, 208] GeV. The ML classifiers learn to separate the jet classes from the same feature space, and the comparison isolates what the learned decision boundary adds beyond the hand-made cuts.","core_discovery":"The central result is a direct performance comparison on MadGraph5 simulated events with a parameterized detector simulation. For both top-tagging and W-tagging, the ML classifiers give the same real efficiencies as the cut-based algorithm—high, about 80%, and mostly flat across jet mass—while the fake efficiencies are significantly lower. In the four-top exercise, the fitted signal significance is slightly lower for the ML tagger (5.6 versus 6.1), but the mass resolution of the signal peak is better (sigma about 80 GeV versus 106 GeV). The authors’ stated conclusion is that ML tagging is a viable lower-mistag alternative to cut-based tagging in hadronic four-top final states.","pith_inferences":["The reported fake-rate suppression may partly be the model learning the label’s own reconstructed-mass window, since the truth labels require m_J in the same range used by the cut-based tagger; if the labels were changed to parton-level matching only, the ML advantage could be smaller.","A testable extension is to decorrelate the ML tagger from m_J, for example by removing the mass feature or adding an adversarial loss; if fake suppression persists without the mass feature, the separation is genuinely substructure-based.","The feature set is minimal, so energy-flow polynomials or graph-based jet representations could plausibly push the fake rate lower still, although that remains to be demonstrated."],"forward_implications":["Replacing the cut-based top and W tagger with the ML tagger in hadronic four-top searches would keep the true-tagging rate near 80% while cutting the light-jet mistag rate, directly reducing the dominant QCD background.","The sharper signal peak in the reconstructed top-pair mass (sigma about 80 GeV versus 106 GeV) means mass-window analyses could get better signal-to-background separation even though the ML significance is slightly lower.","Because the taggers are trained on combined pp and Z-prime datasets, the same ML tagger can be applied to both SM four-top background and BSM signal simulations without separate tuning.","The ML approach is implemented with off-the-shelf classifiers and undersampling, so it can be reproduced and adapted to other boosted-object tagging tasks without custom network architectures."],"supporting_citations":[{"why":"Motivates and supplies the k-NN based undersampling approach used to rebalance the heavily top-jet-dominated training sets.","marker":"[2]"},{"why":"Supplies the gradient boosting classifier implementation used as one of the two ML taggers.","marker":"[3]"},{"why":"Supplies the machine-learning toolkit used to build and evaluate the MLP and GBC classifiers.","marker":"[4]"},{"why":"Supplies the cluster-based undersampling method tested for handling the imbalanced jet-label distributions.","marker":"[5]"}],"fun_headline_variants":["ML jet taggers cut fakes while keeping 80% real-tag rate","For four-top events, ML tagging beats cut-based on false positives","ML W and top taggers: fewer fakes, same real efficiency","Lower fake rates, comparable signal: ML tagging in four-top","ML vs cut-based jet tagging: same real, fewer fake in four-top"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The conclusions depend on the truth-label definition in Section 3, which requires a jet’s reconstructed mass to fall in the same window used by the cut-based tagger; because that mass is also an ML input feature, the reported fake-rate advantage may partly reflect the model applying the label’s own cut.","fun_headline_variants_meta":{"raw":{"variants":["ML jet taggers cut fakes while keeping 80% real-tag rate","For four-top events, ML tagging beats cut-based on false positives","ML W and top taggers: fewer fakes, same real efficiency","Lower fake rates, comparable signal: ML tagging in four-top","ML vs cut-based jet tagging: same real, fewer fake in four-top"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000685,"raw_usage":{"total_tokens":3002,"prompt_tokens":736,"completion_tokens":2266,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":352,"completion_tokens_details":{"reasoning_tokens":2170}},"tokens_in":352,"tokens_out":2266,"duration_ms":13905,"temperature":1.0,"reasoning_tokens":2170,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T21:35:34.958042+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Retrain both taggers with truth labels defined only by angular matching to the partonic top or W, without the m_J window, and re-measure real and fake efficiencies as a function of jet mass; if the ML fake-rate suppression persists, the central claim is robust, and if it disappears, the reported gain was an artifact of the label’s mass requirement.","supporting_citations":[{"cited_title":"k-NN Approach to Unbalanced Data Distributions: A Case Study Involving Information Extraction","cited_arxiv_id":null,"evidence_quote":"Supplies the machine-learning toolkit used to build and evaluate the MLP and GBC classifiers."},{"cited_title":"Gradient Boosting Classifier","cited_arxiv_id":null,"evidence_quote":"Supplies the cluster-based undersampling method tested for handling the imbalanced jet-label distributions."}],"review_version":1}