REVIEW 4 major objections 5 minor 29 references
Investigation of the performance of a GNN-based b-jet tagging method in heavy-ion collisions
T0 review · 4 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read A graph-neural-network tagger still beats classical b-tagging when jets are buried in a Pb-Pb background.
desk verdict A careful feasibility study of GNN b-tagging in Pb-Pb-like background whose internal centrality trends are credible, but the 'outperforms classical' claim rests on a non-matched external baseline and a best-case background model. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the GN1 graph neural network, adapted from pp use: each jet is a fully connected graph whose nodes are up to 40 tracks with the highest impact-parameter significance, carrying track $p_T$, $\eta$, $\phi$, charge, and signed $d_{xy}$/$d_z$; jet-level $p_T$, $\eta$, $\phi$, mass are concatenated to each node. Three output heads predict jet flavour (b/c/light), each track's origin (background, primary, charm decay, beauty decay, other secondary), and pairwise vertex assignment, with a combined loss. The vertex and track-origin tasks force the network to internalize displaced-vertex structure, which is exactly what the Pb-Pb background tries to hide; the graph edges let it compare many track pairs at once. The evaluation machinery is the overlay: background tracks from central (0-20%) to peripheral (80-100%) classes are added within $\Delta R<0.4$ of the jet axis, and the discriminant $D_b = \log[p_b/((1-f_c)p_{lf}+f_c p_c)]$ with $f_c=0.018$ turns the flavour probabilities into a tagger.
What would settle it
Generate a test set of the same pp jets overlaid with background from a full Pb-Pb simulation that includes weak-decay and material-interaction tracks with nonzero impact parameters (or a realistic Pb-Pb event generator with jet energy loss), and recompute the ROC AUC for the pp-trained and centrality-trained models. If, at $20<p_{T,jet}<70$ GeV/c and 0-20% centrality, the GNN AUC drops below the classical secondary-vertex tagger's AUC, the paper's main superiority claim would be refuted under realistic conditions.
Extended reading notes
Core claim
The paper's claim is that the GNN tagger's learned track relations survive contamination, not just degrade gracefully. When a pp-trained model is applied to jets overlaid with 0-20% central background, the b-jet/c-jet/light-flavour separation in the discriminant variable $D_b$ narrows but remains distinct; for $20<p_{T,jet}<70$ GeV/c the network still reaches the kind of purity-efficiency working points that classical SV tagging provides in pp. Training dedicated models on jets overlaid with background from the same centrality class (0-20%, 20-40%, etc.) improves AUC at every centrality, with the largest gain at low $p_T$ and in the most central collisions. The paper concludes that background-aware training largely compensates for the contamination, and that the residual degradation is mainly the model having never seen background rather than an intrinsic failure of graph-based tagging.
Load-bearing premise
The paper's central comparison assumes that Pb-Pb background can be represented by primary-vertex-only tracks with zero impact parameter overlaid on unchanged pp jets; if real background contains many long-lived or material-interaction tracks with large impact parameters, the GNN's separation advantage would be weaker than measured here.
Editorial extensions
If this is right
- Low-$p_T$ b-jet measurements in Pb-Pb, which classical taggers cannot reach, are a realistic target for a GNN tagger trained with centrality-matched backgrounds.
- A single pp-trained model is not enough; the paper's per-centrality retraining recipe should be adopted in heavy-ion analyses.
- The working points at $20<p_{T,jet}<70$ GeV/c suggest the tagger can supply both high efficiency and high purity where the classical SV tagger sits around 0.2-0.3 efficiency at 0.3-0.5 purity.
- At the highest background density (0-20% centrality) the model keeps a usable separation, implying the background is not washing out b-hadron displacement signatures completely.
Reading between the lines
- The overlay treats all background tracks as coming from the primary vertex with zero impact parameter, so the GNN is not tested against large-$d_0$ background tracks from weak decays or detector material; real Pb-Pb events could push the tagger's false-tag rate up, which is a direct extension of the paper's stated limitation.
- In actual heavy-ion collisions the jet itself is modified by energy loss, so the input jet kinematics and substructure differ from a pp jet plus background; whether the advantage survives that change is not addressed by this controlled overlay.
- The comparison baseline is a classical SV tagger in pp; a dedicated Pb-Pb-tuned classical tagger might close part of the gap, so the claim of 'outperforms classical methods' is strongest for the implemented comparison.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper adapts the ATLAS GN1 graph neural network for b-jet tagging in the heavy-ion environment of ALICE. Pythia8 pp jets are simulated with Delphes and overlaid with AMPT Pb-Pb background particles in five centrality classes. The authors evaluate a model trained on pure pp jets and separately trained models on centrality-specific background-overlaid jets, reporting purity-efficiency curves, ROC AUC values, and D_b distributions. They claim the GNN outperforms classical tagging methods such as the secondary-vertex (SV) approach and that centrality-specific retraining mitigates the performance loss. The internal training and evaluation are standard, but the headline comparison to classical methods relies on an external, non-matched baseline, and the background model is simplified to primary-vertex-only particles with zero impact parameter, which is favorable for a displaced-vertex tagger.
Significance. If the claims hold, the paper offers a promising direction for low-pT b-jet measurements at ALICE in Pb-Pb collisions, where classical taggers suffer from high background. The study is clearly presented, uses standard simulation tools, and provides a plausible qualitative result: background contamination degrades tagging performance, and retraining on centrality-specific background improves robustness. The centrality-dependent trend and the benefit of background-aware training are internally consistent. However, the quantitative outperformance claim over classical methods is not established because the comparison is not apples-to-apples, and the simplified background model means the reported purity is likely an upper bound. These limitations are acknowledged in Sec. 5, but they directly affect the paper's central claim and need to be addressed before the results can be considered conclusive.
major comments (4)
- [Sec. 4.1, Fig. 6; Sec. 5] The claim that the GNN 'continues to outperform traditional tagging methods' is based on a comparison to external ALICE SV results from Ref. [16], which were obtained with a different simulation chain (full Geant4 vs. Delphes), different detector modeling, and likely different jet and track selections and purity/efficiency definitions. The GNN curves in Fig. 6 are evaluated on Pythia8+AMPT overlaid jets, while the SV numbers come from a separate ALICE analysis. To support the headline claim, the authors should run a classical tagger (e.g., a simple displaced-vertex or impact-parameter tagger) on the same overlaid jets and compare at matched working points and kinematic bins. As presented, the outperformance claim is not quantitatively established.
- [Sec. 2 and Sec. 5] The background model assumes every AMPT particle originates from the primary vertex with zero impact parameter, and the summary in Sec. 5 explicitly notes that long-lived background particles producing secondary vertices and jet energy loss are not incorporated. Because the GN1 discriminants rely on dxy, dz, and vertex structure, this is the most favorable possible background for a displaced-vertex tagger: the only added confusion comes from primary-vertex tracks, exactly the population the model is designed to reject. Real Pb-Pb events contain weak-decay and material-interaction tracks with large impact parameters that could mimic b-hadron decay products and reduce purity. The authors should either include a more realistic background model or explicitly present the current results as an upper bound on achievable performance, tempering the quantitative claims accordingly.
- [Figs. 6 and 7] No statistical uncertainties are shown on the purity-efficiency curves or the ROC AUC values. The claim that centrality-specific retraining 'significantly mitigates' the performance loss requires an estimate of the statistical precision of these curves (e.g., number of jets per pT bin, or bootstrap/binomial errors). Without error bars, the reader cannot assess whether the differences between the pp-trained and background-trained models are statistically significant, which is a standard requirement for machine-learning performance comparisons in this context.
- [Sec. 3.1] The GNN input is truncated to the 40 tracks with the highest impact-parameter significance when a jet contains more than 40 constituent tracks. In 0-20% central Pb-Pb collisions, the average number of background tracks inside the jet cone is large (Fig. 1 shows values approaching or exceeding 100), so many jets will exceed this cap. The paper does not report how often the 40-track limit is reached, whether signal tracks from the b-hadron decay are retained under this truncation, or how the truncation affects the reported performance. This is a load-bearing detail for the robustness claim at high centrality and should be quantified.
minor comments (5)
- [Title/header] The title line contains a typo: 'GNN-basedb-jet' should read 'GNN-based b-jet'.
- [Fig. 5] The caption says 'Distributions of Db and pT, jet', but the axes in the figure appear to show N_bkg_trks and Delta Db. The caption and axis labels should be made consistent.
- [Sec. 3.2] The sentence '3.63 million pp jets are used for training and validation, with a 0.8:0.2 split' could be clarified to state explicitly that this is an 80%/20% training/validation split.
- [Sec. 4.1, Eq. (3)] The discriminant mixing parameter fc is set to 0.018, inherited from the ATLAS GN1 study. A brief sensitivity check showing how the results depend on fc would increase confidence that the reported conclusions are not driven by this choice.
- [General] The paper states it is a 'comprehensive evaluation', but the study is limited to charged-particle jets with track pT > 0.15 GeV/c and only stable generated particles. This should be explicitly stated in the abstract or introduction so that readers do not overinterpret the scope.
Circularity Check
No significant circularity: reported efficiencies, purities, and AUCs are measured on held-out simulated jets, and the classical-baseline comparison is taken from an external ALICE publication rather than generated by the paper's own pipeline.
full rationale
This is an empirical machine-learning evaluation whose claims do not reduce to their inputs. Training labels (jet flavour from MC-truth heavy-quark content within ΔR<0.4, the five track-origin classes, and vertex clusters from truth positions at a 100 µm threshold) are independent of the reported performance metrics, and every efficiency, purity, and AUC value is computed on jets not used in training, so the numbers are genuine held-out measurements rather than fitted parameters renamed as predictions. The discriminant parameter fc = 0.018 and the loss weights α=1.5, β=0.5 are inherited from the external ATLAS GN1 note [17], not tuned to this paper's outcome; the classical comparison (SV efficiency 0.2–0.3, purity 0.3–0.5 for 20 < pT,jet < 70 GeV/c) is quoted from the independent ALICE publication [16] and not produced by this paper's pipeline, so the 'outperforms traditional tagging' claim rests on an external reference point. The benchmark is imperfectly matched (different simulation chain and detector model), but a mismatched baseline is a validity caveat, not circularity. The centrality-dependent performance loss and its mitigation are measured on held-out overlaid jets; the improvement from domain-matched training is an empirical magnitude, not enforced by construction. The paper itself flags its principal limitations in Sec. 2 ('For simplicity, all background particles were assumed to originate from the primary vertex with zero impact parameter') and Sec. 5 ('This approach does not incorporate medium-induced effects such as jet energy loss, deformation, or the presence of long-lived background particles that may produce secondary vertices'), explicitly conditioning the conclusions on the controlled overlay framework; these assumptions weaken the external validity of extrapolating to real Pb–Pb collisions but do not make the framework-relative results circular. There are no load-bearing self-citations: the authors' prior work is not invoked, and Refs [16] and [17] are external collaborations.
Assumptions & free parameters
free parameters (4)
- Discriminant mixing weight fc =
0.018
- Auxiliary loss weights alpha, beta =
alpha=1.5, beta=0.5
- GNN input track selection =
pT > 0.5 GeV/c, max 40 tracks
- Vertex clustering threshold =
100 um
assumptions (5)
- domain assumption Pythia8 and AMPT Monte Carlo generators provide an adequate description of pp jets and Pb-Pb underlying-event particles at sqrt(s_NN)=5.02 TeV.
- domain assumption Delphes fast simulation with ALICE-like tracking efficiency and resolution adequately approximates the ALICE detector response for the purpose of b-tagging.
- ad hoc to paper Overlaying AMPT background particles inside the jet cone onto Pythia pp jets is a valid proxy for the heavy-ion environment for evaluating tagger robustness.
- ad hoc to paper All background particles originate from the primary vertex with zero impact parameter.
- domain assumption Jet flavour can be assigned by the presence of heavy quarks within dR<0.4 of the particle-level jet axis.
Cite this review
Pith. "Pith review of Investigation of the performance of a GNN-based b-jet tagging method in heavy-ion collisions." pith.science (2026). https://pith.science/paper/DEGHUT2T
@misc{pith2026250622691,
author = {Pith},
title = {Pith review of: Investigation of the performance of a GNN-based b-jet tagging method in heavy-ion collisions},
year = {2026},
howpublished = {\url{https://pith.science/paper/DEGHUT2T}},
note = {Machine review of arXiv:2506.22691}
}
abstract
Beauty-tagged jets (b-jets)-collimated sprays of particles originating from the fragmentation of beauty quarks produced in the initial hard scatterings-provide a unique probe of parton dynamics in the quark-gluon plasma (QGP) created in ultrarelativistic heavy-ion collisions. In particular, energy loss patterns of low-$p_T$ b-jets traversing the QGP offer valuable insight into the strong interaction in its nonperturbative regime. CMS and ATLAS Collaborations at the LHC have studied b-jet production in Pb-Pb collisions. The results were limited to a high-$p_T$ region, because a major challenge at low-$p_T$ is the overwhelming number of background particles from QGP hadronisation, which severely hinders the effectiveness of conventional b-jet tagging techniques. To enable precise measurements in such complex environments, advanced tagging methods are required. Graph Neural Networks (GNNs), capable of learning relational structures among jet constituents, represent a promising deep learning approach for b-jet identification. In this study, we adopt and adapt the GN1 model, initially developed by ATLAS, for use in Pb-Pb collision environments. We investigate the model's performance by applying it to jets embedded with Pb-Pb background particles, evaluating both tagging decisions and robustness against background contamination. This work presents a comprehensive evaluation of GNN-based b-jet tagging under heavy-ion collision conditions, aiming to advance future precision studies of QGP-induced partonic energy loss.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[16]
ALICE Collaboration (ALICE) 2022 JHEP 2201 178 35 pages, 16 captioned fig- ures, 2 tables, authors from page 30, published version, figures at http://alice- publications.web.cern.ch/node/7411 ( Preprint 2110.06104) URL https://cds.cern.ch/ record/2783306
arXiv 2022
-
[1]
Busza W, Rajagopal K and van der Schee W 2018 Ann. Rev. Nucl. Part. Sci.68 339–376 (Preprint 1802.04801) Investigation of performance of a GNN-basedb-jet tagging method in heavy-ion collisions13
arXiv 2018
-
[2]
(Particle Data Group) 2024 Phys
Navas S et al. (Particle Data Group) 2024 Phys. Rev. D110 030001
work page 2024
-
[3]
Dong X, Lee Y J and Rapp R 2019 Ann. Rev. Nucl. Part. Sci.69 417–445 (Preprint 1903.07709)
arXiv 2019
-
[4]
Dokshitzer Y L and Kharzeev D E 2001 Phys. Lett. B519 199–206 (Preprint hep-ph/0106202)
arXiv 2001
-
[5]
ALICE Collaboration (ALICE) 2022 Nature 605 440–446 [Erratum: Nature 607, E22 (2022)] (Preprint 2106.05713)
arXiv 2022
-
[6]
Zigic D, Salom I, Auvinen J, Djordjevic M and Djordjevic M 2019 Phys. Lett. B 791 236–241 (Preprint 1805.04786)
arXiv 2019
-
[7]
Wicks S, Horowitz W, Djordjevic M and Gyulassy M 2007 Nucl. Phys. A784 426–442 (Preprint nucl-th/0512076)
arXiv 2007
Show all 29 references
-
[8]
Ke W, Xu Y and Bass S A 2018 Phys. Rev. C98 064901 (Preprint 1806.08848)
2018 arXiv
- [9]
-
[10]
CMS Collaboration (CMS Collaboration) 2014 Phys. Rev. Lett. 113(13) 132301 URL https: //link.aps.org/doi/10.1103/PhysRevLett.113.132301
2014 doi
-
[11]
ALICE Collaboration 2024 European Physical Journal C84 ISSN 1434-6044 publisher Copyright: © The Author(s) 2024
2024
-
[12]
ALICE Collaboration 2024 Journal of Instrumentation 19 P05062 URL https://dx.doi.org/ 10.1088/1748-0221/19/05/P05062
2024 doi
-
[13]
(CMS) 2013 JINST 8 P04013 (Preprint 1211.4462)
Chatrchyan S et al. (CMS) 2013 JINST 8 P04013 (Preprint 1211.4462)
2013 arXiv
-
[14]
Nguyen M 2013 Nuclear Physics A904-905 705c–708c ISSN 0375-9474 the Quark Matter 2012 URL https://www.sciencedirect.com/science/article/pii/S0375947413002388
2013
-
[15]
(ATLAS) 2016 JINST 11 P04008 (Preprint 1512.01094)
Aad G et al. (ATLAS) 2016 JINST 11 P04008 (Preprint 1512.01094)
2016 arXiv
-
[17]
ATLAS Collaboration (ATLAS) 2022 Neural Network Jet Flavour Tagging with the Upgraded ATLAS Inner Tracker Detector at the High-Luminosity LHC Tech. rep. CERN Geneva all figures including auxiliary figures are available at https://atlas.web.cern.ch/Atlas/GROUPS/PHYSICS/PUBNOTES...
2022
-
[18]
Brody S, Alon U and Yahav E 2021 ( Preprint 2105.14491)
2021 arXiv
-
[19]
Mondal S and Mastrolorenzo L 2024 Eur. Phys. J. ST233 2657–2686 (Preprint 2404.01071)
2024 arXiv
- [20]
-
[21]
Bierlich C, Chakraborty S, Desai N, Gellersen L, Helenius I, Ilten P, L¨ onnblad L, Mrenna S, Prestel S, Preuss C T, Sj¨ ostrand T, Skands P, Utheim M and Verheyen R 2022 A comprehensive guide to the physics and usage of pythia 8.3 ( Preprint 2203.11601) URL https://arxiv.org/...
2022 arXiv
-
[22]
de Favereau J, Delaere C, Demin P, Giammanco A, Lema ˆ ıtre V, Mertens A and Selvaggi M (DELPHES 3) 2014 JHEP 02 057 (Preprint 1307.6346)
2014 arXiv
-
[23]
(GEANT4) 2003 Nucl
Agostinelli S et al. (GEANT4) 2003 Nucl. Instrum. Meth. A506 250–303
2003
-
[24]
Cacciari M, Salam G P and Soyez G 2012 Eur. Phys. J. C72 1896 (Preprint 1111.6097)
2012 arXiv
-
[25]
Cacciari M, Salam G P and Soyez G 2008 Journal of High Energy Physics 2008 063 URL https://dx.doi.org/10.1088/1126-6708/2008/04/063
2008 doi
-
[26]
Lin Z W, Ko C M, Li B A, Zhang B and Pal S 2005 Phys. Rev. C 72 064901 ( Preprint nucl-th/0411110)
2005 arXiv
-
[27]
(ALICE) 2013 Phys
Abelev B et al. (ALICE) 2013 Phys. Rev. C88 044909 (Preprint 1301.4361)
2013 arXiv
-
[28]
Paszke A, Gross S, Massa F, Lerer A, Bradbury J, Chanan G, Killeen T, Lin Z, Gimelshein N, Antiga L, Desmaison A, K¨ opf A, Yang E, DeVito Z, Raison M, Tejani A, Chilamkurthy S, Steiner B, Fang L, Bai J and Chintala S 2019 Pytorch: An imperative style, high-performance deep le...
2019 arXiv
-
[29]
Falcon W and The PyTorch Lightning team 2019 PyTorch Lightning URL https://github.com/ Investigation of performance of a GNN-basedb-jet tagging method in heavy-ion collisions14 Lightning-AI/lightning
2019
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.