Pith. sign in

REVIEW 3 major objections 5 minor 13 references

Machine Learning based jet momentum reconstruction in Pb-Pb collisions measured with the ALICE detector

T0 review · 3 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read Machine-learning background correction extends jet measurements in Pb-Pb to record-low momenta and to R=0.6 jets.

desk verdict Solid proceedings with first Pb-Pb results from the ML jet correction; credible where checkable, but low-pT and R=0.6 claims rest on a training background and fragmentation systematic that are not fully closed. read the letter →

arxiv 1909.01639 v1 pith:YAWOXIO6 submitted 2019-09-04 nucl-ex hep-ex

classification nucl-exhep-ex
keywords jetreconstructionheavy-ioncollisionsmachinelearningbackgroundfluctuationstransversemomentumnuclearmodificationfactorPb-Pbradiusdependence
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that a machine-learning estimator, trained on simulated jets embedded in a thermal background, can correct the transverse momentum of each jet individually in Pb-Pb collisions, reducing the large background fluctuations that limit the standard area-based method. Using this jet-by-jet correction, the ALICE detector measures track-based jets at $\sqrt{s_\mathrm{NN}}=5.02$ TeV down to $40$ GeV/$c$ in 0-10% central collisions and $30$ GeV/$c$ in 30-50% central collisions for resolution parameter $R=0.4$, and reports the first measurement of $R=0.6$ jets in heavy-ion collisions at the LHC. The resulting nuclear modification factors are compatible with the area-based method and show no significant dependence on jet radius. The payoff of the approach is that it lowers the momentum threshold and extends the radius reach of jet measurements, opening a part of phase space that studies of jet quenching in heavy-ion collisions could not previously access.

What carries the argument

The machinery is a supervised regression model, in this analysis a neural network with three hidden layers of 100, 100, and 50 neurons, that takes a set of per-jet observables and outputs a single corrected jet $p_{\mathrm{T}}$. The input features are the jet $p_{\mathrm{T}}$ corrected by the standard area-based method, the jet angularity, the number of jet constituents, and the transverse momenta of the eight hardest constituents; the training target is the detector-level true jet momentum, approximated as the reconstructed jet momentum times the PYTHIA-carried momentum fraction. The trained model is applied jet by jet, and a response matrix built from embedding vacuum jets into real Pb-Pb background handles any residual smearing and detector effects before an unfolding step produces the final spectra.

What would settle it

Embed simulated events with a known true jet $p_{\mathrm{T}}$ but with a fragmentation pattern deliberately different from the PYTHIA training sample (for example, quark-only jets or jets with strong medium-induced energy loss), run the trained ML estimator on them, and test whether the residual between corrected and true jet momentum shifts systematically with jet $p_{\mathrm{T}}$ or radius $R$; a shift that the response-matrix and unfolding procedure cannot remove would show the estimator is not robust to the fragmentation assumption.

Watch

Extended reading notes

Core claim

The central discovery is that a neural-network regressor can learn the mapping from raw, background-contaminated jet observables to the true jet transverse momentum well enough to replace the statistical area-based background subtraction with a per-jet correction. The paper demonstrates this by training the model on PYTHIA jets reconstructed to detector level and embedded in a thermal background, with the regression target defined as the reconstructed jet momentum multiplied by the momentum fraction carried by PYTHIA particles in the jet. The trained estimator is then applied to Pb-Pb data, and the residual fluctuations are removed by unfolding through a response matrix built from embedding vacuum jets into real data backgrounds. The result is that track-based jets are measured down to $p_{\mathrm{T}}=40$ GeV/$c$ (0-10% central) and $30$ GeV/$c$ (30-50%) for $R=0.4$, and $R=0.6$ jets are measured for the first time in heavy-ion collisions at the LHC, with spectra and nuclear modification factors compatible with the area-based method and with no significant $R$ dependence.

Load-bearing premise

The method's reliability rests on the assumption that jets embedded in a simplified thermal background, trained with the PYTHIA momentum fraction as the target, represent the true relationship between the measured jet features and the actual jet momentum in real Pb-Pb collisions; if heavy-ion jet fragmentation or background structure differs from this training model in ways the embedding tests do not capture, the corrected jet spectra will be biased.

Editorial extensions

If this is right

  • Jets with resolution parameter $R=0.6$ can be measured in Pb-Pb collisions at the LHC for the first time, extending jet quenching studies to larger angular scales.
  • Track-based jets at $R=0.4$ become measurable down to $40$ GeV/$c$ in 0-10% and $30$ GeV/$c$ in 30-50% central Pb-Pb collisions, a much larger low-$p_{\mathrm{T}}$ reach than the area-based method.
  • The ML-corrected nuclear modification factors agree with the area-based estimator where both are available, confirming the method's consistency with an established baseline.
  • No significant $R$ dependence of the nuclear modification factor is observed between $R=0.2$, $0.4$, and $0.6$, and the jet cross-section ratios are consistent with the pp baseline.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because the estimator is trained on the PYTHIA momentum fraction, its jet-by-jet correction inherits the fragmentation model; training instead on a mixture of quark- and gluon-initiated jets or on jets with varied energy loss could make the correction less model-dependent.
  • The same regression approach could be pushed to $R=0.8$ or $0.9$ within the ALICE acceptance, exactly the regime where the area-based method's background fluctuations are most severe and where the ML method's advantage would be largest.
  • The jet-by-jet idea is not limited to jet $p_{\mathrm{T}}$: similar features could estimate groomed jet momentum, subjet momentum, or the background under a jet, which would extend the method to jet-structure observables in heavy-ion collisions.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. This conference proceeding reports a machine-learning-based jet momentum correction for Pb-Pb collisions at sqrt(s_NN)=5.02 TeV, applied to track-based jets with R=0.2, 0.4, and 0.6. A neural network is trained on PYTHIA jets embedded in a thermal background to map raw jet features to an estimate of the true detector-level jet pT on a jet-by-jet basis. The authors present nuclear modification factors for R=0.4 (down to 40 GeV/c in 0-10% and 30 GeV/c in 30-50% central collisions) and, for the first time, for R=0.6, together with jet cross-section ratios sigma(R=0.2)/sigma(R=0.4) and sigma(R=0.2)/sigma(R=0.6). They claim compatibility with the area-based estimator at higher pT, no significant R-dependence of R_AA, and no centrality or pp deviation in the cross-section ratios.

Significance. If the ML correction is unbiased, the method would extend ALICE jet measurements to lower pT and larger R than the standard area-based method, and the R=0.6 result would be a first in heavy-ion collisions at the LHC. The paper contains genuine validation steps: an embedding-based comparison of residual widths against the area-based estimator (Fig. 1) and a direct comparison of R_AA at R=0.4 in two centrality classes (Fig. 2). The work is a proceedings contribution and the underlying method is documented in a separate paper (ref. [4]), which is appropriately cited. However, the new physics claims (low-pT reach and first R=0.6 measurement) rest on the untested assumption that a model trained on a simplified thermal background generalizes to real Pb-Pb events; the evidence presented in this manuscript is not sufficient to establish that. The central results are therefore plausible but not yet convincingly validated.

major comments (3)
  1. [Sec. 3.1 and 3.3] The training background in Sec. 3.1 uses an exponential momentum tail above roughly 4 GeV/c, whereas real Pb-Pb events contain a power-law tail of high-pT particles from semi-hard and hard processes. Because the input features include the transverse momenta of the first eight leading jet constituents, a high-pT background or underlying-event particle in a real event could be misidentified by the estimator as part of the signal jet, producing a positive bias in the corrected pT. The embedding validation in Sec. 3.3 reports only the width of residual distributions; it does not show that the mean residual is consistent with zero, especially in the lowest pT bins or for R=0.6. To support the advertised low-pT reach, the authors should provide a closure test with a realistic background model that includes a power-law component, or at least report the mean residual as a function of pT for each R and centrality.
  2. [Sec. 4, fragmentation systematic] The fragmentation systematic variation replaces only the response matrix used in unfolding with a quark-jet response. It does not retrain or re-evaluate the ML estimator itself under altered fragmentation, even though the estimator is trained on PYTHIA jets and its input-output mapping is likely sensitive to jet fragmentation. Changing only the unfolding response cannot capture a fragmentation-dependent bias in the estimator itself. The authors should either retrain the estimator on samples with modified fragmentation and compare the final corrected spectra, or provide an argument based on the estimator's features as to why it is insensitive to fragmentation.
  3. [Sec. 4, Figs. 2-4] The compatibility with the area-based method shown in Fig. 2 is established only in the pT region where the area-based method is already reliable. The new claims -- jets at 40 (30) GeV/c for R=0.4 and the R=0.6 measurement -- lie precisely in the region where no such cross-check exists. The cross-section ratios in Fig. 4 and the R=0.6 R_AA in Fig. 3 therefore rest entirely on the ML estimator's extrapolation. To support these claims, the authors should provide an alternative validation in this region, for example a closure test on full heavy-ion simulations with realistic underlying events, or a comparison with an independent background subtraction technique such as constituent subtraction or soft-drop grooming.
minor comments (5)
  1. [Abstract and Sec. 1] The abstract and introduction state that transverse momentum spectra will be presented, but the figures show only nuclear modification factors and cross-section ratios; the spectra themselves are not displayed. Please clarify what is shown or add the spectra.
  2. [Fig. 1] The caption of Fig. 1 is incomplete: the left panel is described as 'the comparison of the different background estimators for R=0.4' but the reader cannot identify which curves correspond to which estimator or centrality. The right panel lacks axis labels and a legend. Please improve the figure captions and labels.
  3. [Sec. 4] The systematic uncertainties are described only qualitatively; no numeric values are given for any observable. For a measurement-oriented proceeding, at least the dominant systematic uncertainties at representative pT values should be stated.
  4. [Sec. 3.2] The regression target is defined as the reconstructed jet momentum multiplied by the momentum fraction carried by PYTHIA particles in the jet. This definition is confusing and should be clarified: it appears to approximate the true detector-level jet momentum, but the potential bias introduced by this target choice is not discussed. Please rephrase and add a sentence on the validity of this approximation.
  5. [Sec. 4, Eq. (4.1)] Equation (4.1) is not fully typeset: the differentials and symbols are inconsistently formatted (e.g., d2Nch jet/dpT,ch jetdηjet). Also, the definition T_AA = N_coll/sigma_tot is a rough shorthand; please use the standard definition and cite a reference.

Circularity Check

0 steps flagged · score 2.0 of 10

No circular derivation: the ML estimator is trained on simulation, applied to real Pb-Pb events, and the final spectra are checked against the area-based method and pp reference; the only self-citation is the method paper [4], which is not load-bearing for the new data results.

full rationale

The central chain is not circular. In Section 3.2 the regression target is defined as 'the reconstructed jet momentum multiplied by the momentum fraction that is carried by PYTHIA particles in the jet'; this is a simulation-truth training label, not the final measured observable. The trained estimator is then applied to real Pb-Pb data, and the resulting spectra are unfolded with a PYTHIA-based response matrix and compared, in Figs. 2-4, to the area-based estimator, to pp data from [7], and to Hybrid Model calculations. None of the final observables (spectra, R_AA, cross-section ratios) is equal by construction to the training target or to the ML input features. The method paper [4] is cited as the source of the ML technique and its first author overlaps with the present paper, but the proceeding independently describes the training data, input parameters, and validation, and the data results are new; this is a normal method reference rather than a load-bearing self-citation. The fragmentation limitation in Section 4 - the uncertainty is estimated only by replacing the unfolding response matrix with a quark-jet response, without retraining the ML estimator - is a genuine model-robustness weakness, but it is a validity concern, not circularity, because the measurement does not reduce to the training simulation by definition. Likewise, the exponentially falling high-pT tail of the thermal toy background in Section 3.1 is an assumption about the training distribution, not an equation that forces the final answer.

Assumptions & free parameters 6 free parameters · 6 assumptions · 0 invented entities

The measurement rests on the fidelity of the PYTHIA+GEANT simulation chain used for training and response, on the thermal background model for training, and on the definition of the regression target. These are reasonable within the field but are not independently validated here beyond embedding checks.

free parameters (6)
  • Neural network architecture = three hidden layers: 100, 100, 50 neurons
    Model hyperparameters tuned on toy model in the method paper [4]; the regression mapping depends on this choice.
  • Track pT threshold = 0.150 GeV/c
    Choice of minimum constituent pT affects both jet finding and ML input features; hand-selected analysis cut.
  • Jet area cut = Ajet > 0.557 pi R^2
    Rejects jets not fully contained; hand-selected and affects selected jet sample.
  • Thermal background multiplicity range = flat 0 to 3000 tracks
    Training background is a toy model with flat multiplicity; the true Pb-Pb multiplicity distribution is not used.
  • Unfolding regularization parameter = not stated
    Varied as a systematic uncertainty in Section 4; the central value is not given in this proceeding.
  • Low pT reach cut = 40 GeV/c (0-10%), 30 GeV/c (30-50%) for R=0.4
    The reported measurement range is chosen based on where results are considered reliable; effectively a hand-set boundary.
assumptions (6)
  • domain assumption PYTHIA 8 with GEANT3 simulates jet fragmentation, hadronization, and detector response accurately enough for training and response matrices.
    Used in Sections 2 and 3.1 for the training data and in the response matrix embedding; no independent validation against data in this paper.
  • ad hoc to paper The thermal background toy model (flat multiplicity 0-3000 tracks, thermal pT spectrum, exponential tail above ~4 GeV/c) is representative of the real Pb-Pb underlying event for the purpose of training the ML estimator.
    Section 3.1; the ML model is trained on this toy background rather than on real Pb-Pb events.
  • ad hoc to paper The regression target, defined as reconstructed jet pT times the PYTHIA particle momentum fraction in the jet, is a valid approximation of the true detector-level jet momentum.
    Section 3.2; this definition is specific to the method and carries PYTHIA fragmentation assumptions into the correction.
  • domain assumption A supervised ML model trained on embedded PYTHIA jets generalizes to real Pb-Pb jets in the feature space used.
    This is the transfer assumption at the core of the method; validated indirectly by embedding probes into real events and by compatibility with the area-based method.
  • domain assumption The measured pp reference from [7] and the TAA centrality scaling from [11] are valid for computing R_AA.
    Standard inputs in heavy-ion jet measurements; taken from cited ALICE results.
  • standard math Standard jet clustering algorithms (anti-kT, kT) and the area-based background subtraction provide a valid baseline.
    Used in Sections 2 and 3; standard methods in the field.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Machine Learning based jet momentum reconstruction in Pb-Pb collisions measured with the ALICE detector." pith.science (2026). https://pith.science/paper/YAWOXIO6

@misc{pith2026190901639,
  author       = {Pith},
  title        = {Pith review of: Machine Learning based jet momentum reconstruction in Pb-Pb collisions measured with the ALICE detector},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/YAWOXIO6}},
  note         = {Machine review of arXiv:1909.01639}
}
abstract

The precise reconstruction of jet transverse momenta in heavy-ion collisions is a challenging task. A major obstacle is the large number of uncorrelated (mainly) low-$p_\mathrm{T}$ particles overlaying the jets. Strong region-to-region fluctuations of this background complicate the jet measurement and lead to significant uncertainties. We developed a novel approach to correct jet momenta (or energies) for the underlying background in heavy-ion collisions. The approach allows the measurement of jets down to extremely low transverse momenta and for large resolution $R$ by making use of common Machine Learning techniques to estimate the jet transverse momentum based on several parameters. In this conference proceeding, we will present transverse momentum spectra and nuclear modification factors of track-based jets that have been corrected by this Machine Learning approach and comparisons to published results where possible. The analysis was performed on Pb-Pb collisions at $\sqrt{s_\mathrm{NN}} = 5.02$ TeV recorded with the ALICE detector and measures jets with large resolution parameters for low momenta, unprecedented thus far in data on heavy-ion collisions.

Figures

Figures reproduced from arXiv: 1909.01639 by the authors.

Figure 1
Figure 1. Residual pT-distributions of embedded jet probes of known transverse momentum. 3 [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Nuclear modification factor for R = 0.4, comparing spectra corrected with either ML-based or area-based estimator. 0-10% (left) and 30-50% (right) most central collisions [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Nuclear modification factor for R = 0.4 and R = 0.6 for 0-10% (left) and 30-50% (right). 5. Conclusions In this paper, we presented transverse momentum spectra, nuclear modification factors, and cross-section ratios of track-based jets in Pb–Pb collisions at √ sNN = 5.02 TeV that have been corrected by our novel Machine-Learning-based background correction approach. Thanks to the new background estimation method, je… view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Jet cross-section ratios for R = 0.2/R = 0.4 (left) and R = 0.2/R = 0.6 (right). References [1] ALICE Collaboration: The ALICE experiment at the CERN LHC, JINST 3 (2008) S08002. [2] ALICE Collaboration: Measurement of event background fluctuations for charged particle …

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

13 extracted references · 10 canonical work pages

  1. [4]

    Haake, C

    R. Haake, C. Loizides: Machine-learning-based jet momentum reconstruction in heavy-ion collisions, Phys. Rev. C 99, 064904 (2019), cf. [nucl-ex/1810.06324]

  2. [1]

    ALICE Collaboration: The ALICE experiment at the CERN LHC, JINST 3 (2008) S08002

  3. [2]

    ALICE Collaboration: Measurement of event background fluctuations for charged particle jet reconstruction in Pb–Pb collisions at√sNN = 2.76 TeV , JHEP 03 (2012) 053

  4. [3]

    ALICE Collaboration: Measurement of charged jet suppression in Pb–Pb collisions at √sNN = 2.76 TeV , JHEP 03 (2014) 013

  5. [5]

    Sjostrand et al.: PYTHIA 6.4 – Physics and Manual, JHEP 0605 (2006) 026

    T. Sjostrand et al.: PYTHIA 6.4 – Physics and Manual, JHEP 0605 (2006) 026

  6. [6]

    Brun et al.: GEANT Detector Description and Simulation Tool, CERN Program Library Long Writeup CERN-W-5013 (1994)

    R. Brun et al.: GEANT Detector Description and Simulation Tool, CERN Program Library Long Writeup CERN-W-5013 (1994)

  7. [7]

    [nucl-ex/1905.02536]

    ALICE Collaboration: Measurement of charged jet cross section in pp collisions at √sNN = 5.02 TeV , submitted to PRD, cf. [nucl-ex/1905.02536]

  8. [8]

    Cacciari, G.P

    M. Cacciari, G.P. Salam: Dispelling the N3 myth for the kt jet-finder, Phys. Lett. B 641, 57-61 (2006)

Show all 13 references
  1. [9]

    Cacciari, G.P

    M. Cacciari, G.P. Salam, and G. Soyez: The anti- kT jet clustering algorithm, JHEP 0804 (2008) 063

  2. [10]

    Pedregosa et al.: Scikit-learn: Machine learning in Python, Journal of Machine Learning Research 12(2011) 2825–2830

    F. Pedregosa et al.: Scikit-learn: Machine learning in Python, Journal of Machine Learning Research 12(2011) 2825–2830

  3. [11]

    ALICE Collaboration: Centrality determination in heavy-ion collisions, Public note

  4. [12]

    Casalderrey-Solana et al., A hybrid strong/weak coupling approach to jet quenching, JHEP 10 (2014) 019

    J. Casalderrey-Solana et al., A hybrid strong/weak coupling approach to jet quenching, JHEP 10 (2014) 019

  5. [13]

    [hep-ph/1907.12301]

    Daniel Pablos, Jet suppression from small to intermediate to large radius, cf. [hep-ph/1907.12301]. 6

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.