REVIEW 2 major objections 5 minor 1 cited by
Run 3 performance and advances in heavy-flavor jet tagging in CMS
T0 review · 2 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read UParT, a transformer-based tagger, now gives CMS its best charm-jet identification in Run 3, and flavor-enriched scale factors bring data and simulation into agreement for heavy-flavor and boosted jets.
desk verdict A useful, honest proceedings summary of CMS Run 3 heavy-flavor tagging; calibration closure is in-sample, but that is a presentation gap, not a fatal flaw. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Two mechanisms carry the argument. The first is UParT itself: a transformer tagger in which a multi-head attention mechanism consumes pairwise interaction features between every jet constituent and every secondary vertex, and adversarial training on distorted input features keeps the model from latching onto simulation-specific details; UParT also returns jet energy regression and resolution estimates alongside flavor probabilities. The second is the calibration machinery: scale factors derived by comparing data and simulation for flavor-enriched selections, applied either per working point or as shape corrections to the full discriminant distributions. For boosted jets, the load-bearing object is the ParticleNetMD discriminant, combined with soft-drop mass in a simultaneous likelihood fit over score regions.
What would settle it
Take the scale factors described in the paper and apply them to a flavor-enriched sample outside the calibration selections, for example $Z\to b\bar{b}$ or $Z\to c\bar{c}$ events from the same 2022-2023 run, and compare the tagger discriminant and soft-drop mass distributions to simulation: if the data-to-simulation ratio departs from unity by more than the quoted systematic uncertainties in any kinematic bin, the transferability assumption fails.
Extended reading notes
Core claim
On its own terms, the paper's central result is that the latest tagger generation works better and can be calibrated. UParT, built on the ParticleTransformer architecture for AK4 jets, outperforms all previous CMS taggers in c-jet identification and reduces light-jet or c-jet mistagging at fixed b-jet efficiency; it also introduces the first s-jet classifier in CMS and a hadronic-tau classifier. The calibration study shows that scale factors computed from b-enriched top-pair and c-enriched W+jets samples correct the main data-versus-simulation shifts in the BvsAll, CvsL, and CvsB discriminators for 2022–2023 data, with post-correction agreement visibly better than before. For merged jets, the paper reports that ParticleNetMD is the best of the Run 2 boosted taggers, and its Run 3 validation in $Z\to b\bar{b}$ events reproduces the Z boson mass peak after a likelihood fit, which the paper reads as confirmation that the tagger works in data.
Load-bearing premise
The load-bearing premise is that the correction factors derived from top-pair dilepton events for bottom jets and W+jets events for charm jets work everywhere else in Run 3; the paper only shows post-correction agreement in those enriched samples, not in other kinematic regions.
Editorial extensions
If this is right
- Run 3 CMS analyses can use UParT discriminators with better light- and charm-jet rejection than previous taggers, reducing backgrounds in top and Higgs measurements.
- Calibrated scale factors from top-pair and W+jets selections make b- and c-tagging usable in data, though the paper notes some performance loss after calibration.
- UParT's s-jet and hadronic-tau classifiers open new tagging channels, including a low-efficiency s-tagger that reaches 20% efficiency at a $10^{-3}$ misidentification probability.
- ParticleNetMD's Run 3 $Z\to b\bar{b}$ validation, with a visible Z mass peak after the fit, supports boosted analyses such as $H\to b\bar{b}$ and $H\to c\bar{c}$.
- The b-hive and BTVNanoCommissioning frameworks let the training and commissioning cycle run automatically, which keeps up with changing detector conditions in Run 3.
Reading between the lines
- If adversarial training makes UParT insensitive to input distortions, future taggers may need fewer retraining cycles after detector or simulation changes; the paper does not quantify how far this transfers.
- A working s-jet tagger would give LHC analyses a new probe of strange-quark physics, such as strange-quark Yukawa or fragmentation measurements, but only simulated ROC performance is reported here.
- The calibration claim is the most natural point to stress-test: applying the same corrections in a very different regime, for example very high transverse momentum or boosted topologies, is a testable extension rather than something the paper demonstrates.
- UParT's per-jet flavor probabilities and ParticleNetMD's boosted-jet discriminant could eventually be combined into one unified resolved-plus-boosted tagger, but the paper stops short of that.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This ICHEP 2024 conference proceeding summarizes recent CMS heavy-flavor jet tagging results: the evolution of taggers from CSV to UParT, the UParT architecture and its simulated performance (including first s-jet tagging and hadronic tau classification), data-to-simulation comparisons and scale-factor calibration using 2022-2023 Run 3 data, boosted-jet tagging for H to bb and H to cc, and the b-hive and BTVNanoCommissioning frameworks. The central claims are that UParT provides the best c-jet tagging performance in CMS so far and that the derived scale factors bring data and simulation into agreement, making the calibrated taggers usable for Run 3 physics analyses.
Significance. If the claims hold, UParT represents a clear advance in c-jet identification and robust transformer-based tagging, and the described calibration framework would be directly useful for Run 3 analyses. The paper is a readable and generally accurate summary of the current CMS program, and it appropriately cites the official CMS detector performance summaries. The paper's main quantitative support, however, is limited: the ROC curves in Figures 2 and 3 are shown without uncertainty bands, and the calibration demonstration in Figure 5 is largely an in-sample consistency check without independent closure. The strengths are the clear architectural description of UParT, the public availability of the b-hive and BTVNanoCommissioning frameworks, and the explicit references to the underlying CMS DP notes.
major comments (2)
- [Section 3, Figures 2 and 3] The headline claim that UParT is 'the most performing model so far in c-jet identification' rests on ROC curves that are shown without any statistical or systematic uncertainty bands. Because the comparison is made on a single simulated ttbar sample and no uncertainties are given, the reader cannot judge whether the apparent improvement over ParticleNet is significant. Please add uncertainties or, at a minimum, state explicitly that the conclusion is qualitative and refer the reader to the official CMS detector performance summary for the quantitative comparison.
- [Section 5, Figure 5] The demonstration that scale factors bring data and simulation into agreement is an in-sample consistency check for b tagging: the left panel uses the e-mu+jets ttbar phase space, which is the same phase space typically used to derive b-tagging scale factors. The right panel provides a hadronic-ttbar cross-check, but no quantitative closure metric or uncertainty budget is quoted, and the c-tagging scale factors derived from W+jets are not validated against an independent c-enriched selection. Consequently, the extrapolation of these scale factors to arbitrary Run 3 analyses is not established by the evidence shown; please add an independent closure test or explicitly restrict the claim to the validated phase space and cite the official CMS calibration note.
minor comments (5)
- [Section 2 and Figure 2 caption] The label 'BvsC' appears in the Figure 2 caption but is not defined among the discriminators in Section 2; either define it or use the already-defined 'CvsB' consistently.
- [Section 4 and Figure 4] The text and figure caption use both 'BvsAll' and 'BvAll' for the same discriminant; please use a single notation throughout.
- [Section 6, Figure 8 caption] The caption says 'under the 2018 data-taking conditions' while the surrounding text describes a Run-3 validation; clarify that the left panel is Run 2 and the right panel is Run 3.
- [Sections 2, 5, and 8] There are several typographical errors, including 'Multi-Layer-Percrepton', 'the the light-jet misidentification rate', and 'utilisessophisticated machine learning techniques'; these should be corrected in a final proofreading pass.
- [Section 3 and Figure 3] The claim that UParT provides the 'first attempt' at s-quark jet tagging in CMS is stated without a citation; if this is the first public result, please add a reference to the relevant CMS DP note or CMS-PAS.
Circularity Check
Core UParT performance comparisons are independent, but the Section 5 post-calibration agreement is an in-sample fit shown as validation.
-
fitted input called prediction
[Section 5, Figure 5 and surrounding text]
"Scale Factors (SFs) are introduced to correct discrepancies between the tagging efficiencies observed in data and those predicted by simulations. These SFs are derived by comparing the performance of tagging algorithms on data and simulations for various jet flavors... The comparison between pre- and post-calibration distributions is shown in figure 5. After applying the SFs, the agreement between data and simulation improves considerably, mitigating the effects of mismodeling."
The SFs are fitted to data/MC differences in flavor-enriched selections. Figure 5 (left) then demonstrates improved agreement in the e-mu+jets phase space, which is the same ttbar-dilepton topology used to derive b-tagging SFs. The post-SF improvement is therefore the in-sample result of the fit, not an independent validation; a shape-based SF that successfully reduces the residual would produce a better ratio by construction. The right panel uses hadronic ttbar as a cross-check but quotes no closure metric or uncertainty budget, and the c-tagging SFs from W+jets are not validated against an independent c-enriched selection. Thus the claim that calibrated taggers are usable for Run 3 physics is supported mainly by an in-sample calibration check.
full rationale
The paper is a performance summary, not a derivation-based analysis, so most claimed results are empirical comparisons on simulation. The central UParT performance claim (Figures 2 and 3) is an ROC comparison against prior taggers on simulated samples; this is an independent algorithmic benchmark and is not circular. The boosted-jet ROC curves and the Z->bb post-fit mass peak are similarly externally meaningful. The only genuinely in-sample circularity is in Section 5: SFs are derived from data/MC comparisons, and the post-SF agreement shown in Figure 5 (left) is displayed in the same flavor-enriched phase space used for the SF derivation, so the improvement is a property of the fitted SFs rather than a predictive validation. The hadronic-ttbar right panel is a partial cross-check but no quantitative closure statistic or uncertainty is provided. Because the main tagging-performance claims do not depend on this calibration step and retain independent content, the overall circularity is limited to this one fitted-input-as-validation step.
Assumptions & free parameters
free parameters (1)
- Flavor tagging scale factors (b, c, light) =
Not specified in this proceeding
assumptions (2)
- domain assumption Residual simulation mismodeling can be corrected by per-flavor, per-phase-space scale factors derived from data.
- domain assumption Flavor-enriched calibration selections (ttbar for b, W+jets for c) are pure enough that the derived scale factors are not dominated by background contamination.
Cite this review
Pith. "Pith review of Run 3 performance and advances in heavy-flavor jet tagging in CMS." pith.science (2026). https://pith.science/paper/DUJNRHAP
@misc{pith2026241205863,
author = {Pith},
title = {Pith review of: Run 3 performance and advances in heavy-flavor jet tagging in CMS},
year = {2026},
howpublished = {\url{https://pith.science/paper/DUJNRHAP}},
note = {Machine review of arXiv:2412.05863}
}
read the original abstract
Identification of hadronic jets originating from heavy-flavor quarks is extremely important to several physics analyses in High Energy Physics, such as studies of the properties of the top quark and the Higgs boson, and searches for new physics. Recent algorithms used in the CMS experiment were developed using state-of-the-art machine-learning techniques to distinguish jets emerging from the decay of heavy flavour (charm and bottom) quarks from those arising from light-flavor (udsg) ones. Increasingly complex deep neural network architectures, such as graphs and transformers, have helped achieve unprecedented accuracies in jet tagging. New advances in tagging algorithms, along with new calibration methods using flavour-enriched selections of proton-proton collision events, allow us to estimate flavour tagging performances with the CMS detector during early Run 3 of the LHC.
Figures
Figures from the paper (7 more)
Forward citations
Cited by 1 Pith paper
-
Searching for Di-Higgs Signatures of Light Charged Scalars
A roughly 130 GeV charged Higgs explaining the ATLAS t to bH+ excess and B-meson anomalies can be constrained, and perhaps discovered, by recasting SM di-Higgs 4b searches at the LHC.
Reference graph
Works this paper leans on
-
[4]
AunifiedapproachforjettagginginRun3at √𝑠=13.6TeVinCMS
CMSCollaboration. AunifiedapproachforjettagginginRun3at √𝑠=13.6TeVinCMS. CMS Detector Performance Summary CMS-DP-2024-066, 2024. URLhttps://cds.cern.ch/ record/2904702
arXiv 2024
-
[1]
The CMS Experiment at the CERN LHC.JINST, 3:S08004, 2008
CMS Collaboration. The CMS Experiment at the CERN LHC.JINST, 3:S08004, 2008. doi: 10.1088/1748-0221/3/08/S08004
-
[2]
A.M.Sirunyanetal. Identificationofheavy-flavourjetswiththeCMSdetectorinppcollisions at 13 TeV.JINST, 13(05):P05011, 2018. doi: 10.1088/1748-0221/13/05/P05011
-
[3]
CMS collaboration. Identification of b-quark jets with the CMS experiment.Journal of Instrumentation, 8(04):P04013, apr 2013. doi: 10.1088/1748-0221/8/04/P04013. URL https://dx.doi.org/10.1088/1748-0221/8/04/P04013
-
[5]
CMS collaboration. Development of the CMS detector for the CERN LHC run 3.Journal of Instrumentation, 19(05):P05064, May 2024. doi: 10.1088/1748-0221/19/05/P05064. URL https://dx.doi.org/10.1088/1748-0221/19/05/P05064
-
[6]
A Combined Secondary Vertex Based B-Tagging Algorithm in CMS
Christian Weiser. A Combined Secondary Vertex Based B-Tagging Algorithm in CMS. Technical report, CERN, Geneva, 2006. URLhttp://cds.cern.ch/record/927399
work page 2006
-
[7]
Jet flavor classification in high-energy physics with deep neural networks.Phys
Daniel Guest, Julian Collado, Pierre Baldi, Shih-Chieh Hsu, Gregor Urban, and Daniel Whiteson. Jet flavor classification in high-energy physics with deep neural networks.Phys. 11 Heavy-flavor jet tagging in CMS Uttiya Sarkar on behalf of the CMS collaboration Rev. D, 94:112002, Dec 2016. doi: 10.1103/PhysRevD.94.112002. URLhttps://link. aps.org/doi/10.110...
-
[8]
E. Bols, J. Kieseler, M. Verzetti, M. Stoye, and A. Stakia. Jet flavour classification using deepjet. Journal of Instrumentation, 15(12):P12012, dec 2020. doi: 10.1088/1748-0221/15/ 12/P12012. URL https://dx.doi.org/10.1088/1748-0221/15/12/P12012
Show all 21 references
-
[9]
ParticleNet: Jet Tagging via Particle Clouds.Phys
Huilin Qu and Loukas Gouskos. ParticleNet: Jet Tagging via Particle Clouds.Phys. Rev. D, 101(5):056019, 2020. doi: 10.1103/PhysRevD.101.056019
2020 doi
-
[10]
Improving Robustness of Jet Tagging Algorithms with Adversarial Training.Comput
Annika Stein, Xavier Coubez, Spandan Mondal, Andrzej Novak, and Alexander Schmidt. Improving Robustness of Jet Tagging Algorithms with Adversarial Training.Comput. Softw. Big Sci., 6(1):15, 2022. doi: 10.1007/s41781-022-00087-1
2022 doi
-
[11]
Transformer models for heavy flavor jet identification
CMS Collaboration. Transformer models for heavy flavor jet identification. CMS Detector Performance Summary CMS-DP-2022-050, 2022. URLhttps://cds.cern.ch/record/ 2839920
2022
-
[12]
Particle Transformer for Jet Tagging.arXiv, 2022
Huilin Qu, Congqiao Li, and Sitian Qian. Particle Transformer for Jet Tagging.arXiv, 2022. URL https://arxiv.org/abs/2202.03772
2022 arXiv
-
[13]
Run3commissioningresultsofheavy-flavorjettaggingat √𝑠 =13.6TeV with CMS data using a modern framework for data processing
CMSCollaboration. Run3commissioningresultsofheavy-flavorjettaggingat √𝑠 =13.6TeV with CMS data using a modern framework for data processing. CMS Detector Performance Summary CMS-DP-2024-024, 2024. URLhttps://cds.cern.ch/record/2898463
2024
-
[14]
PerformancesummaryofAK4jetbtaggingwithdatafrom2022proton- proton collisions at 13.6 TeV with the CMS detector
CMSCollaboration. PerformancesummaryofAK4jetbtaggingwithdatafrom2022proton- proton collisions at 13.6 TeV with the CMS detector. CMS Detector Performance Summary CMS-DP-2024-025, 2024. URLhttps://cds.cern.ch/record/2898464
2024
-
[15]
Performance of heavy-flavour jet identification in boosted topologies in proton-proton collisions at√𝑠 = 13 TeV
CMS Collaboration. Performance of heavy-flavour jet identification in boosted topologies in proton-proton collisions at√𝑠 = 13 TeV. CMS Detector Performance Summary CMS-PAS- BTV-22-001, CERN, Geneva, 2023. URLhttps://cds.cern.ch/record/2866276
2023
-
[16]
Performance of boosted𝑏 ¯𝑏 jet tagging at√𝑠 = 13.6 TeV with Run 3 CMS data
CMS Collaboration. Performance of boosted𝑏 ¯𝑏 jet tagging at√𝑠 = 13.6 TeV with Run 3 CMS data. CMS Detector Performance Summary CMS-DP-2024-055, 2024. URLhttps: //cds.cern.ch/record/2904691
2024
-
[17]
Larkoski, Simone Marzani, Gregory Soyez, and Jesse Thaler
Andrew J. Larkoski, Simone Marzani, Gregory Soyez, and Jesse Thaler. Soft Drop.JHEP, 05:146, 2014. doi: 10.1007/JHEP05(2014)146
2014 doi
-
[18]
b-hive: a modular training framework for state-of-the-art object-tagging within the Python ecosystem at the CMS experiment
CMS Collaboration. b-hive: a modular training framework for state-of-the-art object-tagging within the Python ecosystem at the CMS experiment. CMS Detector Performance Summary CMS-DP-2024-020, 2024. URLhttps://cds.cern.ch/record/2896100
2024
-
[19]
root-project/root: v6.18/02, June 2020
Rene Brun and Fons Rademakers. root-project/root: v6.18/02, June 2020. URLhttps: //doi.org/10.5281/zenodo.3895860. 12 Heavy-flavor jet tagging in CMS Uttiya Sarkar on behalf of the CMS collaboration
2020 doi
-
[20]
coffea, April 2024
Lindsey Gray, Nicholas Smith, Andrzej Novak, Peter Fackeldey, Benjamin Tovar, Yi-Mu Chen, Gordon Watts, and Iason Krommydas. coffea, April 2024. URLhttps://doi.org/ 10.5281/zenodo.10977418
2024 doi
-
[21]
scikit-hep/awkward-array: 0.13.0, July 2020
JimPivarski,CharlesEscott,NicholasSmith,MichaelHedges,MasonProffitt,CharlieEscott, JaydeepNandi,JonasRembser,bfis, benkrikler, LindseyGray, DougDavis, HenrySchreiner, Nollde, Peter Fackeldey, and Pratyush Das. scikit-hep/awkward-array: 0.13.0, July 2020. URL https://doi.org/10...
2020 doi
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.