REVIEW 2 major objections 5 minor 14 references
A Particle Transformer jet tagger cuts heavy-flavor background acceptance by a factor of 5–10 for Higgs-factory detectors.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-13 22:22 UTC pith:IAO7FIQQ
load-bearing objection Solid ILD application of ParT with real multi-category and PID results; the 5–10× vs LCFIPlus is useful but not same-sample controlled, and full-sim vs SGV gaps are still open. the 2 major comments →
Jet flavor tagging with Particle Transformer for Higgs factories
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
When Particle Transformer is trained on ILD full-simulation jets (and on larger SGV fast-simulation samples) that include per-particle impact parameters, kinematics and Comprehensive PID outputs, background acceptance for b- and c-tagging drops by a factor of approximately 5–10 relative to the conventional LCFIPlus BDT taggers, reaching O(10⁻³) charm and O(10⁻³)–O(10⁻⁴) light-flavor contamination at 80 percent signal efficiency; the same architecture also yields workable strange tagging and quark–antiquark discrimination for heavy flavors.
What carries the argument
Particle Transformer (ParT): a set-to-set Transformer that embeds each reconstructed particle, injects pairwise kinematic interaction biases into the attention weights, and classifies the whole jet into 3, 6 or 11 flavor categories.
Load-bearing premise
The quoted factor-of-5–10 improvement remains meaningful even though absolute performance differs non-negligibly between full simulation and the fast-simulation samples used for the largest statistics, and those differences are still under investigation.
What would settle it
A side-by-side evaluation of the identical ParT model and the previous BDT tagger on the same fully simulated ILD sample, with identical working-point definitions and no fast-simulation mixture, that fails to recover at least a several-fold reduction in background acceptance at fixed signal efficiency.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper evaluates the Particle Transformer (ParT) for jet flavor tagging under Higgs-factory conditions using ILD full simulation (~1M jets) and SGV fast simulation (1M and 10M jets). It trains 3-category (b/c/d), 6-category (b/c/d/u/s/g), and 11-category (including quark–antiquark) classifiers that ingest per-particle kinematics, impact parameters, track uncertainties, and CPID outputs from dE/dx and time-of-flight. The central claim is a factor-of-5–10 reduction in background acceptance relative to previous BDT-based LCFIPlus taggers for b/c tagging (e.g., c-background ~O(10^{-3}) at 80% b efficiency), with additional results on strange tagging and heavy-flavor q/q̄ separation, plus a clear gain from increasing SGV training statistics from 1M to 10M jets.
Significance. If the improvement holds under controlled comparison, this is a substantial advance for Higgs-factory jet tagging: end-to-end attention models with PID inputs can replace or strongly outperform traditional secondary-vertex + BDT pipelines, with direct impact on H o bb/cc/ss and related coupling measurements. Strengths include multi-category training (including q/q̄), explicit use of CPID, full-sim plus large-statistics fast-sim scaling, tabulated working points (Table 2), and an ONNX deployment path. The work is empirical and reproducible in spirit; the main open issue is whether the headline 5–10 imes factor is apples-to-apples against LCFIPlus on the same samples.
major comments (2)
- [Abstract; §4; Table 2(a); Fig. 1] Abstract and §4 claim a factor-of-5–10 improvement over previous BDT-based (LCFIPlus) taggers, but Fig. 1 and Table 2(a) report only ParT background efficiencies. No LCFIPlus ROC or working points are shown on the same mc-2020 full-sim or SGV samples, with the same Durham jets, CPID inputs, and efficiency definition. Because §4 already notes non-negligible full-sim vs SGV absolute differences that are still under investigation, the quoted factor mixes classifier, simulation model, and reconstruction version. A same-sample LCFIPlus baseline (or a clear citation of the exact historical numbers and conditions) is needed to make the central claim controlled.
- [§4; Table 2; Fig. 1–2] Table 2 and Fig. 1–2 give point estimates of background efficiency with no statistical uncertainties, bootstrap errors, or systematic variations (e.g., tracking/PID resolution, jet energy scale, train/test fluctuations). At the O(10^{-3})–O(10^{-4}) level the test-set size is limited, so the significance of the full-sim vs SGV gap and of the 1M→10M improvement cannot be judged. Error bars or a short uncertainty discussion are load-bearing for the performance claims.
minor comments (5)
- [§2] §2: “full simualtion” → “full simulation”.
- [Table 1; Ref. [11]] Table 1 cites “[11]” for variable definitions; the reference is a GitHub link. A short self-contained list or appendix of exact feature definitions would improve reproducibility.
- [Fig. 1] Fig. 1 caption says three panels side-by-side but the text layout shows only (a)–(c); ensure the published figure includes all panels and a clear LCFIPlus reference curve if retained.
- [§3; Fig. 2(b)] §3: the 11-category parton-matching procedure (color-singlet grouping then angular assignment) is important for q/q̄ labels; a brief validation of purity/misassignment rate would strengthen Fig. 2(b).
- [Table 1; §3–4] Clarify whether CPID is used only in 6- and 11-category trainings (as Table 1 suggests) and how that affects the 3-category vs 11-category comparison in Fig. 1.
Circularity Check
No circularity: empirical supervised tagging study; metrics are held-out test efficiencies, not inputs renamed as predictions.
full rationale
The paper is a standard machine-learning performance evaluation of Particle Transformer for jet flavor tagging on ILD full-sim and SGV samples. Labels are assigned from generated sample type (3-/6-category) or MC-truth parton matching (11-category); models are trained on 80% splits and evaluated on held-out 10% test sets via background efficiency at fixed signal efficiency (Fig. 1, Table 2, Fig. 2). No equation or claim reduces a reported quantity to a fitted parameter or definitional identity by construction. Self-citations (LCFIPlus [1], CPID [7], ONNX pipeline) supply the conventional baseline and tooling; they are not load-bearing uniqueness theorems that force the performance numbers. The narrative 5–10× improvement versus historical BDT taggers is an external comparison, not a circular derivation. Score 0 is therefore appropriate.
Axiom & Free-Parameter Ledger
free parameters (4)
- ECAL single-hit ToF resolution =
100 ps
- Train/validation/test split fractions =
80%/10%/10%
- ParT architecture and training hyperparameters
- Working-point signal efficiencies =
0.8 (primary tables)
axioms (6)
- domain assumption Geant4 ILD full simulation plus PandoraPFA/Durham reconstruction faithfully represent detector response for flavor-tagging observables.
- domain assumption SGV fast simulation is adequate for heavy-flavor (b/c) discrimination studies even without hadron PID.
- domain assumption Event-level sample labels (3- and 6-category) and MC-truth parton matching (11-category) define the correct jet flavor for supervised training.
- domain assumption Comprehensive PID BDT outputs (π/K/p probabilities from dE/dx and ToF) are valid per-particle features for the tagger.
- domain assumption Transformer self-attention with pairwise kinematic interaction bias is an appropriate model class for jet constituents.
- standard math Standard supervised classification metrics (background efficiency at fixed signal efficiency) measure tagging performance relevant to physics analyses.
read the original abstract
We study the performance of the Particle Transformer (ParT) for jet flavor tagging using ILD full simulation events (1M jets) as well as fast simulation samples (10M and 1M jets). We perform 3-category ($b/c/d$), 6-category ($b/c/d/u/s/g$), and 11-category trainings (including quark--antiquark separation), incorporating multivariate hadron particle identification information from $dE/dx$ and time-of-flight. For $b$/$c$ tagging, we observe a factor of 5--10 improvement over previous BDT-based taggers, and we obtain reasonable performance for strange tagging and quark/antiquark separation. The 10M-jet fast simulation study indicates that further gains are possible with higher training statistics.
Reference graph
Works this paper leans on
-
[1]
Suehara, T., Tanabe, T.: LCFIPlus: A Framework for Jet Analysis in Linear Collider Studies. Nucl. Instrum. Meth. A808, 109–116 (2016) https://doi.org/ 10.1016/j.nima.2015.11.054 [physics.ins-det]
-
[2]
Gouskos Loukas: Jet flavour tagging for future colliders with fast simulation
Bedeschi Franco, S.M. Gouskos Loukas: Jet flavour tagging for future colliders with fast simulation. The European Physical Journal C (2022) https://doi.org/ 10.1140/epjc/s10052-022-10609-1
-
[3]
Qu, H., Li, C., Qian, S.: Particle Transformer for Jet Tagging (2022) arXiv:arXiv:2202.03772 [hep-ph]
Pith/arXiv arXiv 2022
-
[4]
Abramowicz, H., et al.: International Large Detector: Interim Design Report (2020) arXiv:2003.01116 [physics.ins-det]
Pith/arXiv arXiv 2020
-
[5]
https://github.com/ilcsoft
iLCSoft. https://github.com/ilcsoft
-
[6]
Thomson, M.A.: Particle Flow Calorimetry and the PandoraPFA Algorithm. Nucl. Instrum. Meth. A611, 25–40 (2009) https://doi.org/10.1016/j.nima.2009. 09.009 [physics.ins-det]
-
[7]
Einhaus, U.: CPID: A Comprehensive Particle Identification Framework for Future e+e− Colliders (2023) arXiv:2307.15635 [hep-ex]
Pith/arXiv arXiv 2023
-
[8]
Kilian, W., Ohl, T., Reuter, J.: WHIZARD: Simulating Multi-Particle Processes at LHC and ILC. Eur. Phys. J. C71, 1742 (2011) https://doi.org/10.1140/epjc/ s10052-011-1742-y arXiv:0708.4233 [hep-ph]
-
[9]
JHEP05, 026 (2006) https://doi.org/10.1088/1126-6708/2006/05/026 arXiv:hep-ph/0603175
Sjostrand, T., Mrenna, S., Skands, P.Z.: PYTHIA 6.4 Physics and Man- ual. JHEP05, 026 (2006) https://doi.org/10.1088/1126-6708/2006/05/026 arXiv:hep-ph/0603175
-
[10]
Berggren, M.: SGV 3.0 - a fast detector simulation (2012) arXiv:1203.0217 [physics.ins-det]
Pith/arXiv arXiv 2012
-
[11]
https://github.com/suehara/ LCFIPlus/
LCFIPlus implementation adapted to ParT. https://github.com/suehara/ LCFIPlus/. See MLInputGenerator for exact definition of input variables. 7
-
[12]
https://onnxruntime.ai/
ONNX Runtime. https://onnxruntime.ai/
-
[13]
Seino, T., Sugawara, R., Suehara, T., Hosokawa, R., Ootani, W., Narita, S.: Eval- uation of measurement accuracy of higgs coupling constant to strange quarks in the international linear collider. Proc. LCWS2025, in preparation. (2025)
2025
-
[14]
Berggren, C.M., Bliewert, B., List, J., Ntounis, D., Suehara, T., Tian, J., Torndal, J.M., Vernieri, C.: A modern reconstruction and analysis framework for the ild zhh study. Proc. LCWS2025, in preparation. (2025) 8
2025
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.