REVIEW 3 major objections 4 minor 9 references
The paper estimates that current jet taggers are nearly optimal for boosted W, Z, and H→gg jets, while a large gap remains for top jets.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-03 05:35 UTC pith:W2PL4GQS
load-bearing objection A concise, honest proceedings note extending the generative-model jet-tagging limit to W/Z/H, but the central process-dependence is only as solid as the unvalidated learned densities. the 3 major comments →
The fundamental limit of jet tagging: Beyond top jets
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Using autoregressive transformers trained on a large simulated jet sample, the authors construct explicit density estimates for top, W, Z, and H→gg jets and for QCD background. From these densities they form the log-likelihood ratio log p_S(x)/p_QCD(x), which by the fundamental lemma of hypothesis testing is the optimal discriminator, and they compare it with a baseline transformer classifier trained on the same synthetic data. Their central observation is that the gap between the baseline tagger and this estimated optimum is strongly jet-type dependent: for W→qq′, Z→qq̄, and H→gg jets the gap is substantially reduced and in some cases nearly closed, whereas for top jets the previously obser
What carries the argument
The central object is the estimated log-likelihood ratio log[p_S(x)/p_QCD(x)], where p_S and p_QCD are the probability densities learned by autoregressive transformers, one per jet class, after discretizing each jet's constituents into tokens. Because the synthetic events have exact model likelihoods, this ratio yields the optimal ROC curve for the model-defined problem. The work compares that curve with a baseline transformer classifier sharing the same backbone but with a classification head, on identical synthetic samples, so the only difference between the estimator and the classifier is the loss used to train them.
Load-bearing premise
The estimated optimal limit is only as reliable as the learned generative densities: the likelihood-ratio classifier is provably optimal for p_S and p_QCD as modeled, but not guaranteed to be optimal for the true physical jet distributions, so the reported gap is a property of the models, not necessarily of nature.
What would settle it
Train a substantially more expressive or better-validated generative model for W, Z, or H→gg jets and recompute the likelihood-ratio ROC curve; if the estimated optimal curve shifts to significantly higher background rejection at fixed signal efficiency, the near-closed gap reported here is an artifact of generative-model insufficiencies. Alternatively, a two-sample statistical test rejecting the learned densities as inconsistent with an independent simulated sample would invalidate the estimated limit.
If this is right
- For boosted W, Z, and H→gg tagging, current transformer classifiers are already close to the estimated optimal limit, so large algorithmic improvements are not expected for these channels; remaining gains likely require richer input information.
- Top-tagging retains a sizeable gap to the estimated optimum, indicating that more headroom exists for future architecture or training improvements in that channel.
- The gap estimate is process-dependent, so benchmarking taggers against process-specific estimated limits is more informative than a single global limit.
- The framework allows the separation of information-limited tasks (W/Z/H) from model-limited tasks (top) within a common setup.
- Using a common generative-model and dataset setup, the ranking of tagger headroom across jet types is a stable qualitative result even if the absolute limit may shift with better generators.
Where Pith is reading between the lines
- The near-closed gap for W/Z/H could be partly an artifact of the generative models under-representing true jet stochasticity for these classes; a stronger test would validate the learned densities against real data with two-sample tests, and if the LLR curve moves upward, the reported near-optimality is model-limited.
- A direct extrapolation: if a generative model were trained with more expressive architectures or more data and the W/Z/H LLR curve remained stable, then the current taggers are effectively at the information limit for the constituent-level representation used here, implying only input-level changes (e.g., particle identification, vertexing) could push further.
- The same per-class generative-optimum construction could be applied to other signal-background pairs (e.g., tau vs QCD, or quark vs gluon discrimination), giving a general map of where modern taggers sit relative to their theoretical limits.
- The interpretation implies a testable asymmetry: for W/Z/H, doubling the generative model capacity should not materially change the estimated optimal ROC, whereas for top jets it should, if the top gap is due to model capacity rather than intrinsic information.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This proceedings note extends the framework of Ref. [1] for estimating the optimal jet-tagging limit from generative models to boosted W→qq′, Z→q̄q, and H→gg jets, in addition to the original top-quark case. Autoregressive transformers are trained on JetClass constituent-level data; after discretization, each class model supplies an explicit density p_S, and the log-likelihood ratio log[p_S(x)/p_QCD(x)] in Eq. (1) is taken as the estimated optimal classifier. A baseline transformer classifier (BTC) is trained on the same synthetic samples, and the ROC curves are compared in Fig. 1. The paper reports that the gap between BTC and the estimated optimal LLR is strongly process-dependent: large for top jets, but substantially reduced and in some cases nearly closed for W, Z, and H→gg jets. The authors are careful to call this an 'estimated' limit, cite Ref. [4] on generative-model dependence, and describe ongoing work on validation, scaling, and interpretation.
Significance. If the central empirical claim survives scrutiny, it is a valuable benchmark for the jet-tagging community. It would show that, for several standard tagging tasks, modern transformer-based classifiers are already close to the information-theoretic limit defined by the generative model, while top tagging remains an outlier. The paper's honesty about the model-dependent character of the limit, its explicit citation of Ref. [4], and its extension to four benchmark processes are strengths. The potential impact, however, is limited by the lack of any validation of the learned densities and by the absence of uncertainties on the reported ROC comparison; the headline process dependence could be a property of the generative models rather than of jet physics.
major comments (3)
- [§2–§3, Eq. (1), Fig. 1] The central claim—that the gap to the optimal limit is small for W/Z/H and large for top—rests entirely on the fidelity of the transformer densities p_S and p_QCD used in Eq. (1) and on the comparability of their accuracy across processes. No quantitative validation of these densities is provided. Training curves in Fig. 2 show only loss values, not whether the learned densities match held-out JetClass events. Since Ref. [4] demonstrated that the inferred limit depends on the choice and quality of the generative model, and since top jets have a more complex three-prong plus b-quark structure, one cannot exclude that the larger top gap is an artifact of a less accurate generative model for top jets. The paper should include a two-sample test or a comparison of classifier performance on real vs. synthetic events, at least for one signal class and QCD, before drawing the cross-process concl
- [Fig. 1, §3] The ROC comparison has no statistical or systematic uncertainties and no numerical performance table. Without error bars, multiple training seeds, or a measure of training variability, the statements that the gap is 'substantially smaller' and 'nearly closed' for W/Z/H are not quantitatively supported. With 10 million synthetic events, statistical fluctuations on the ROC points may be small, but the BTC is trained on a finite sample and the generative models are stochastic; a table with, e.g., background rejection at a few signal efficiencies, including standard deviations over seeds, is necessary to assess whether the observed jet dependence is significant.
- [§3, last paragraph] The interpretation that W/Z/H jets are 'closer to QCD' morphologically and in stochasticity, thereby making the optimum easier to approach, is presented as a plausible explanation but is not tested. If the central claim is established after addressing the density-fidelity concern, this interpretation would benefit from a quantitative measure of closeness, such as a divergence between learned signal and QCD densities or a direct comparison of the LLR distributions. As written, it is a post hoc narrative, not a result.
minor comments (4)
- [Title] The title says 'fundamental limit' while the text carefully says 'estimated optimal limit.' Given the acknowledged model dependence, the title overstates the result; consider 'model-estimated limit' or similar.
- [Fig. 2] Representative training curves are shown only for top and H→gg. The reader cannot tell whether W and Z models converged similarly. Adding all four or at least mentioning that they are representative would be useful.
- [§2] The discretization into (40,30,30) bins and the approximate 39k-token vocabulary are stated, but no details are given on the transformer architecture, training hyperparameters, or the generation procedure. For a proceedings note this may be acceptable, but citing the full method or an appendix would improve reproducibility.
- [References] Refs. [5–7] are cited as ongoing validation work, but the manuscript does not state whether any of those methods have been applied to the present models. A sentence clarifying the status would help avoid the impression that validation is deferred indefinitely.
Circularity Check
No circularity found: the comparison is explicitly model-relative and internally consistent.
full rationale
The paper's central object is the likelihood ratio in Eq. (1), built from the generative models' own densities p_S and p_QCD. The 'optimal' classifier is optimal for those learned densities, not for nature; the paper repeatedly says 'estimated optimal limit' and cites Ref. [4] for the dependence on model choice and quality. The comparison between the LLR and the baseline transformer classifier is a measured gap on the same synthetic data, not a fitted parameter renamed as a prediction. The process-dependent gap is an empirical result within that controlled framework, and the paper explicitly disclaims assigning a definitive physical optimum. The generative-model fidelity concern raised in the reader's take is a substantive validation question, but it does not make the derivation circular: no step reduces an output to an input by construction. Refs. [1] and [2] are methodological self-citations, but they supply an independently reproducible technique (training transformers and computing log-likelihoods) rather than a uniqueness theorem or a forced conclusion. No circular step can be quoted or exhibited.
Axiom & Free-Parameter Ledger
free parameters (2)
- Per-class generative transformer parameters =
Not reported
- Discretization grid and token vocabulary =
(40,30,30) bins; ~39k tokens
axioms (4)
- standard math Neyman-Pearson lemma: the likelihood ratio is the optimal binary classifier when true densities are known.
- domain assumption The trained generative transformers provide sufficiently accurate estimates of the true JetClass jet densities.
- domain assumption Synthetic events generated by the models are statistically representative enough to train a baseline classifier that reflects real-data performance.
- domain assumption JetClass simulated jets faithfully represent physical LHC jets for the purposes of this study.
read the original abstract
Jet tagging, i.e. determining the origin of high-energy hadronic jets, is a key challenge in particle physics. Machine-learning-based taggers have achieved remarkable progress, raising the question of how close current methods are to the theoretical limit of performance. Previous work addressed this question for boosted top-quark jets using transformer-based generative models that provide realistic synthetic jet data with known probability density functions. This enables a direct comparison between modern taggers and the optimal likelihood-ratio classifier. In this note, we summarize the approach and extend the study to boosted W, Z, and H$\rightarrow gg$ jets. We find that the gap to the estimated optimal limit is strongly jet dependent and is substantially reduced for these seemingly more challenging tagging tasks. Ongoing work aimed at understanding the interpretation, robustness, and scaling of these limits is also briefly discussed.
Figures
Reference graph
Works this paper leans on
-
[1]
J. Geuskens, N. Gite, M. Krämer, V. Mikuni, A. Mück, B. Nachman et al.,Fundamental limit of jet tagging,Phys. Rev. D112(2025) L091901 [2411.02628]
Pith/arXiv arXiv 2025
-
[2]
T. Finke, M. Krämer, A. Mück and J. Tönshoff,Learning the language of QCD jets with transformers,JHEP06(2023) 184 [2303.07364]
Pith/arXiv arXiv 2023
-
[3]
H. Qu, C. Li and S. Qian,Particle Transformer for Jet Tagging,2202.03772
-
[4]
I.Pang, D.A.Faroughy, D.Shih, R.DasandG.Kasieczka,SURFingtotheFundamentalLimit of Jet Tagging,2511.15779
-
[5]
S. Grossi, M. Letizia and R. Torre,Refereeing the referees: evaluating two-sample tests for validating generators in precision sciences,Mach. Learn. Sci. Tech.6(2025) 015052 [2409.16336]
Pith/arXiv arXiv 2025
-
[6]
S. Grossi, M. Letizia and R. Torre,Comparing generative models with the new physics learning machine,Nucl. Phys. B1024(2026) 117349 [2508.02275]
Pith/arXiv arXiv 2026
-
[7]
P. Cappelli, G. Grosso, M. Letizia, H. Reyes-González and M. Zanetti,Learning to validate generative models: a goodness-of-fit approach,Mach. Learn. Sci. Tech.7(2026) 045011 [2511.09118]
arXiv 2026
-
[8]
O. Amram, D.A. Faroughy, T. Gerdes, A. Hallin, G. Kasieczka, M. Krämer et al.,Neural Scaling Laws for Jet Generation,2605.28940
-
[9]
M. Vigl, N. Hartman, M. Kagan and L. Heinrich,Neural Scaling Laws for Boosted Jet Tagging,2602.15781. 4
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.