Pith. sign in

REVIEW 2 major objections 7 minor 36 references

A jet-transformer scaling law fit only on small models predicts the loss of models trained with over 100 imes more compute to within one percent, and that loss tracks tagging performance.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-30 23:51 UTC pith:VT66SVYM

load-bearing objection First real prospective scaling forecast in jet ML, cleanly executed, with a real but disclosed L∞ degeneracy and a soft fine-tune reporting choice. the 2 major comments →

arxiv 2607.23377 v1 pith:VT66SVYM submitted 2026-07-25 hep-ex cs.AI

Predict before you train: Scaling Laws for particle physics foundation models

classification hep-ex cs.AI
keywords scaling lawsfoundation modelsjet taggingparticle physicstransformerspretrainingcompute-optimal allocation
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper shows that for a generic transformer pretrained on collider jets, you can forecast large-model performance before spending the compute. The authors fit a joint model-and-data scaling law only on small ParticleViT runs below 10^19 FLOPs, then train five larger models afterward—up to more than one hundred times that compute—and the measured pretraining loss lands within about one percent of the prediction. They further show that lower pretraining loss systematically yields lower fine-tuning loss and higher background rejection on two standard tagging benchmarks. Within this model family and these tasks, a compute budget can therefore be turned into an expected physics result before any large model is trained. The planned frontier model matches published numbers for a physics-aware foundation model trained on the same corpus on accuracy, AUC, and quark/gluon rejection, with a residual edge for the physics-aware model only in the high-purity tail of top tagging. Five pretrained models and the full recipe are released.

Core claim

Fitting the Chinchilla-style law L(N,D)=A/N^α+B/D^β+L∞ only on ParticleViT runs below 10^19 FLOPs predicts the measured pretraining loss of five held-out models trained afterward at up to 3.4×10^20 FLOPs to within one percent, and lower pretraining loss systematically tracks lower fine-tuning loss and higher background rejection on top tagging and quark/gluon tagging—so compute can be translated into expected physics performance before any large model is trained.

What carries the argument

The joint model-and-data scaling law L(N,D)=A/N^α+B/D^β+L∞, fit only on the small-model IsoFLOP ladder and then used to forecast held-out large runs; pretraining cross-entropy is the common currency that links the forecast to downstream rejection.

Load-bearing premise

That pretraining loss on this jet taxonomy keeps predicting the physics metrics a practitioner actually cares about, including on tasks and purity cuts beyond the two benchmarks shown.

What would settle it

Train a larger ParticleViT (or the same family on a new held-out compute budget) and check whether measured pretraining loss still lands within roughly one percent of the law’s prediction, and whether the loss-to-rejection ordering on top and quark/gluon tagging breaks.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • A practitioner can fit the law on a modest small-model sweep and request compute against a target rejection before training the large model.
  • Within this family, compute is best spent overwhelmingly on data rather than model size (compute-optimal exponent a≈0.21).
  • A generic transformer at scale reaches performance consistent with a same-corpus physics-aware foundation model on bulk metrics, with residual physics-aware advantage only in the high-purity top-tagging tail.
  • The same fit-then-forecast protocol can be tested in other data-rich scientific domains before budgeting large runs.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If particle tokens truly carry low information density, other sparse scientific modalities (event records, spectra, sensor streams) may likewise favor small compute-optimal models and strong data scaling.
  • The residual high-purity top-tagging gap is a natural target for hybrid designs that add physics bias only where rare radiation patterns dominate.
  • Repeating the prospective holdout test on a physics-aware architecture would show whether planning power is architecture-specific or general.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 7 minor

Summary. The manuscript fits a Chinchilla-form joint model-and-data scaling law, L(N,D)=A/N^α+B/D^β+L∞, to 136 IsoFLOP cells of a generic vision-style transformer (ParticleViT) pretrained on the ~1.06B-jet OmniLearned corpus, using only runs below 10^19 FLOPs. It then reports that five models trained afterward at up to 3.4×10^20 FLOPs (113× the largest fit budget) land within 1% of the predicted pretraining loss (max 0.66%, mean 0.26%; Table II). Across 141 checkpoints × 5 fine-tunes, lower pretraining loss is shown to track lower fine-tuning loss and higher background rejection on top tagging and quark/gluon tagging (Fig. 2), closing a compute→loss→physics-performance planning chain. The frontier model (397M) is benchmarked against the physics-aware OmniLearned model trained on the same corpus and found consistent on accuracy, AUC, and quark/gluon rejection, with a residual deficit only in the high-purity top-tagging tail. Five checkpoints, the full recipe, and code are released.

Significance. If the extrapolation result holds up under the requested robustness checks, this is the first fit-then-forecast (rather than retrospective) scaling-law validation in particle physics, and it directly addresses a practical problem: the field's largest models are its most expensive to train, and prior jet scaling studies could not be used to budget runs in advance. The study is well designed for falsifiability: the fit is committed on small models, the held-out models are trained afterward, and the downstream transfer is measured on two community benchmarks disjoint from the pretraining mix. The release of five pretrained checkpoints, preprocessing constants, fine-tuning protocol, and full code under MIT license is a concrete community asset and makes the central claims independently checkable. The compute-optimal allocation result (a=β/(α+β)≈0.21, versus ≈0.5 for language models) is a robust and useful qualitative finding even where the precise optimum is an extrapolation, and the authors are commendably careful in labeling the information-density interpretation as a hypothesis.

major comments (2)
  1. [§III.D, Table II (with Table I and §III.B)] The headline 'within one percent' forecast is evaluated at a single, arbitrarily selected point on a degenerate parameter manifold, and the sensitivity of the held-out errors to that choice is never shown. The manuscript itself states that L∞ is unresolved (rails to 10^-12, but refits under least-squares, soft-L1, or 80% subsamples place it anywhere from 0 to ≈0.2 with 'negligible change in fit quality'), that β slides from 0.059 to ≲0.09 along the L∞ degeneracy, and that α spans 0.14–0.29. Table II's predictions are then computed with the railed near-zero L∞ carried 'through to the extrapolation.' This is not a cosmetic gap: the data term dominates the frontier loss (normalized amplitude 0.93 at D_0=6×10^7 versus ≈0.023 for the model term at 397M), and even the α uncertainty alone moves the model term at N=397M by roughly ±0.008 in absolute loss — about 1% of the measured 0.758, i.e. co
  2. [§V, Tables III–IV; §VII.D.e] The ParticleViT benchmark numbers are summarized by the best test value reached along each fine-tuning trajectory (averaged over five runs), while the OmniLearned and reference-tagger numbers they are compared against follow their source publications' protocols, which do not use test-trajectory selection. Best-over-training test selection is optimistically biased, so the 'consistent with the published numbers' conclusion of §V — an abstract-level claim — is drawn under a protocol asymmetry that favors ParticleViT. The authors disclose this limitation clearly (§VI.e, §VII.D.e), which is to their credit, but disclosure does not quantify the effect. Since test metrics were already logged every 1000 steps, the incremental cost of also logging validation metrics and re-summarizing at least the five frontier models with validation-only checkpoint selection is modest. At minimum, please report
minor comments (7)
  1. [§VII.A.b / §III.D] The scaling quantity is the Gaussian-smoothed training loss. Sweep runs make <1 pass over the corpus, so their train loss is on unseen examples, but the five held-out runs make ~2.8 passes, so their smoothed train loss partly reflects repeated data and is not the same statistical object as in the fit set. The agreement in Table II is empirical and stands, but a one-time evaluation of the five frontier models on a held-out sample of pretraining data would confirm that the law predicts a generalization loss, not a memorization-discounted one, and would preempt an obvious reader question.
  2. [§VII.A.c, Eq. (2)] The compute accounting C=6N(fL+1)BS drops the sequence-length-dependent attention FLOPs. This is a standard convention and is applied consistently across the ladder, but at L=151 and widths up to 1536 the neglected term is not negligible for the smaller models; a sentence quantifying the fractional correction (and noting it cancels in the IsoFLOP comparisons) would help.
  3. [§VII.D.c] Each of the 136 fit cells keeps the lowest-loss run over a learning-rate grid and up to three seeds. This best-of selection biases the fit points slightly low relative to a fixed-recipe run and is inherited from the Chinchilla procedure, but it is worth one sentence acknowledging that the law describes best-tuned runs, since a practitioner using the law to plan a single untuned run may land above it.
  4. [§III.C] The claim that the compute-optimal size is 'at or below 10^5 parameters' sits below the smallest grid model (3.6×10^5) and at the search floor; the text already flags this, but the abstract-level phrase 'far smaller than the rule of thumb' would be better anchored to the robust quantity a=0.21 (0.17, 0.30) rather than the extrapolated optimum.
  5. [Fig. 2] The y-axes are zoomed to narrow ranges (e.g., quark/gluon rejection 38–43), which visually amplifies the trends. Reporting a Spearman/Pearson coefficient per panel, or a fit of downstream metric versus pretraining loss, would make the 'systematically' claim quantitative and would also give practitioners the second link of the planning chain in usable form.
  6. [Table II] Please add a column with D (training jets/tokens) for each held-out model; the predictions depend on both N and D, and D is currently only inferable from C via Eq. (2).
  7. [Notation] β is overloaded for the data exponent and the AdamW β1, β2 (§VII.D.b); harmless in context but worth disambiguating. Also, 'meaningully' (§III.C) is a typo.

Circularity Check

0 steps flagged

No significant circularity: held-out loss forecasts and downstream transfer are independent measurements, not forced by the fit inputs.

full rationale

The central claim is an empirical extrapolation test, not a definitional identity. The Chinchilla form (Eq. 1) is fit only on IsoFLOP runs below 10^19 FLOPs (136 cells); five later models at up to 3.4×10^20 FLOPs are trained afterward, and their measured smoothed train losses are compared to the formula evaluated at those (N, D). Agreement within ~1% (Table II) is a genuine out-of-sample check: nothing in the fit procedure constrains those held-out losses to match. Downstream results are separate fine-tunes on held-out top-tagging and quark/gluon benchmarks; the monotonic link from pretraining loss to fine-tuning loss and rejection (Fig. 2) is observational, not by construction. Adopting Hoffmann et al.'s parametric form is a standard external ansatz, not a self-citation that forces the exponents or the forecast. Comparison to OmniLearned uses published same-corpus numbers as a baseline and does not underwrite the scaling-law claim. The unresolved L∞–β degeneracy (Sec. III B) is a robustness/correctness concern about which point on a flat manifold is extrapolated, not circularity: the paper still reports an independent measured-vs-predicted comparison at the chosen point. No step reduces the claimed prediction to its inputs by definition.

Axiom & Free-Parameter Ledger

7 free parameters · 6 axioms · 1 invented entities

The central forecast rests on the empirical Chinchilla parametric form, standard transformer FLOP accounting, the OmniLearned corpus and its 210-class label space, and the assumption that smoothed pretraining CE is the right scalar to extrapolate and to link to fine-tuned rejection. Free parameters are the five fit coefficients (with known degeneracies). No new physical entities are postulated; ParticleViT is an architecture choice, not a claimed new particle or force.

free parameters (7)
  • A (parameter coefficient) = 1.95
    Amplitude in L=A/N^α+…; strongly degenerate with α and quoted only as a point estimate.
  • α (parameter exponent) = 0.225 (0.14, 0.29)
    Controls model-size scaling and the compute-optimal allocation; 10th–90th percentile (0.14,0.29) from 80% subsamples.
  • B (data coefficient) = 2.67
    Amplitude of the data term in the Chinchilla form, fit on the 136-cell ladder.
  • β (data exponent) = 0.0591 (0.057, 0.063)
    Data-scaling exponent; small and tightly constrained when L∞ is pinned, widens if L∞ is freed.
  • L∞ (irreducible loss) = ≲0.2 (unresolved)
    Rails to search floor; unresolved and reported only as upper limit ≲0.2; carried as near-zero into extrapolation.
  • Huber δ and L-BFGS-B initialization grid = δ=10^-3
    Fitting hyperparameters that select the retained optimum among multiple starts.
  • Mean token occupancy f and sequence length L = f≈0.31, L=150 → ~47.5 tokens/jet
    Enter the FLOP formula C=6N(fL+1)BS; measured f≈0.31, L=150 particle slots.
axioms (6)
  • ad hoc to paper Chinchilla parametric form L(N,D)=A/N^α+B/D^β+L∞ is an adequate description of pretraining loss over the probed range
    Adopted from Hoffmann et al. and fit in Sec. III B; not derived from jet physics.
  • domain assumption Training compute ≈ 6 FLOPs per parameter per token
    Standard LM accounting used in Sec. VII A; defines the IsoFLOP grid.
  • domain assumption Smoothed training cross-entropy on the 210-class OmniLearned label space is the correct scalar for scaling and for predicting downstream rejection
    Sec. VII A and IV; sweep runs see each example at most once, held-out runs ~2.8 epochs.
  • domain assumption Jets are unordered sets of particles; no positional encoding is required
    Architectural premise in Sec. II A / VII B.
  • domain assumption Width-to-depth aspect ratio near 100 is a suitable one-parameter family for scaling N
    Taken from Kaplan et al. and held fixed across the ladder (Sec. II A).
  • domain assumption Downstream benchmarks are sufficiently disjoint from pretraining for the transfer claim to be meaningful
    Stated in Sec. II C and VII D; authors also note semantic overlap of top/QCD and q/g concepts with the pretraining taxonomy.
invented entities (1)
  • ParticleViT model family independent evidence
    purpose: Generic low-inductive-bias transformer on jet constituents used as the vehicle for the scaling law and public checkpoints.
    Standard transformer blocks with split particle stem, class token, RMSNorm, QK-norm, SwiGLU; not a new physical entity. independent_evidence is true insofar as checkpoints and code are released and can be evaluated on public benchmarks.

pith-pipeline@v1.2.0-grok45-kimik3 · 21695 in / 3877 out tokens · 68648 ms · 2026-07-30T23:51:38.577941+00:00 · methodology

0 comments
read the original abstract

The largest machine learning models in particle physics are also the most expensive to train, yet the return on scaling a given architecture cannot be estimated before that compute is spent. Scaling laws have been fit for jets, but none has yet been shown to predict the performance of models it was not fit on. We show that, for a generic transformer pretrained on collider jets, it can be forecast. Fitting a joint model-and-data scaling law on small models alone, spanning three orders of magnitude of training compute, we predict the loss of models trained afterward with more than one hundred times more compute to within one percent. We then connect the forecast to downstream physics performance: across two standard tagging benchmarks, lower pretraining loss yields systematically lower fine-tuning loss and higher background rejection after fine-tuning. Within this model family and these tasks, a compute budget can therefore be translated into expected physics performance before any large model is trained. The final frontier model is consistent with the published numbers for current state-of-the-art physics-aware foundation models trained on the same corpus, on accuracy, AUC, and quark/gluon rejection, with a residual edge for the physics-aware model only in the high-purity tail of top tagging. We release five pretrained models spanning multiple sizes, together with the complete training recipe and code.

Figures

Figures reproduced from arXiv: 2607.23377 by Benjamin Nachman, Christopher Re, Jan-Lucas Uslu.

Figure 1
Figure 1. Figure 1: FIG. 1 [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: FIG. 2 [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

36 extracted references · 29 linked inside Pith

  1. [1]

    R. S. Sutton, The bitter lesson,http://www. incompleteideas.net/IncIdeas/BitterLesson.html (2019), incomplete Ideas (blog), 13 March 2019

  2. [2]

    Kaplan, S

    J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei, arXiv preprint arXiv:2001.08361 (2020), arXiv:2001.08361 [cs.LG]

  3. [3]

    Henighan, J

    T. Henighan, J. Kaplan, M. Katz, M. Chen, C. Hesse, J. Jackson, H. Jun, T. B. Brown, P. Dhariwal, S. Gray, C. Hallacy, B. Mann, A. Radford, A. Ramesh, N. Ry- der, D. M. Ziegler, J. Schulman, D. Amodei, and S. McCandlish, arXiv preprint arXiv:2010.14701 (2020), arXiv:2010.14701 [cs.LG]

  4. [4]

    Hoffmann, S

    J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. de Las Casas, L. A. Hendricks, J. Welbl, A. Clark, T. Hennigan, E. Noland, K. Millican, G. van den Driessche, B. Damoc, A. Guy, S. Osindero, K. Simonyan, E. Elsen, O. Vinyals, J. Rae, and L. Sifre, inAdvances in Neural Information Processing Systems (NeurIPS), Vol. 35 (2022) pp...

  5. [5]

    C. L. Cheng, S. Corrodi, T. J. Hobbs, A. S. Mete, and B. Nachman, arXiv preprint arXiv:2606.00224 (2026), arXiv:2606.00224 [hep-ex]

  6. [6]

    Bhimji, C

    W. Bhimji, C. Harris, V. Mikuni, and B. Nachman, Phys. Rev. D113, 032020 (2026), phys. Rev. D 113, 032020 (2026), arXiv:2510.24066 [hep-ph]

  7. [7]

    H. Qu, C. Li, and S. Qian, inProceedings of the 39th International Conference on Machine Learning (ICML), Proceedings of Machine Learning Research, Vol. 162 (PMLR, 2022) pp. 18281–18292, particle Trans- former (ParT); also introduces the JetClass dataset., arXiv:2202.03772 [hep-ph]

  8. [8]

    J. Birk, A. Hallin, and G. Kasieczka, Mach. Learn.: Sci. Technol.5, 035031 (2024), arXiv:2403.05618 [hep-ph]

  9. [9]

    Golling, L

    T. Golling, L. Heinrich, M. Kagan, S. Klein, M. Leigh, M. Osadchy, and J. A. Raine, Mach. Learn.: Sci. Technol. 5, 035074 (2024), arXiv:2401.13537 [hep-ph]

  10. [10]

    Batson and Y

    J. Batson and Y. Kahn, SciPost Phys. Core8, 034 (2025), arXiv:2312.02264 (Dec 2023); SciPost Phys. Core 8, 034 (2025), arXiv:2312.02264 [hep-ph]

  11. [11]

    M. Vigl, N. Hartman, M. Kagan, and L. Heinrich, arXiv preprint arXiv:2602.15781 (2026), arXiv:2602.15781 [hep-ex]

  12. [12]

    Amram, D

    O. Amram, D. A. Faroughy, T. Gerdes, A. Hallin, G. Kasieczka, M. Kr¨ amer, H. Reyes-Gonzalez, and 10 D. Shih, arXiv preprint arXiv:2605.28940 (2026), arXiv:2605.28940 [hep-ph]

  13. [13]

    Kasieczka, T

    G. Kasieczka, T. Plehn, A. Butter, K. Cranmer, D. Deb- nath, B. M. Dillon, M. Fairbairn, D. A. Faroughy, W. Fedorko, C. Gay, L. Gouskos, J. F. Kamenik, P. T. Komiske, S. Leiss, A. Lister, S. Macaluso, E. M. Metodiev, L. Moore, B. Nachman, K. Nord- str¨ om, J. Pearkes, H. Qu, Y. Rath, M. Rieger, D. Shih, J. M. Thompson, and S. Varma, SciPost Phys.7, 014 (2...

  14. [14]

    P. T. Komiske, E. M. Metodiev, and J. Thaler, JHEP 01, 121, introduces the Energy Flow Network (EFN) and Particle Flow Network (PFN), and the Pythia8 quark/gluon classification dataset., arXiv:1810.05165 [hep-ph]

  15. [15]

    Ettinger, A

    Team Olmo, A. Ettinger, A. Bertsch, B. Kuehl, D. Gra- ham, D. Heineman, D. Groeneveld, F. Brahman, F. Tim- bers, H. Ivison, J. Morrison, J. Poznanski, K. Lo, L. Sol- daini, M. Jordan, M. Chen, M. Noukhovitch, N. Lam- bert, P. Walsh, P. Dasigi, R. Berry, S. Malik, S. Shah, S. Geng, S. Arora, S. Gupta, T. Anderson, T. Xiao, T. Murray, T. Romero, V. Graf, A....

  16. [16]

    Zhang and R

    B. Zhang and R. Sennrich, inAdvances in Neural Infor- mation Processing Systems (NeurIPS), Vol. 32 (2019) pp. 12360–12371, arXiv:1910.07467 [cs.LG]

  17. [17]

    Dehghani, J

    M. Dehghani, J. Djolonga, B. Mustafa, P. Padlewski, J. Heek, J. Gilmer, A. P. Steiner, M. Caron, R. Geirhos, I. Alabdulmohsin, R. Jenatton, L. Beyer, M. Tschannen, A. Arnab, X. Wang, C. Riquelme Ruiz, M. Minderer, J. Puigcerver, U. Evci, M. Kumar, S. Van Steenkiste, G. F. Elsayed, A. Mahendran, F. Yu, A. Oliver, F. Huot, J. Bastings, M. Collier, A. A. Gri...

  18. [18]

    Shazeer, arXiv preprint arXiv:2002.05202 (2020), arXiv:2002.05202 [cs.LG]

    N. Shazeer, arXiv preprint arXiv:2002.05202 (2020), arXiv:2002.05202 [cs.LG]

  19. [19]

    S. Gong, Q. Meng, J. Zhang, H. Qu, C. Li, S. Qian, W. Du, Z.-M. Ma, and T.-Y. Liu, JHEP07, 030, lorentzNet., arXiv:2201.08187 [hep-ph]

  20. [20]

    Bogatskiy, T

    A. Bogatskiy, T. Hoffman, D. W. Miller, and J. T. Of- fermann, Pelican: Permutation equivariant and lorentz invariant or covariant aggregator network for parti- cle physics, Machine Learning and the Physical Sci- ences Workshop at NeurIPS (2022), pELICAN. Origi- nal workshop paper; the full study was published as Bo- gatskiy et al., JHEP 03 (2024) 113 (ar...

  21. [21]

    Spinner, V

    J. Spinner, V. Bres´ o, P. de Haan, T. Plehn, J. Thaler, and J. Brehmer, inAdvances in Neural Information Process- ing Systems (NeurIPS), Vol. 37 (2024) pp. 22178–22205, l-GATr; source of the base L-GATr top-tagging results., arXiv:2405.14806 [physics.data-an]

  22. [22]

    Qu and L

    H. Qu and L. Gouskos, Phys. Rev. D101, 056019 (2020), particleNet; also reports the ResNeXt-50 and P-CNN baselines., arXiv:1902.08570 [hep-ph]

  23. [23]

    E. A. Moreno, O. Cerri, J. M. Duarte, H. B. New- man, T. Q. Nguyen, A. Periwal, M. Pierini, A. Serikova, M. Spiropulu, and J.-R. Vlimant, Eur. Phys. J. C80, 58 (2020), jEDI-net., arXiv:1908.05318 [hep-ex]

  24. [24]

    Mikuni and F

    V. Mikuni and F. Canelli, Mach. Learn.: Sci. Tech- nol.2, 035027 (2021), point Cloud Transformer (PCT)., arXiv:2102.05073 [physics.data-an]

  25. [25]

    Bogatskiy, B

    A. Bogatskiy, B. Anderson, J. T. Offermann, M. Roussi, D. W. Miller, and R. Kondor, inProceedings of the 37th International Conference on Machine Learning, Proceed- ings of Machine Learning Research, Vol. 119 (PMLR,

  26. [26]

    Shimmin, Particle convolution for high energy physics, ML4Jets 2021 (2021), reduced Particle Cloud Network (rPCN)., arXiv:2107.02908 [hep-ph]

    C. Shimmin, Particle convolution for high energy physics, ML4Jets 2021 (2021), reduced Particle Cloud Network (rPCN)., arXiv:2107.02908 [hep-ph]

  27. [27]

    Hammad and M

    A. Hammad and M. M. Nojiri, JHEP06, 176, mixer., arXiv:2404.14677 [hep-ph]

  28. [28]

    Y. Wu, K. Wang, C. Li, H. Qu, and J. Zhu, Chin. Phys. C 49, 013110 (2025), more-Interaction Particle Transformer (MIParT)., arXiv:2407.08682 [hep-ph]

  29. [29]

    Brehmer, V

    J. Brehmer, V. Bres´ o, P. de Haan, T. Plehn, H. Qu, J. Spinner, and J. Thaler, SciPost Phys.19, 108 (2025), l-GATr; source of the fine-tuned L-GATr top-tagging re- sults., arXiv:2411.00446 [hep-ph]

  30. [30]

    Mikuni and B

    V. Mikuni and B. Nachman, Phys. Rev. D111, 054015 (2025), point Edge Transformer (PET)., arXiv:2502.14652 [hep-ph]

  31. [31]

    Mikuni and B

    V. Mikuni and B. Nachman, Phys. Rev. D111, L051504 (2025), omniLearn foundation model., arXiv:2404.16091 [hep-ph]

  32. [32]

    J.-L. Uslu, K. Greif, D. Whiteson, and B. Nach- man, arXiv preprint arXiv:2606.19781 (2026), arXiv:2606.19781 [hep-ex]

  33. [33]

    Geuskens, N

    J. Geuskens, N. Gite, M. Kr¨ amer, V. Mikuni, A. M¨ uck, B. Nachman, and H. Reyes-Gonz´ alez, Phys. Rev. D112, L091901 (2025), arXiv:2411.02628 [hep-ph]

  34. [34]

    Muennighoff, A

    N. Muennighoff, A. Rush, B. Barak, T. Le Scao, N. Tazi, A. Piktus, S. Pyysalo, T. Wolf, and C. A. Raffel, inAdvances in Neural Information Processing Systems (NeurIPS), Vol. 36 (2023) pp. 50358–50376, arXiv:2305.16264 [cs.CL]

  35. [35]

    Chowdhery, S

    A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sut- ton, S. Gehrmann, P. Schuh, K. Shi, S. Tsvyashchenko, J. Maynez, A. Rao, P. Barnes, Y. Tay, N. Shazeer, V. Prabhakaran, E. Reif, N. Du, B. Hutchinson, R. Pope, J. Bradbury, J. Austin, M. Isard, G. Gur-Ari, P. Yin, T. Duke, A. Levskaya, S. Ghemawat, S. Dev,...

  36. [2020]

    992–1002, lorentz Group Network (LGN)., arXiv:2006.04780 [hep-ph]

    pp. 992–1002, lorentz Group Network (LGN)., arXiv:2006.04780 [hep-ph]