Pith. sign in

REVIEW 2 major objections 1 minor 1 cited by

Pre-trained LLMs supply embeddings that let one autoregressive transformer generate events for any LHC process.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.3

2026-06-30 10:10 UTC pith:MH7DT7B5

load-bearing objection The paper proposes LLM embeddings to condition one transformer on many LHC processes, but the abstract gives no metrics or ablations to show the embeddings actually help. the 2 major comments →

arxiv 2606.23791 v2 pith:MH7DT7B5 submitted 2026-06-22 hep-ph

One Generator, Any Process: LLM-Conditioning for the LHC

classification hep-ph
keywords LHC event generationgenerative networksLLM conditioningautoregressive transformerFeynman diagramsphysics inductive biasprocess generalization
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper proposes novel conditioning schemes that feed continuous parameters, process labels, and Feynman diagrams into an autoregressive transformer via embeddings drawn from pre-trained LLMs. These embeddings are intended to inject high-level physics patterns as inductive bias. The resulting networks are claimed to converge faster, produce higher-quality events, and generalize to processes absent from the training set. A reader would care because standard LHC simulation pipelines train separate networks for each process, which multiplies computational cost.

Core claim

Employing pre-trained LLMs as multi-modal foundation models to provide descriptive embeddings for continuous parameters, process labels, and Feynman diagrams equips an autoregressive transformer with high-level physics-inductive bias; with this conditioning the generative networks converge faster, deliver better results, and generalize to unseen processes.

What carries the argument

LLM-derived embeddings that encode continuous parameters, process labels, and Feynman diagrams and serve as conditioning input to the autoregressive transformer.

Load-bearing premise

Pre-trained LLMs already produce embeddings that meaningfully capture relevant high-level physics patterns for LHC processes without further domain-specific training.

What would settle it

Train the conditioned transformer on a set of processes and evaluate generation quality and training speed on a new process; if performance matches or falls below an unconditioned baseline, the benefit of the LLM embeddings is absent.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • A single trained network can replace multiple process-specific generators.
  • Training time decreases because common patterns across processes are reused via the embeddings.
  • Event quality improves because the model receives explicit high-level physics structure rather than learning it from scratch.
  • The same architecture applies to new processes by simply supplying the appropriate LLM embeddings at inference time.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The approach could be extended by feeding the same LLM embeddings into other architectures such as diffusion models or flow-based generators.
  • If the embeddings capture diagram topology, the method might reduce the need for explicit graph neural network layers in future generators.
  • Process generalization might allow rapid prototyping of simulations for beyond-Standard-Model scenarios whose diagrams were never shown during training.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 1 minor

Summary. The manuscript proposes novel conditioning schemes for continuous parameters, process labels, and Feynman diagrams in an autoregressive transformer for LHC event generation. Pre-trained LLMs are used as multi-modal foundation models to supply descriptive embeddings, with the central claim that this imparts high-level physics-inductive bias resulting in faster convergence, better results, and generalization to unseen processes.

Significance. If the quantitative claims hold after proper validation, the work could enable a single generator architecture across multiple LHC processes by transferring inductive bias from LLMs, reducing the computational overhead of training separate models per process. This would be a notable contribution to generative modeling in high-energy physics, provided the embeddings demonstrably encode relevant QFT structures beyond generic text patterns.

major comments (2)
  1. Abstract: the claims of faster convergence, better results, and generalization to unseen processes are stated without any quantitative metrics, baselines, error bars, ablation studies, or experimental details, rendering the central claim unevaluable from the provided text.
  2. Abstract: the load-bearing assumption that pre-trained LLM embeddings meaningfully encode high-level physics patterns for Feynman diagrams, process labels, and continuous parameters (without domain adaptation) is not supported by any control experiments or comparisons to non-physics baselines, which is required to substantiate the claimed inductive bias over standard conditioning.
minor comments (1)
  1. Abstract: 'provide better result' should be pluralized to 'results' for grammatical consistency.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for the constructive feedback. We address the two major comments on the abstract below and will make targeted revisions to strengthen the presentation of our results and claims.

read point-by-point responses
  1. Referee: Abstract: the claims of faster convergence, better results, and generalization to unseen processes are stated without any quantitative metrics, baselines, error bars, ablation studies, or experimental details, rendering the central claim unevaluable from the provided text.

    Authors: We agree that the abstract would be strengthened by including quantitative support for the claims. In the revised version we will expand the abstract to report key metrics (e.g., relative convergence speed-up, performance deltas versus baselines with error bars, and generalization accuracy on held-out processes) while preserving brevity, with explicit pointers to the relevant figures and tables. revision: yes

  2. Referee: Abstract: the load-bearing assumption that pre-trained LLM embeddings meaningfully encode high-level physics patterns for Feynman diagrams, process labels, and continuous parameters (without domain adaptation) is not supported by any control experiments or comparisons to non-physics baselines, which is required to substantiate the claimed inductive bias over standard conditioning.

    Authors: The paper demonstrates the benefit of LLM conditioning via direct empirical comparisons against standard (non-LLM) conditioning on the same architecture, showing faster convergence, higher fidelity, and generalization. These results provide indirect support for the inductive bias. We acknowledge that dedicated ablations against generic or randomly initialized text embeddings are absent; we will add a brief discussion of this point and, if space permits, a control comparison in the revised manuscript. revision: partial

Circularity Check

0 steps flagged

No circularity: architecture proposal with no derivations or self-referential reductions

full rationale

The manuscript proposes an LLM-conditioned autoregressive transformer for LHC event generation and asserts faster convergence plus generalization from high-level inductive bias. No equations, fitted parameters, or derivation steps appear in the provided text. The central claim is an empirical hypothesis about LLM embeddings supplying physics patterns; it does not reduce by construction to any input, self-citation chain, or renamed ansatz. The architecture is therefore self-contained against external benchmarks and receives the default non-circularity finding.

Axiom & Free-Parameter Ledger

0 free parameters · 1 axioms · 0 invented entities

Abstract-only review; no explicit free parameters, axioms, or invented entities are stated beyond reliance on pre-trained LLMs.

axioms (1)
  • domain assumption Pre-trained LLMs encode useful high-level physics patterns in embeddings for processes and diagrams
    Central to the proposed conditioning scheme.

pith-pipeline@v0.9.1-grok · 5593 in / 1087 out tokens · 33636 ms · 2026-06-30T10:10:03.719874+00:00 · methodology

0 comments
read the original abstract

Neural network training for LHC event generation should, ideally, benefit from common high-level patterns in different processes. We propose novel conditioning schemes for continuous parameters, process labels, and Feynman diagrams. We employ pre-trained LLMs as multi-modal foundation models to provide descriptive embeddings for an autoregressive transformer. With such high-level physics-inductive bias the generative networks converge faster, provide better result, and generalize to unseen processes.

Figures

Figures reproduced from arXiv: 2606.23791 by Daniel Schiller, Henning Bahl, Thanush Sivagnanalingam, Tilman Plehn.

Figure 1
Figure 1. Figure 1: Autoregressive generative architecture. The prefix tokens [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Left: Drell-Yan invariant mass distribution for [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Predicted Z mass, extracted from the generated invariant mass peak, as a function of the true Z mass. The shaded regions indicate the in-training masses. The uncertainty bars indicate the standard deviation over five independent runs. case of text input, we test the Qwen3 model with 2B parameters in addition. Each conditioning scheme is trained independently five times. Both LLMs are partially fine-tuned, … view at source ↗
Figure 4
Figure 4. Figure 4: Loss comparison over 20 epochs for different conditioning mechanisms. [PITH_FULL_IMAGE:figures/full_fig_p009_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: AUCs for different conditioning schemes. The dashed line indicates the [PITH_FULL_IMAGE:figures/full_fig_p010_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Kinematics for uu¯ → t¯tH (in-training) and g g → t¯tH (hold-out) produc￾tion. We show the rapidity of the t¯tH system (left) and mt¯t (right). The histograms depict the mean and standard deviation of 5 independent runs. In the lower panels we also show the uu¯ → t¯t distributions for comparison. of [PITH_FULL_IMAGE:figures/full_fig_p011_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: AUCs for different conditioning schemes trained for 120 epochs. [PITH_FULL_IMAGE:figures/full_fig_p012_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: AUCs for different conditioning schemes for a non-LLM transformer back [PITH_FULL_IMAGE:figures/full_fig_p013_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Classifier AUC for g g → t¯tH as a function of the number of finetuning events, for the one-hot and edge-list conditioning compared to training from scratch. The left panel shows the zero-shot AUC of the pretrained networks. Lines and bands are the mean and standard deviation over five independent runs. batch size and the number of epochs fixed and use a cosine-annealing learning-rate schedule, such that a… view at source ↗
Figure 10
Figure 10. Figure 10: Rapidity of the t¯tH system (left) and mt¯t (right) for g g → t¯tH, comparing the edge-list network finetuned on 2048 events to the network trained from scratch on 2048 and 10000 events. Histograms show the mean and standard deviation over five independent runs, with the ratio to the truth in the lower panels. 14 [PITH_FULL_IMAGE:figures/full_fig_p014_10.png] view at source ↗
Figure 11
Figure 11. Figure 11: Classifier AUC versus number of finetuning events for high-multiplicity [PITH_FULL_IMAGE:figures/full_fig_p015_11.png] view at source ↗
Figure 12
Figure 12. Figure 12: Classifier AUC versus number of finetuning events for [PITH_FULL_IMAGE:figures/full_fig_p015_12.png] view at source ↗
Figure 13
Figure 13. Figure 13: Intermediate mbW+ for in-training uu¯ → b¯tW+ (left) and hold-out g g → b¯tW+ (right), the latter finetuned on 2048 events. Green and orange lines denote short and long pretraining, black is the truth. Bands show the run-to-run standard deviation, or Poisson uncertainties for the single long-pretraining run. but the network has only learned to generate the multiplicities seen during pretraining. After fin… view at source ↗
Figure 14
Figure 14. Figure 14: Exemplary Feynman diagram image used as input. [PITH_FULL_IMAGE:figures/full_fig_p021_14.png] view at source ↗
Figure 15
Figure 15. Figure 15: Left: Drell-Yan invariant mass distribution. Right: same for the positron [PITH_FULL_IMAGE:figures/full_fig_p023_15.png] view at source ↗
Figure 16
Figure 16. Figure 16: Drell-Yan invariant mass (left) and positron transverse momentum (right) [PITH_FULL_IMAGE:figures/full_fig_p024_16.png] view at source ↗
Figure 17
Figure 17. Figure 17: Comparison of the AUCs of different conditioning schemes for the 2 [PITH_FULL_IMAGE:figures/full_fig_p025_17.png] view at source ↗
Figure 18
Figure 18. Figure 18: Comparison of the AUCs of different conditioning schemes for the 2 [PITH_FULL_IMAGE:figures/full_fig_p026_18.png] view at source ↗
Figure 19
Figure 19. Figure 19: AUCs for the different input conditioning schemes using the Qwen3 back [PITH_FULL_IMAGE:figures/full_fig_p027_19.png] view at source ↗
Figure 20
Figure 20. Figure 20: Overview of individual train process AUCs for the different input repre [PITH_FULL_IMAGE:figures/full_fig_p028_20.png] view at source ↗
Figure 21
Figure 21. Figure 21: Overview of individual hold-out process AUCs for the different input [PITH_FULL_IMAGE:figures/full_fig_p029_21.png] view at source ↗
Figure 22
Figure 22. Figure 22: Validation loss over 120 epochs for the in-interm-out and edge-list condi [PITH_FULL_IMAGE:figures/full_fig_p030_22.png] view at source ↗
Figure 23
Figure 23. Figure 23: Invariant mass mbW+ reconstructing the intermediate top resonance, for the in-training uu¯ → b¯tW+ (left) and the hold-out g g → b¯tW+ (right), the latter shown pretrained (zero-shot) and finetuned on 2048 and 10000 events. Line colour denotes the pretraining length (short/long) and the finetuning dataset size, as in the legend, black denotes the truth. Bands show the run-to-run standard deviation, or Poi… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Neural Control Variates at LO and NLO

    hep-ph 2026-07 accept novelty 7.0

    Signed neural control variates from normalizing flows, combined with neural importance sampling, reduce weight ranges and negative weights for LO and NLO phase-space integration and event generation.

Reference graph

Works this paper leans on

296 extracted references · 296 canonical work pages · cited by 1 Pith paper · 54 internal anchors

  1. [1]

    Learning Feynman Diagrams using Graph Neural Networks

    Mitchell, Harrison and Norcliffe, Alexander and Li \`o , Pietro. Learning Feynman Diagrams using Graph Neural Networks. 2022. arXiv:2211.15348

  2. [2]

    Parameterized Machine Learning for High-Energy Physics

    Baldi, Pierre and Cranmer, Kyle and Faucett, Taylor and Sadowski, Peter and Whiteson, Daniel. Parameterized neural networks for high-energy physics. Eur. Phys. J. C. 2016. doi:10.1140/epjc/s10052-016-4099-4. arXiv:1601.07913

  3. [3]

    Monte Carlo Event Generation with Continuous Normalizing Flows

    Bothmann, Enrico and Jan en, Timo and Knobbe, Max and Schmitzer, Bernhard and Sinz, Fabian. Monte Carlo Event Generation with Continuous Normalizing Flows. 2026. arXiv:2604.03511

  4. [4]

    Sampling NNLO QCD phase space with normalizing flows

    Jan en, Timo and Poncelet, Rene and Schumann, Steffen. Sampling NNLO QCD phase space with normalizing flows. JHEP. 2025. doi:10.1007/JHEP09(2025)194. arXiv:2505.13608

  5. [5]

    and Jan en, T

    Bothmann, E. and Jan en, T. and Knobbe, M. and Schmitzer, B. and Sinz, F. Efficient many-jet event generation with Flow Matching. 2025. arXiv:2506.18987

  6. [6]

    Point cloud transformers applied to collider physics

    Mikuni, Vinicius and Canelli, Florencia. Point cloud transformers applied to collider physics. Mach. Learn. Sci. Tech. 2021. doi:10.1088/2632-2153/ac07f6. arXiv:2102.05073

  7. [7]

    Reconstructing particles in jets using set transformer and hypergraph prediction networks

    Di Bello, Francesco Armando and others. Reconstructing particles in jets using set transformer and hypergraph prediction networks. Eur. Phys. J. C. 2023. doi:10.1140/epjc/s10052-023-11677-7. arXiv:2212.01328

  8. [8]

    Learning the language of QCD jets with transformers

    Finke, Thorben and Kr. Learning the language of QCD jets with transformers. JHEP. 2023. doi:10.1007/JHEP06(2023)184. arXiv:2303.07364

  9. [9]

    Journal of High Energy Physics , year =

    Komiske, Patrick T. and Metodiev, Eric M. and Thaler, Jesse. Energy Flow Networks: Deep Sets for Particle Jets. JHEP. 2019. doi:10.1007/JHEP01(2019)121. arXiv:1810.05165

  10. [10]

    Jet tagging via particle clouds

    Qu, Huilin and Gouskos, Loukas. ParticleNet: Jet Tagging via Particle Clouds. Phys. Rev. D. 2020. doi:10.1103/PhysRevD.101.056019. arXiv:1902.08570

  11. [11]

    Hydrogen reionization ends by z = 5.3: Lyman-alpha optical depth measured by the XQR-30 sample , volume=

    Bosman, Sarah E I and Davies, Frederick B and Becker, George D and Keating, Laura C and Davies, Rebecca L and Zhu, Yongda and Eilers, Anna-Christina and D’Odorico, Valentina and Bian, Fuyan and Bischetti, Manuela and Cristiani, Stefano V and Fan, Xiaohui and Farina, Emanuele P and Haehnelt, Martin G and Hennawi, Joseph F and Kulkarni, Girish and Mesinger,...

  12. [12]

    Spina, Benedetta and Bosman, Sarah E. I. and Davies, Frederick B. and Gaikwad, Prakash and Zhu, Yongda , year=. Damping wings in the Lyman- forest: A model-independent measurement of the neutral fraction at 5.4 < z < 6.1 , volume=. doi:10.1051/0004-6361/202450798 , journal=

  13. [13]

    and Spannowsky, Michael

    Bhardwaj, Akanksha and Englert, Christoph and Naskar, Wrishik and Ngairangbam, Vishal S. and Spannowsky, Michael. Equivariant, safe and sensitive graph networks for new physics. JHEP. 2024. doi:10.1007/JHEP07(2024)245. arXiv:2402.12449

  14. [14]

    2022, MNRAS, 511, 3, 3446

    Neutsch, Steffen and Heneka, Caroline and Br\"uggen, Marcus , title = ". Mon. Not. Roy. Astron. Soc. 2022. doi:10.1093/mnras/stac218. arXiv:2201.07587

  15. [15]

    Proceedings of the National Academy of Sciences , year =

    Cranmer, Kyle and Brehmer, Johann and Louppe, Gilles. The frontier of simulation-based inference. Proc. Nat. Acad. Sci. 2020. doi:10.1073/pnas.1912789117. arXiv:1911.01429

  16. [16]

    Machine Learning and Cosmology

    Dvorkin, Cora and others. Machine Learning and Cosmology. Snowmass 2021. 2022. arXiv:2203.08056

  17. [17]

    Enhancing Gravitational-Wave Science with Machine Learning

    Cuoco, Elena and others. Enhancing Gravitational-Wave Science with Machine Learning. Mach. Learn. Sci. Tech. 2021. doi:10.1088/2632-2153/abb93a. arXiv:2005.03745

  18. [18]

    RAS Techniques and Instruments , author =

    Slijepcevic, Inigo V and Scaife, Anna M M and Walmsley, Mike and Bowles, Micah and Wong, O Ivy and Shabala, Stanislav S and White, Sarah V , title =. RAS Techniques and Instruments , volume =. 2023 , month =. doi:10.1093/rasti/rzad055 , url =

  19. [19]

    AstroCLIP: a cross-modal foundation model for galaxies , volume=

    Parker, Liam and Lanusse, Francois and Golkar, Siavash and Sarra, Leopoldo and Cranmer, Miles and Bietti, Alberto and Eickenberg, Michael and Krawezik, Geraud and McCabe, Michael and Morel, Rudy and Ohana, Ruben and Pettee, Mariel and Régaldo-Saint Blancard, Bruno and Cho, Kyunghyun and Ho, Shirley and The Polymathic AI Collaboration , title =. Monthly No...

  20. [20]

    and Mishra-Sharma, Siddharth and Villar, V

    Zhang, Gemma and Helfer, Thomas and Gagliano, Alexander T. and Mishra-Sharma, Siddharth and Villar, V. Ashley. Maven: a multimodal foundation model for supernova science. Mach. Learn. Sci. Tech. 2024. doi:10.1088/2632-2153/ad990d. arXiv:2408.16829

  21. [21]

    and Roussi, Marwah and Miller, David W

    Bogatskiy, Alexander and Anderson, Brandon and Offermann, Jan T. and Roussi, Marwah and Miller, David W. and Kondor, Risi. Lorentz Group Equivariant Neural Network for Particle Physics. 2020. arXiv:2006.04780

  22. [22]

    uller, David I. and Schuh, Daniel , title =

    Favoni, Matteo and Ipp, Andreas and M\"uller, David I. and Schuh, Daniel , title = ". Phys. Rev. Lett. 2022. doi:10.1103/PhysRevLett.128.032003. arXiv:2012.12901

  23. [23]

    & Viviani, M

    Bulusu, Srinath and Favoni, Matteo and Ipp, Andreas and M\"uller, David I. and Schuh, Daniel , title = ". EPJ Web Conf. 2022. doi:10.1051/epjconf/202225809001. arXiv:2112.12493

  24. [24]

    An efficient Lorentz equivariant graph neural network for jet tagging

    Gong, Shiqi and Meng, Qi and Zhang, Jue and Qu, Huilin and Li, Congqiao and Qian, Sitian and Du, Weitao and Ma, Zhi-Ming and Liu, Tie-Yan. An efficient Lorentz equivariant graph neural network for jet tagging. JHEP. 2022. doi:10.1007/JHEP07(2022)030. arXiv:2201.08187

  25. [25]

    uller, David I. , title =

    Favoni, Matteo and Ipp, Andreas and M\"uller, David I. , title = ". EPJ Web Conf. 2022. doi:10.1051/epjconf/202227409001. arXiv:2212.00832

  26. [26]

    Symmetry Group Equivariant Architectures for Physics

    Bogatskiy, Alexander and others. Symmetry Group Equivariant Architectures for Physics. Snowmass 2021. 2022. arXiv:2203.06153

  27. [27]

    and Offermann, Jan T

    Bogatskiy, Alexander and Hoffman, Timothy and Miller, David W. and Offermann, Jan T. PELICAN: Permutation Equivariant and Lorentz Invariant or Covariant Aggregator Network for Particle Physics. 2022. arXiv:2211.00454

  28. [28]

    Equivariant Graph Neural Networks for Charged Particle Tracking

    Murnane, Daniel and Thais, Savannah and Thete, Ameya. Equivariant Graph Neural Networks for Charged Particle Tracking. 21th International Workshop on Advanced Computing and Analysis Techniques in Physics Research : AI meets Reality. 2023. arXiv:2304.05293

  29. [29]

    Learning broken symmetries with approximate invariance

    Nabat, Seth and Ghosh, Aishik and Witkowski, Edmund and Kasieczka, Gregor and Whiteson, Daniel. Learning broken symmetries with approximate invariance. Phys. Rev. D. 2025. doi:10.1103/PhysRevD.111.072002. arXiv:2412.18773

  30. [30]

    and Offermann, Jan T

    Bogatskiy, Alexander and Hoffman, Timothy and Miller, David W. and Offermann, Jan T. and Liu, Xiaoyang. Explainable equivariant neural networks for particle physics: PELICAN. JHEP. 2024. doi:10.1007/JHEP03(2024)113. arXiv:2307.16506

  31. [31]

    and Spannowsky, Michael

    Ma\^ tre, Daniel and Ngairangbam, Vishal S. and Spannowsky, Michael. Optimal equivariant architectures from the symmetries of matrix-element likelihoods. Mach. Learn. Sci. Tech. 2025. doi:10.1088/2632-2153/adbab1. arXiv:2410.18553

  32. [32]

    and Hallin, Anna and Kasieczka, Gregor and Kr

    Amram, Oz and Anzalone, Luca and Birk, Joschka and Faroughy, Darius A. and Hallin, Anna and Kasieczka, Gregor and Kr. Aspen Open Jets: unlocking LHC data for foundation models in particle physics. Mach. Learn. Sci. Tech. 2025. doi:10.1088/2632-2153/ade58f. arXiv:2412.10504

  33. [33]

    2023 , eprint=

    LIMA: Less Is More for Alignment , author=. 2023 , eprint=

  34. [34]

    RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback , author=

    RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback , author=. 2024 , eprint=

  35. [35]

    2021 , eprint=

    LoRA: Low-Rank Adaptation of Large Language Models , author=. 2021 , eprint=

  36. [36]

    2024 , eprint=

    RLHF Deciphered: A Critical Analysis of Reinforcement Learning from Human Feedback for LLMs , author=. 2024 , eprint=

  37. [37]

    Terry , journal =

    Ralph Allan Bradley and Milton E. Terry , journal =. Rank Analysis of Incomplete Block Designs: I. The Method of Paired Comparisons , urldate =

  38. [38]

    2024 , eprint=

    Direct Preference Optimization: Your Language Model is Secretly a Reward Model , author=. 2024 , eprint=

  39. [39]

    2024 , eprint=

    Investigating the Synergistic Effects of Dropout and Residual Connections on Language Model Training , author=. 2024 , eprint=

  40. [40]

    2020 , eprint=

    Scaling Laws for Neural Language Models , author=. 2020 , eprint=

  41. [41]

    2022 , eprint=

    Emergent Abilities of Large Language Models , author=. 2022 , eprint=

  42. [42]

    Permutationless many-jet event reconstruction with symmetry preserving attention networks

    Fenton, Michael James and Shmakov, Alexander and Ho, Ta-Wei and Hsu, Shih-Chieh and Whiteson, Daniel and Baldi, Pierre. Permutationless many-jet event reconstruction with symmetry preserving attention networks. Phys. Rev. D. 2022. doi:10.1103/PhysRevD.105.112008. arXiv:2010.09206

  43. [43]

    ABCNet: An attention-based method for particle tagging

    Mikuni, Vinicius and Canelli, Florencia. ABCNet: An attention-based method for particle tagging. Eur. Phys. J. Plus. 2020. doi:10.1140/epjp/s13360-020-00497-3. arXiv:2001.05311

  44. [44]

    Particle Transformer for Jet Tagging

    Qu, Huilin and Li, Congqiao and Qian, Sitian. Particle Transformer for Jet Tagging. 2022. arXiv:2202.03772

  45. [45]

    Automated Approach to Accurate, Precise, and Fast Detector Simulation and Reconstruction

    Dreyer, Etienne and Gross, Eilam and Kobylianskii, Dmitrii and Mikuni, Vinicius and Nachman, Benjamin and Soybelman, Nathalie. Automated Approach to Accurate, Precise, and Fast Detector Simulation and Reconstruction. Phys. Rev. Lett. 2024. doi:10.1103/PhysRevLett.133.211902. arXiv:2406.01620

  46. [46]

    Generating variable length full events from partons

    Qu\'etant, Guillaume and Raine, John Andrew and Leigh, Matthew and Sengupta, Debajyoti and Golling, Tobias. Generating variable length full events from partons. Phys. Rev. D. 2024. doi:10.1103/PhysRevD.110.076023. arXiv:2406.13074

  47. [47]

    Generating particle physics Lagrangians with transformers

    Koay, Yong Sheng and Enberg, Rikard and Moretti, Stefano and Camargo-Molina, Eliel. Generating particle physics Lagrangians with transformers. 2025. arXiv:2501.09729

  48. [48]

    and Zhang, Xiaoyuan

    Dersy, Aur\'elien and Schwartz, Matthew D. and Zhang, Xiaoyuan. Simplifying Polylogarithms with Machine Learning. Int. J. Data Sci. Math. Sci. 2024. doi:10.1142/S2810939223500028. arXiv:2206.04115

  49. [49]

    Learning the simplicity of scattering amplitudes

    Cheung, Clifford and Dersy, Aur\'elien and Schwartz, Matthew D. Learning the simplicity of scattering amplitudes. SciPost Phys. 2025. doi:10.21468/SciPostPhys.18.2.040. arXiv:2408.04720

  50. [50]

    OmniJet- _ C : Learning point cloud calorimeter simulations using generative transformers

    Birk, Joschka and Gaede, Frank and Hallin, Anna and Kasieczka, Gregor and Mozzanica, Martina and Rose, Henning. OmniJet- _ C : Learning point cloud calorimeter simulations using generative transformers. 2025. arXiv:2501.05534

  51. [51]

    HEP-JEPA: A foundation model for collider physics using joint embedding predictive architecture

    Bardhan, Jai and Agrawal, Radhikesh and Tilak, Abhiram and Neeraj, Cyrin and Mitra, Subhadip. HEP-JEPA: A foundation model for collider physics using joint embedding predictive architecture. 2025. arXiv:2502.03933

  52. [52]

    and Rodgers, Jack P

    Wildridge, Andrew J. and Rodgers, Jack P. and Colbert, Ethan M. and yao, Yao and Jung, Andreas W. and Liu, Miaoyuan. Bumblebee: Foundation Model for Particle Physics Discovery. 38th conference on Neural Information Processing Systems. 2024. arXiv:2412.07867

  53. [53]

    Solving key challenges in collider physics with foun- dation models.Phys

    Mikuni, Vinicius and Nachman, Benjamin. Solving key challenges in collider physics with foundation models. Phys. Rev. D. 2025. doi:10.1103/PhysRevD.111.L051504. arXiv:2404.16091

  54. [54]

    Resimulation-based self-supervised learning for pretraining physics foundation models.Phys

    Harris, Philip and Krupa, Jeffrey and Kagan, Michael and Maier, Benedikt and Woodward, Nathaniel. Resimulation-based self-supervised learning for pretraining physics foundation models. Phys. Rev. D. 2025. doi:10.1103/PhysRevD.111.032010. arXiv:2403.07066

  55. [55]

    OmniJet-α: The first cross-task foundation model for particle physics.Machine Learning: Science and Technology, 5:035031, 2024

    Birk, Joschka and Hallin, Anna and Kasieczka, Gregor. OmniJet- : the first cross-task foundation model for particle physics. Mach. Learn. Sci. Tech. 2024. doi:10.1088/2632-2153/ad66ad. arXiv:2403.05618

  56. [56]

    Physics event classification using Large Language Models

    Fanelli, Cristiano and Giroux, James and Moran, Patrick and Nayak, Hemalata and Suresh, Karthik and Walter, Eric. Physics event classification using Large Language Models. JINST. 2024. doi:10.1088/1748-0221/19/07/C07011. arXiv:2404.05752

  57. [57]

    Masked particle modeling on sets: towards self-supervised high energy physics foundation models.Mach

    Golling, Tobias and Heinrich, Lukas and Kagan, Michael and Klein, Samuel and Leigh, Matthew and Osadchy, Margarita and Raine, John Andrew. Masked particle modeling on sets: towards self-supervised high energy physics foundation models. Mach. Learn. Sci. Tech. 2024. doi:10.1088/2632-2153/ad64a8. arXiv:2401.13537

  58. [58]

    Is Tokenization Needed for Masked Particle Modelling?

    Leigh, Matthew and Klein, Samuel and Charton, Fran c ois and Golling, Tobias and Heinrich, Lukas and Kagan, Michael and Ochoa, In\^es and Osadchy, Margarita. Is Tokenization Needed for Masked Particle Modelling?. Mach. Learn. Sci. Tech. 2025. doi:10.1088/2632-2153/addb98. arXiv:2409.12589

  59. [59]

    Finetuning foundation models for joint analysis optimization in High Energy Physics

    Vigl, Matthias and Hartman, Nicole and Heinrich, Lukas. Finetuning foundation models for joint analysis optimization in High Energy Physics. Mach. Learn. Sci. Tech. 2024. doi:10.1088/2632-2153/ad55a3. arXiv:2401.13536

  60. [60]

    Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , month =

    Assran, Mahmoud and Duval, Quentin and Misra, Ishan and Bojanowski, Piotr and Vincent, Pascal and Rabbat, Michael and LeCun, Yann and Ballas, Nicolas , title =. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , month =. 2023 , pages =

  61. [61]

    International Conference on Learning Representations , year=

    An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale , author=. International Conference on Learning Representations , year=

  62. [62]

    2024 , eprint=

    Can Large Language Models Learn the Physics of Metamaterials? An Empirical Study with ChatGPT , author=. 2024 , eprint=

  63. [63]

    2023 , eprint=

    Chain-of-Thought Prompting Elicits Reasoning in Large Language Models , author=. 2023 , eprint=

  64. [64]

    2020 , eprint=

    Language Models are Few-Shot Learners , author=. 2020 , eprint=

  65. [65]

    2023 , eprint=

    Large Language Models are Zero-Shot Reasoners , author=. 2023 , eprint=

  66. [66]

    The Impact of AI in Physics Education: A Comprehensive Review from GCSE to University Levels The Impact of AI in Physics Education , doi =

    Yeadon, Will and Hardy, Tom , year =. The Impact of AI in Physics Education: A Comprehensive Review from GCSE to University Levels The Impact of AI in Physics Education , doi =

  67. [67]

    2023 , eprint=

    Physics simulation capabilities of LLMs , author=. 2023 , eprint=

  68. [68]

    2024 , eprint=

    Xiwu: A Basis Flexible and Learnable LLM for High Energy Physics , author=. 2024 , eprint=

  69. [69]

    Irˇ siˇ c, M

    Unveiling Dark Matter free-streaming at the smallest scales with high redshift Lyman-alpha forest. doi:10.48550/arXiv.2309.04533 , archivePrefix =. 2309.04533 , primaryClass =

  70. [70]

    , keywords =

    New constraints on warm dark matter from the Lyman- forest power spectrum. , keywords =. doi:10.1103/PhysRevD.108.023502 , archivePrefix =. 2209.14220 , primaryClass =

  71. [71]

    High Mass X-ray Binaries and the Cosmic 21-cm Signal: Impact of Host Galaxy Absorption

    High-mass X-ray binaries and the cosmic 21-cm signal: impact of host galaxy absorption. , keywords =. doi:10.1093/mnras/stx943 , archivePrefix =. 1702.00409 , primaryClass =

  72. [72]

    , keywords =

    Cosmology with One Galaxy?. , keywords =. doi:10.3847/1538-4357/ac5d3f , archivePrefix =. 2201.02202 , primaryClass =

  73. [73]

    Radio Galaxy Zoo: Compact and extended radio source classification with deep learning

    Radio Galaxy Zoo: compact and extended radio source classification with deep learning. , keywords =. doi:10.1093/mnras/sty163 , archivePrefix =. 1801.04861 , primaryClass =

  74. [74]

    ML4Astro International Conference , pages=

    Deep Learning 21 cm Lightcones in 3D , author=. ML4Astro International Conference , pages=. 2022 , organization=. doi:10.48550/arXiv.2311.17553 , archivePrefix =. 2311.17553 , primaryClass =

  75. [75]

    , keywords =

    Quantifying uncertainty in deep learning approaches to radio galaxy classification. , keywords =. doi:10.1093/mnras/stac223 , archivePrefix =. 2201.01203 , primaryClass =

  76. [76]

    Galaxy Spectra neural Network (GaSNet). II. Using Deep Learning for Spectral Classification and Redshift Predictions. doi:10.48550/arXiv.2311.04146 , archivePrefix =. 2311.04146 , primaryClass =

  77. [77]

    doi:10.48550/arXiv.2310.02684 , archivePrefix =

    The LoReLi database: 21 cm signal inference with 3D radiative hydrodynamics simulations. doi:10.48550/arXiv.2310.02684 , archivePrefix =. 2310.02684 , primaryClass =

  78. [78]

    2023, Reports on Progress in Physics, 86, 7, 076901

    Machine learning for observational cosmology. Reports on Progress in Physics , keywords =. doi:10.1088/1361-6633/acd2ea , archivePrefix =. 2303.15794 , primaryClass =

  79. [79]

    Machine Learning and Cosmology

    Machine Learning and Cosmology. doi:10.48550/arXiv.2203.08056 , archivePrefix =. 2203.08056 , primaryClass =

  80. [80]

    Simulating the 21-cm signal from reionisation including non-linear ionisations and inhomogeneous recombinations

    Simulating the 21 cm signal from reionization including non-linear ionizations and inhomogeneous recombinations. , keywords =. doi:10.1093/mnras/stv3001 , archivePrefix =. 1510.04280 , primaryClass =

Showing first 80 references.