REVIEW 10 cited by
OmniJet-$\alpha$: The first cross-task foundation model for particle physics
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
Foundation models are multi-dataset and multi-task machine learning methods that once pre-trained can be fine-tuned for a large variety of downstream applications. The successful development of such general-purpose models for physics data would be a major breakthrough as they could improve the achievable physics performance while at the same time drastically reduce the required amount of training time and data. We report significant progress on this challenge on several fronts. First, a comprehensive set of evaluation methods is introduced to judge the quality of an encoding from physics data into a representation suitable for the autoregressive generation of particle jets with transformer architectures (the common backbone of foundation models). These measures motivate the choice of a higher-fidelity tokenization compared to previous works. Finally, we demonstrate transfer learning between an unsupervised problem (jet generation) and a classic supervised task (jet tagging) with our new OmniJet-$\alpha$ model. This is the first successful transfer between two different and actively studied classes of tasks and constitutes a major step in the building of foundation models for particle physics.
Forward citations
Cited by 10 Pith papers
-
Predict before you train: Scaling Laws for particle physics foundation models
A Chinchilla-style law fit on ParticleViT runs below 10^19 FLOPs predicts held-out pretraining loss within ~1% at >100× compute and tracks downstream jet-tagging rejection.
-
Learning Standard Model structure from LHC data with Riemannian flow matching
ShellFlow, a Riemannian flow-matching transformer fed only on-shell and invariant-mass priors and ~8×10^8 recorded ATLAS events, reproduces the SM's dilepton resonances, Weinberg angle, and top/W mass peaks in a singl...
-
Neural Scaling Laws for Jet Generation
Scaling laws hold logarithmically for model size in autoregressive jet generation, with next-token loss correlating to physical metrics via sliced Wasserstein distance, but show weaker scaling for dataset size and com...
-
Learning transferable event representations for charmed baryon physics at BESIII
A Particle Transformer pre-trained on simulated Lambda_c events transfers across 12 decay channels, improving classification and momentum-direction regression over training from scratch in low-statistics regimes.
-
SPADE: Split-and-Delay Embeddings for Autoregressive High-Granularity Calorimeter Simulation
SPADE is a split-and-delay embedding technique for multi-feature autoregressive transformers that achieves competitive performance on high-granularity calorimeter shower simulation.
-
Explicit or Implicit? Encoding Physics at the Precision Frontier
On three precision classification tasks — reweighting-based unfolding, likelihood-ratio estimation, and weakly supervised anomaly detection — a Lorentz-equivariant transformer and a pretrained foundation model perform...
-
A universal vision transformer for fast calorimeter simulations
A vision-transformer flow-matching model generates calorimeter showers across regular and irregular detector geometries at millisecond speeds, and pretraining plus fine-tuning cuts training cost by about half.
-
Enhancing next token prediction based pre-training for jet foundation models
Using continuous particle features as input and combining next-token with masked-token pre-training markedly improves classification accuracy of the OmniJet jet foundation model without visibly hurting its generative quality.
-
Are We Ready for AI-Driven Discovery? AI Verification Before the Next Fundamental Physics Breakthrough
Verification of ML in fundamental physics is essential precisely when models enter statistical modeling, inference, or hypothesis testing, and is bounded by unavoidable inductive bias, sample complexity, and experimen...
-
HEPTAPOD: Orchestrating High Energy Physics Workflows Towards Autonomous Agency
HEPTAPOD uses LLM agents to drive FeynRules, MadGraph, Pythia, and analysis tools through schema-validated tool calls and run-card templates, demonstrated on a leptoquark signal scan.
Discussion (0). Sign in to comment.