A Particle Transformer jet tagger contains a sparse six-head circuit whose source-relay-readout structure recovers most performance and whose residual stream preferentially encodes 2-prong energy correlators.
hub Tool reference
Particle Transformer for Jet Tagging
Tool reference. 78% of classified Pith citations use this work as a method, library, or software dependency, not as a substantive claim.
abstract
Jet tagging is a critical yet challenging classification task in particle physics. While deep learning has transformed jet tagging and significantly improved performance, the lack of a large-scale public dataset impedes further enhancement. In this work, we present JetClass, a new comprehensive dataset for jet tagging. The JetClass dataset consists of 100 M jets, about two orders of magnitude larger than existing public datasets. A total of 10 types of jets are simulated, including several types unexplored for tagging so far. Based on the large dataset, we propose a new Transformer-based architecture for jet tagging, called Particle Transformer (ParT). By incorporating pairwise particle interactions in the attention mechanism, ParT achieves higher tagging performance than a plain Transformer and surpasses the previous state-of-the-art, ParticleNet, by a large margin. The pre-trained ParT models, once fine-tuned, also substantially enhance the performance on two widely adopted jet tagging benchmarks. The dataset, code and models are publicly available at https://github.com/jet-universe/particle_transformer.
hub tools
citation-role summary
citation-polarity summary
representative citing papers
PLuM multimodal transformer improves top and H->bb jet tagging by jointly processing particle constituents and Lund plane splittings, yielding 25% higher background rejection at 25% di-Higgs efficiency.
PHAT-JeT combines geometric message-passing with hierarchical patch attention to reach state-of-the-art accuracy and background rejection among resource-constrained jet tagging models on four benchmarks.
LP2B encoding converts Lund plane jet representations into Bloch sphere qubit states, enabling a QTTN that matches classical LundNet performance on polarization tagging and W/top tagging with three orders of magnitude fewer parameters and improved low-data regime results.
IAFormer uses boost-invariant pairwise quantities and differential attention to create a sparse Transformer that achieves state-of-the-art classification on top-quark and quark-gluon jet datasets while using over an order of magnitude fewer parameters than prior Particle Transformer models.
Quantity-grounded multi-agent decomposition makes LLM-generated collider analysis code inspectable and reliable with 14B-scale models, outperforming prior single-prompt approaches.
Point-cloud ML models classify charm- and bottom-origin electrons at ~80% purity for 40% efficiency, outperforming a BDT baseline, with performance limited by intrinsic decay similarity.
LLM embeddings condition a generative transformer to enable faster convergence, better performance, and generalization to unseen LHC processes using a single model.
Transformer models trained on FCC-ee CLD simulation reconstruct hadronic taus with per-mille mis-ID, F1 up to 0.95, sub-per-mille charge errors, and percent-level transverse momentum resolution.
Collider events are represented as multivectors in Cl(1,3) ⊗ V_flav whose grade projections recover standard observables, intended as input for equivariant foundation models.
Neural networks for HEP tasks can be fooled at significant rates by subtle perturbations inside uncertainty envelopes, revealing hidden systematics not captured by conventional methods.
Semi-leptonic h→VV* decays retain an effective two-qutrit quantum description under NLO QCD and electroweak corrections, unlike the fully leptonic h→4ℓ channel.
PaRT achieves >50% tagging efficiency for boosted H->WW jets at 1% background efficiency, decorrelated from jet mass, with data-to-simulation scale factors of 0.9-1.0 on 138 fb^{-1} of 13 TeV collisions.
Higgsformer achieves AUC 0.855 on t tbar H vs t tbar classification from raw hits, matching a Delphes-based Particle Transformer at ~40% b-tagging efficiency.
Equivariant jet taggers suppress frame-dependent pseudorapidity while encoding jet mass and N-subjettiness strongly, with bivector channels negligible and vector channels dominant for top tagging.
Presents a reusable open-source framework for mapping quantized transformer layers to AMD Versal AI Engine tiles for jet tagging at the LHC.
Release of an AI-ready dataset containing approximately 660,000 reconstructed polarized e+e- collision events at 91.2 GeV from the SLD experiment, translated from legacy formats with accompanying digitized documentation.
A hyper-graph neural network improves discrimination of four-top production at 13 TeV, raising expected significance from 5.13 to 9.11 and enabling projected 95% CL limits on five dimension-six SMEFT Wilson coefficients at current and HL-LHC luminosities.
Explainability techniques applied to LundNet show that assigned node importance correlates with classical jet substructure observables such as N-subjettiness ratios and energy correlation functions, with shifts across transverse-momentum regimes.
An edge-weighted multi-graph GNN (E-PCN/KIGNet) that encodes four jet kinematic variables improves JetClass jet-tagging accuracy and attributes most predictions to angular separation and transverse momentum.
SAL-T enhances the linformer with spatially aware kinematic partitioning and convolutions to match full-attention transformer performance on jet tagging while keeping linear complexity and lower latency.
Deep learning on all particles via holistic analysis and Advanced Color Singlet Identification improves Higgs signal extraction up to sixfold in high-energy collisions.
Transformers trained on cosmic ray simulations learn physically plausible features in positional encodings for symmetric air showers and in attention mechanisms for galaxy-origin particles.
Adversarial training enhances robustness of jet tagging classifiers while preserving performance, with loss surface geometry providing insights into correlations and vulnerability.
citing papers explorer
-
Dissecting Jet-Tagger Through Mechanistic Interpretability
A Particle Transformer jet tagger contains a sparse six-head circuit whose source-relay-readout structure recovers most performance and whose residual stream preferentially encodes 2-prong energy correlators.
-
Particle-Lund Multimodality in Jet Taggers
PLuM multimodal transformer improves top and H->bb jet tagging by jointly processing particle constituents and Lund plane splittings, yielding 25% higher background rejection at 25% di-Higgs efficiency.
-
Patch Hierarchical Attention Transformer for Efficient Particle Jet Tagging
PHAT-JeT combines geometric message-passing with hierarchical patch attention to reach state-of-the-art accuracy and background rejection among resource-constrained jet tagging models on four benchmarks.
-
Lund Plane to Bloch (LP2B) Encoding for Object and Polarization Tagging with Quantum Jet Substructure
LP2B encoding converts Lund plane jet representations into Bloch sphere qubit states, enabling a QTTN that matches classical LundNet performance on polarization tagging and W/top tagging with three orders of magnitude fewer parameters and improved low-data regime results.
-
IAFormer: Interaction-Aware Transformer network for collider data analysis
IAFormer uses boost-invariant pairwise quantities and differential attention to create a sparse Transformer that achieves state-of-the-art classification on top-quark and quark-gluon jet datasets while using over an order of magnitude fewer parameters than prior Particle Transformer models.
-
Articulating Assumptions in AI-Generated Scientific Analyses through Task Decomposition
Quantity-grounded multi-agent decomposition makes LLM-generated collider analysis code inspectable and reliable with 14B-scale models, outperforming prior single-prompt approaches.
-
Heavy-Flavor Electron Classification Using Hadronic Environment as Point Cloud
Point-cloud ML models classify charm- and bottom-origin electrons at ~80% purity for 40% efficiency, outperforming a BDT baseline, with performance limited by intrinsic decay similarity.
-
One Generator, Any Process: LLM-Conditioning for the LHC
LLM embeddings condition a generative transformer to enable faster convergence, better performance, and generalization to unseen LHC processes using a single model.
-
ParticleTransformer is all you need for reconstructing hadronic tau leptons
Transformer models trained on FCC-ee CLD simulation reconstruct hadronic taus with per-mille mis-ID, F1 up to 0.95, sub-per-mille charge errors, and percent-level transverse momentum resolution.
-
Geometric algebra as the input language of collider foundation models
Collider events are represented as multivectors in Cl(1,3) ⊗ V_flav whose grade projections recover standard observables, intended as input for equivariant foundation models.
-
Uncovering Hidden Systematics in Neural Network Models for High Energy Physics
Neural networks for HEP tasks can be fooled at significant rates by subtle perturbations inside uncertainty envelopes, revealing hidden systematics not captured by conventional methods.
-
Quantum Tomography and Entanglement in Semi-Leptonic $h\to VV^*$ Decays at Higher Orders
Semi-leptonic h→VV* decays retain an effective two-qutrit quantum description under NLO QCD and electroweak corrections, unlike the fully leptonic h→4ℓ channel.
-
Particle transformers for identifying Lorentz-boosted Higgs bosons decaying to a pair of W bosons
PaRT achieves >50% tagging efficiency for boosted H->WW jets at 1% background efficiency, decorrelated from jet mass, with data-to-simulation scale factors of 0.9-1.0 on 138 fb^{-1} of 13 TeV collisions.
-
Hits to Higgs: Hit-Level Higgs Classification from Raw LHC Detector Data Using Higgsformer
Higgsformer achieves AUC 0.855 on t tbar H vs t tbar classification from raw hits, matching a Delphes-based Particle Transformer at ~40% b-tagging efficiency.
-
What Do Lorentz-Equivariant Jet Taggers Learn?
Equivariant jet taggers suppress frame-dependent pseudorapidity while encoding jet mass and N-subjettiness strongly, with bivector channels negligible and vector channels dominant for top tagging.
-
Reconfigurable Computing Challenge: Transformer for Jet Tagging on Versal AI Engines
Presents a reusable open-source framework for mapping quantized transformer layers to AMD Versal AI Engine tiles for jet tagging at the LHC.
-
An AI-ready, Polarized Electron-Positron Collision Dataset
Release of an AI-ready dataset containing approximately 660,000 reconstructed polarized e+e- collision events at 91.2 GeV from the SLD experiment, translated from legacy formats with accompanying digitized documentation.
-
Probing SMEFT Operators through $t\bar{t}t\bar{t}$ Production with Hyper-Graph Neural Networks at the LHC
A hyper-graph neural network improves discrimination of four-top production at 13 TeV, raising expected significance from 5.13 to 9.11 and enabling projected 95% CL limits on five dimension-six SMEFT Wilson coefficients at current and HL-LHC luminosities.
-
Explainable AI for Jet Tagging: A Comparative Study of GNNExplainer, GNNShap, and GradCAM for Jet Tagging in the Lund Jet Plane
Explainability techniques applied to LundNet show that assigned node importance correlates with classical jet substructure observables such as N-subjettiness ratios and energy correlation functions, with shifts across transverse-momentum regimes.
-
KIGNet: Physics-Motivated Multi-Graph Representation Learning for Explainable Jet Tagging
An edge-weighted multi-graph GNN (E-PCN/KIGNet) that encodes four jet kinematic variables improves JetClass jet-tagging accuracy and attributes most predictions to angular separation and transverse momentum.
-
Spatially Aware Linear Transformer (SAL-T) for Particle Jet Tagging
SAL-T enhances the linformer with spatially aware kinematic partitioning and convolutions to match full-attention transformer performance on jet tagging while keeping linear complexity and lower latency.
-
Learning from all particles in high-energy collisions
Deep learning on all particles via holistic analysis and Advanced Color Singlet Identification improves Higgs signal extraction up to sixfold in high-energy collisions.
-
What exactly did the Transformer learn from our physics data?
Transformers trained on cosmic ray simulations learn physically plausible features in positional encodings for symmetric air showers and in attention mechanisms for galaxy-origin particles.
-
Improving robustness of jet tagging algorithms with adversarial training: exploring the loss surface
Adversarial training enhances robustness of jet tagging classifiers while preserving performance, with loss surface geometry providing insights into correlations and vulnerability.
-
Application of Deep Learning to Jet Charge Discrimination
Graph neural network achieves AUC of 0.883 for up versus anti-up quark jet charge discrimination in controlled QCD simulations.
-
Search for invisible decays of light mesons via $J/\psi \to VP$ $(V=\omega/\phi,P=\eta/\eta')$ decays at STCF
Monte Carlo projections set 90% CL upper limits of 3.7e-7, 8.9e-7, 1.8e-7 and 4.1e-7 on invisible branching fractions of omega, phi, eta and eta' at STCF.
-
Improved results on Higgs boson pair production in the 4b final state
CMS sets an observed upper limit of 4.4 on the HH signal strength μ_HH in the 4b final state at 13.6 TeV, improving prior LHC results by more than a factor of two in the resolved topology.
-
Comprehensive Mass Predictions: From Triply Heavy Baryons to Pentaquarks
Machine learning models trained on known hadron data and an extended Gürsey-Radicati mass formula predict masses for triply heavy baryons and numerous pentaquark states, agreeing with available data and forecasting unobserved states.
-
Data Preservation in High Energy Physics: Global Report 2026
The 2026 DPHEP report records substantial progress in HEP data preservation, including modern reanalyses of LEP legacy data and expanding open-data policies, alongside sustainability challenges.
-
Open LHC Monte Carlo Event Generation
A review of initiatives to make LHC Monte Carlo event generations available as open data to minimize redundant simulations and resource use.
-
Rare top quark production and top quark properties in ATLAS and CMS
ATLAS and CMS have performed analyses of rare top quark production modes that constrain top couplings and search for new physics.