A Particle Transformer jet tagger contains a sparse six-head circuit whose source-relay-readout structure recovers most performance and whose residual stream preferentially encodes 2-prong energy correlators.
Interpreting Transformers for Jet Tagging
5 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
roles
background 2polarities
background 2representative citing papers
Equivariant jet taggers suppress frame-dependent pseudorapidity while encoding jet mass and N-subjettiness strongly, with bivector channels negligible and vector channels dominant for top tagging.
Explainability techniques applied to LundNet show that assigned node importance correlates with classical jet substructure observables such as N-subjettiness ratios and energy correlation functions, with shifts across transverse-momentum regimes.
JEDI-linear is a linear-complexity GNN for FPGA jet tagging that reports sub-60 ns latency, higher accuracy than prior designs, and no DSP usage while meeting HL-LHC CMS Level-1 trigger requirements.
Transformers trained on cosmic ray simulations learn physically plausible features in positional encodings for symmetric air showers and in attention mechanisms for galaxy-origin particles.
citing papers explorer
-
Dissecting Jet-Tagger Through Mechanistic Interpretability
A Particle Transformer jet tagger contains a sparse six-head circuit whose source-relay-readout structure recovers most performance and whose residual stream preferentially encodes 2-prong energy correlators.
-
What Do Lorentz-Equivariant Jet Taggers Learn?
Equivariant jet taggers suppress frame-dependent pseudorapidity while encoding jet mass and N-subjettiness strongly, with bivector channels negligible and vector channels dominant for top tagging.
-
Explainable AI for Jet Tagging: A Comparative Study of GNNExplainer, GNNShap, and GradCAM for Jet Tagging in the Lund Jet Plane
Explainability techniques applied to LundNet show that assigned node importance correlates with classical jet substructure observables such as N-subjettiness ratios and energy correlation functions, with shifts across transverse-momentum regimes.
-
JEDI-linear: Fast and Efficient Graph Neural Networks for Jet Tagging on FPGAs
JEDI-linear is a linear-complexity GNN for FPGA jet tagging that reports sub-60 ns latency, higher accuracy than prior designs, and no DSP usage while meeting HL-LHC CMS Level-1 trigger requirements.
-
What exactly did the Transformer learn from our physics data?
Transformers trained on cosmic ray simulations learn physically plausible features in positional encodings for symmetric air showers and in attention mechanisms for galaxy-origin particles.