J-JEPA pretraining on 1M jets modestly improves top jet tagging versus from-scratch training, but gains are inconsistent for the strongest baseline model.
Is Tokenization Needed for Masked Particle Modelling?
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
In this work, we significantly enhance masked particle modeling (MPM), a self-supervised learning scheme for constructing highly expressive representations of unordered sets relevant to developing foundation models for high-energy physics. In MPM, a model is trained to recover the missing elements of a set, a learning objective that requires no labels and can be applied directly to experimental data. We achieve significant performance improvements over previous work on MPM by addressing inefficiencies in the implementation and incorporating a more powerful decoder. We compare several pre-training tasks and introduce new reconstruction methods that utilize conditional generative models without data tokenization or discretization. We show that these new methods outperform the tokenized learning objective from the original MPM on a new test bed for foundation models for jets, which includes using a wide variety of downstream tasks relevant to jet physics, such as classification, secondary vertex finding, and track identification.
citation-role summary
citation-polarity summary
fields
hep-ph 1years
2024 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Learning Symmetry-Independent Jet Representations via Jet-Based Joint Embedding Predictive Architecture
J-JEPA pretraining on 1M jets modestly improves top jet tagging versus from-scratch training, but gains are inconsistent for the strongest baseline model.