Pith. sign in

REVIEW 5 cited by

Re-Simulation-based Self-Supervised Learning for Pre-Training Foundation Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.07066 v2 pith:KUTLYBKM submitted 2024-03-11 hep-ph cs.LGhep-ex

classification hep-phcs.LGhep-ex
keywords learningdownstreamrs3lself-supervisedtasksavailabledatafoundation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Self-Supervised Learning (SSL) is at the core of training modern large machine learning models, providing a scheme for learning powerful representations that can be used in a variety of downstream tasks. However, SSL strategies must be adapted to the type of training data and downstream tasks required. We propose RS3L ("Re-simulation-based self-supervised representation learning"), a novel simulation-based SSL strategy that employs a method of re-simulation to drive data augmentation for contrastive learning in the physical sciences, particularly, in fields that rely on stochastic simulators. By intervening in the middle of the simulation process and re-running simulation components downstream of the intervention, we generate multiple realizations of an event, thus producing a set of augmentations covering all physics-driven variations available in the simulator. Using experiments from high-energy physics, we explore how this strategy may enable the development of a foundation model; we show how RS3L pre-training enables powerful performance in downstream tasks such as discrimination of a variety of objects and uncertainty mitigation. In addition to our results, we make the RS3L dataset publicly available for further studies on how to improve SSL strategies.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Explicit or Implicit? Encoding Physics at the Precision Frontier

    hep-ph 2026-03 conditional novelty 6.0 of 10

    On three precision classification tasks — reweighting-based unfolding, likelihood-ratio estimation, and weakly supervised anomaly detection — a Lorentz-equivariant transformer and a pretrained foundation model perform...

  2. Enhancing next token prediction based pre-training for jet foundation models

    hep-ph 2025-12 conditional novelty 6.0 of 10

    Using continuous particle features as input and combining next-token with masked-token pre-training markedly improves classification accuracy of the OmniJet jet foundation model without visibly hurting its generative quality.

  3. Pretrained Event Classification Model for High Energy Physics Analysis

    hep-ph 2024-12 unverdicted novelty 6.0 of 10

    A pretrained GNN on 120M simulated LHC events improves downstream event-classification accuracy when training data are scarce, with benefits shrinking as data grow.

  4. Learning Symmetry-Independent Jet Representations via Jet-Based Joint Embedding Predictive Architecture

    hep-ph 2024-12 conditional novelty 6.0 of 10

    J-JEPA pretraining on 1M jets modestly improves top jet tagging versus from-scratch training, but gains are inconsistent for the strongest baseline model.

  5. Are We Ready for AI-Driven Discovery? AI Verification Before the Next Fundamental Physics Breakthrough

    physics.data-an 2026-07 accept novelty 4.0 of 10

    Verification of ML in fundamental physics is essential precisely when models enter statistical modeling, inference, or hypothesis testing, and is bounded by unavoidable inductive bias, sample complexity, and experimen...

Pith tools