Pith. sign in

REVIEW 2 cited by

Towards Foundation Models for Scientific Machine Learning: Characterizing Scaling and Transfer Behavior

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.00258 v1 pith:CZH7SXLT submitted 2023-06-01 cs.LG cs.NAmath.NA

classification cs.LGcs.NAmath.NA
keywords learningdownstreammodelsapplicationsbehaviormachinemodelpre-trained
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Pre-trained machine learning (ML) models have shown great performance for a wide range of applications, in particular in natural language processing (NLP) and computer vision (CV). Here, we study how pre-training could be used for scientific machine learning (SciML) applications, specifically in the context of transfer learning. We study the transfer behavior of these models as (i) the pre-trained model size is scaled, (ii) the downstream training dataset size is scaled, (iii) the physics parameters are systematically pushed out of distribution, and (iv) how a single model pre-trained on a mixture of different physics problems can be adapted to various downstream applications. We find that-when fine-tuned appropriately-transfer learning can help reach desired accuracy levels with orders of magnitude fewer downstream examples (across different tasks that can even be out-of-distribution) than training from scratch, with consistent behavior across a wide range of downstream examples. We also find that fine-tuning these models yields more performance gains as model size increases, compared to training from scratch on new downstream tasks. These results hold for a broad range of PDE learning tasks. All in all, our results demonstrate the potential of the "pre-train and fine-tune" paradigm for SciML problems, demonstrating a path towards building SciML foundation models. We open-source our code for reproducibility.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Towards a Physics Foundation Model

    cs.LG 2025-09 conditional novelty 6.0 of 10

    A single transformer-based model, GPhyT, trained on diverse 2D simulation data, predicts next states across several fluid and heat-transfer systems and extrapolates to similar unseen regimes with plausible results.

  2. HEP-JEPA: A foundation model for collider physics using joint embedding predictive architecture

    cs.LG 2025-02 conditional novelty 6.0 of 10

    A JEPA-style self-supervised transformer for collider jets improves few-shot classification and transfers to top and quark-gluon tagging, yet remains behind specialized taggers.

Pith tools