REVIEW 13 cited by
Poseidon: Efficient Foundation Models for PDEs
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Poseidon: Efficient Foundation Models for PDEs
read the original abstract
We introduce Poseidon, a foundation model for learning the solution operators of PDEs. It is based on a multiscale operator transformer, with time-conditioned layer norms that enable continuous-in-time evaluations. A novel training strategy leveraging the semi-group property of time-dependent PDEs to allow for significant scaling-up of the training data is also proposed. Poseidon is pretrained on a diverse, large scale dataset for the governing equations of fluid dynamics. It is then evaluated on a suite of 15 challenging downstream tasks that include a wide variety of PDE types and operators. We show that Poseidon exhibits excellent performance across the board by outperforming baselines significantly, both in terms of sample efficiency and accuracy. Poseidon also generalizes very well to new physics that is not seen during pretraining. Moreover, Poseidon scales with respect to model and data size, both for pretraining and for downstream tasks. Taken together, our results showcase the surprising ability of Poseidon to learn effective representations from a very small set of PDEs during pretraining in order to generalize well to unseen and unrelated PDEs downstream, demonstrating its potential as an effective, general purpose PDE foundation model. Finally, the Poseidon model as well as underlying pretraining and downstream datasets are open sourced, with code being available at https://github.com/camlab-ethz/poseidon and pretrained models and datasets at https://huggingface.co/camlab-ethz.
Forward citations
Cited by 13 Pith papers
-
Function graph transformers universally approximate operators between function spaces
Function graph transformers use graph measures to provide a measure-theoretic framework where standard transformer components universally approximate operators between function spaces while preserving single-valued fu...
-
OmniMol: Transferring Particle Physics Knowledge to Molecular Dynamics with Point-Edge Transformers
OmniMol transfers a billion-jet pre-trained PET foundation model from HEP to molecular dynamics via an interaction-matrix attention bias, delivering strong performance on the oMol dataset with minimal fine-tuning and ...
-
WLNO: Wavelet-Laplace Neural Operator for Solving Partial Differential Equations
WLNO augments LNO with a parallel Haar wavelet branch and learnable gate to capture multi-scale spatial features, outperforming LNO on five PDE benchmarks especially those with sharp structures.
-
ARC-STAR: Auditable Post-Hoc Correction for PDE Foundation Models
ARC-STAR is a frozen, auditable post-hoc correction method that reduces velocity rollout error by at least 36x over raw Poseidon across five flow benchmarks using global and local stages with budget-aware triage.
-
ARC-STAR: Auditable Post-Hoc Correction for PDE Foundation Models
ARC-STAR is an auditable, budget-aware post-hoc correction method that reduces velocity rollout error by at least 36x over raw Poseidon across five flow benchmarks.
-
ARC-STAR: Auditable Post-Hoc Correction for PDE Foundation Models
ARC-STAR reduces velocity rollout error by at least 36x over raw Poseidon across all tested regime cells via auditable global and local correction stages on five flow benchmarks.
-
A Hybridizable Neural Time Integrator for Stable Autoregressive Forecasting
A hybrid transformer-FEM integrator provides provable discrete energy preservation and gradient bounds for stable autoregressive forecasting of chaotic systems, with 65x fewer parameters and 9000x speedup in a fusion ...
-
SuperWing: a comprehensive transonic wing dataset for data-driven aerodynamic design
SuperWing supplies 4,239 diverse wing shapes and 28,856 flow-field solutions that let Transformer models predict surface aerodynamics to 2.5 drag-count error and generalize zero-shot to DLR-F6 and NASA CRM wings.
-
A Two-Phase Deep Learning Framework for Adaptive Time-Stepping in High-Speed Flow Modeling
ShockCast is a two-phase ML method that predicts adaptive timestep sizes to model high-speed flows with shocks more efficiently than fixed-step approaches.
-
Neuro-Symbolic AI for Analytical Solutions of Differential Equations
SIGS is a neuro-symbolic framework that discovers analytical solutions to PDEs by generating grammar-constrained expressions, embedding them in a topology-regularised latent manifold, and refining structure and coeffi...
-
Sparse probes and murky physics: a case study of interpretability challenges in a foundation model for continuum dynamics
Case study applies SAE probing with enstrophy triage to a continuum-dynamics foundation model and reports intermittent feature consistency that does not align with standard physics while linking some output discrepanc...
-
Sequential Physics-Constrained Neural Operator Forward Modeling for the $\textit{Norne}$ Reservoir System
Presents sequential physics-constrained neural operator models for the Norne reservoir with theoretical stability guarantees and empirical accuracy exceeding 0.99 R² for oil production predictions alongside a 10,000x ...
-
Towards Scaling Law Analysis For Spatiotemporal Weather Data
Scaling laws for weather models exhibit strong cross-channel and cross-horizon heterogeneity, where globally pooled metrics appear favorable while many individual channels degrade at longer leads.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.