Pith. sign in

REVIEW 13 cited by

Poseidon: Efficient Foundation Models for PDEs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.19101 v2 pith:HD7TYSJB submitted 2024-05-29 cs.LG

Poseidon: Efficient Foundation Models for PDEs

classification cs.LG
keywords poseidonpdesdownstreammodelpretrainingfoundationwellcamlab-ethz
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

We introduce Poseidon, a foundation model for learning the solution operators of PDEs. It is based on a multiscale operator transformer, with time-conditioned layer norms that enable continuous-in-time evaluations. A novel training strategy leveraging the semi-group property of time-dependent PDEs to allow for significant scaling-up of the training data is also proposed. Poseidon is pretrained on a diverse, large scale dataset for the governing equations of fluid dynamics. It is then evaluated on a suite of 15 challenging downstream tasks that include a wide variety of PDE types and operators. We show that Poseidon exhibits excellent performance across the board by outperforming baselines significantly, both in terms of sample efficiency and accuracy. Poseidon also generalizes very well to new physics that is not seen during pretraining. Moreover, Poseidon scales with respect to model and data size, both for pretraining and for downstream tasks. Taken together, our results showcase the surprising ability of Poseidon to learn effective representations from a very small set of PDEs during pretraining in order to generalize well to unseen and unrelated PDEs downstream, demonstrating its potential as an effective, general purpose PDE foundation model. Finally, the Poseidon model as well as underlying pretraining and downstream datasets are open sourced, with code being available at https://github.com/camlab-ethz/poseidon and pretrained models and datasets at https://huggingface.co/camlab-ethz.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Function graph transformers universally approximate operators between function spaces

    cs.LG 2026-05 unverdicted novelty 8.0

    Function graph transformers use graph measures to provide a measure-theoretic framework where standard transformer components universally approximate operators between function spaces while preserving single-valued fu...

  2. OmniMol: Transferring Particle Physics Knowledge to Molecular Dynamics with Point-Edge Transformers

    physics.chem-ph 2026-01 unverdicted novelty 7.0

    OmniMol transfers a billion-jet pre-trained PET foundation model from HEP to molecular dynamics via an interaction-matrix attention bias, delivering strong performance on the oMol dataset with minimal fine-tuning and ...

  3. WLNO: Wavelet-Laplace Neural Operator for Solving Partial Differential Equations

    cs.LG 2026-05 unverdicted novelty 6.0

    WLNO augments LNO with a parallel Haar wavelet branch and learnable gate to capture multi-scale spatial features, outperforming LNO on five PDE benchmarks especially those with sharp structures.

  4. ARC-STAR: Auditable Post-Hoc Correction for PDE Foundation Models

    cs.LG 2026-05 unverdicted novelty 6.0

    ARC-STAR is a frozen, auditable post-hoc correction method that reduces velocity rollout error by at least 36x over raw Poseidon across five flow benchmarks using global and local stages with budget-aware triage.

  5. ARC-STAR: Auditable Post-Hoc Correction for PDE Foundation Models

    cs.LG 2026-05 unverdicted novelty 6.0

    ARC-STAR is an auditable, budget-aware post-hoc correction method that reduces velocity rollout error by at least 36x over raw Poseidon across five flow benchmarks.

  6. ARC-STAR: Auditable Post-Hoc Correction for PDE Foundation Models

    cs.LG 2026-05 unverdicted novelty 6.0

    ARC-STAR reduces velocity rollout error by at least 36x over raw Poseidon across all tested regime cells via auditable global and local correction stages on five flow benchmarks.

  7. A Hybridizable Neural Time Integrator for Stable Autoregressive Forecasting

    cs.LG 2026-04 unverdicted novelty 6.0

    A hybrid transformer-FEM integrator provides provable discrete energy preservation and gradient bounds for stable autoregressive forecasting of chaotic systems, with 65x fewer parameters and 9000x speedup in a fusion ...

  8. SuperWing: a comprehensive transonic wing dataset for data-driven aerodynamic design

    cs.LG 2025-12 conditional novelty 6.0

    SuperWing supplies 4,239 diverse wing shapes and 28,856 flow-field solutions that let Transformer models predict surface aerodynamics to 2.5 drag-count error and generalize zero-shot to DLR-F6 and NASA CRM wings.

  9. A Two-Phase Deep Learning Framework for Adaptive Time-Stepping in High-Speed Flow Modeling

    cs.LG 2025-06 unverdicted novelty 6.0

    ShockCast is a two-phase ML method that predicts adaptive timestep sizes to model high-speed flows with shocks more efficiently than fixed-step approaches.

  10. Neuro-Symbolic AI for Analytical Solutions of Differential Equations

    cs.LG 2025-02 unverdicted novelty 6.0

    SIGS is a neuro-symbolic framework that discovers analytical solutions to PDEs by generating grammar-constrained expressions, embedding them in a topology-regularised latent manifold, and refining structure and coeffi...

  11. Sparse probes and murky physics: a case study of interpretability challenges in a foundation model for continuum dynamics

    cs.LG 2026-06 unverdicted novelty 5.0

    Case study applies SAE probing with enstrophy triage to a continuum-dynamics foundation model and reports intermittent feature consistency that does not align with standard physics while linking some output discrepanc...

  12. Sequential Physics-Constrained Neural Operator Forward Modeling for the $\textit{Norne}$ Reservoir System

    cs.LG 2026-05 unverdicted novelty 5.0

    Presents sequential physics-constrained neural operator models for the Norne reservoir with theoretical stability guarantees and empirical accuracy exceeding 0.99 R² for oil production predictions alongside a 10,000x ...

  13. Towards Scaling Law Analysis For Spatiotemporal Weather Data

    cs.LG 2026-04 unverdicted novelty 5.0

    Scaling laws for weather models exhibit strong cross-channel and cross-horizon heterogeneity, where globally pooled metrics appear favorable while many individual channels degrade at longer leads.