Pith. sign in

REVIEW 15 cited by

DPOT: Auto-Regressive Denoising Operator Transformer for Large-Scale PDE Pre-Training

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.03542 v4 pith:7SCM5HV3 submitted 2024-03-06 cs.LG cs.NAmath.NA

classification cs.LGcs.NAmath.NA
keywords pre-trainingmodeldataauto-regressivedenoisingdownstreamdpotlarge-scale
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Pre-training has been investigated to improve the efficiency and performance of training neural operators in data-scarce settings. However, it is largely in its infancy due to the inherent complexity and diversity, such as long trajectories, multiple scales and varying dimensions of partial differential equations (PDEs) data. In this paper, we present a new auto-regressive denoising pre-training strategy, which allows for more stable and efficient pre-training on PDE data and generalizes to various downstream tasks. Moreover, by designing a flexible and scalable model architecture based on Fourier attention, we can easily scale up the model for large-scale pre-training. We train our PDE foundation model with up to 0.5B parameters on 10+ PDE datasets with more than 100k trajectories. Extensive experiments show that we achieve SOTA on these benchmarks and validate the strong generalizability of our model to significantly enhance performance on diverse downstream PDE tasks like 3D data. Code is available at \url{https://github.com/thu-ml/DPOT}.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 15 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TIDE: A Physically Diverse 3D Turbulence Benchmark Dataset for Advancing Scientific Machine Learning

    physics.flu-dyn 2026-08 accept novelty 7.0 of 10

    TIDE is a DNS-verified, physically diverse 3D turbulence benchmark with independent ensembles that shows current neural operators barely beat persistence and that low pointwise error does not guarantee physical fidelity.

  2. From Points to Edges: Edge-Conditioned Spectral Operators for Physics-Sensitive PDE Learning

    cs.AI 2026-08 conditional novelty 6.0 of 10

    ESO conditions spectral mode mixing on local pairwise variation statistics, improving neural operator accuracy on PDEs with sharp local structures.

  3. Hybrid Lagrangian-Eulerian Model for Lagrangian Fluid Simulation

    cs.CE 2026-08 conditional novelty 6.0 of 10

    A hybrid Lagrangian-Eulerian graph neural simulator with adaptive downsampling and cross-attention achieves state-of-the-art accuracy and rollout stability on particle-based fluid benchmarks.

  4. Neural operator discovery from heterogeneous trajectories

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Trajectory grouping plus a low-dimensional latent bottleneck lets a neural operator discover each system's hidden governing factors and extrapolate to unseen systems.

  5. Neptuna: A Comprehensive Machine Learning Framework for Benchmarking Complex Multiphase Flows

    physics.flu-dyn 2026-07 conditional novelty 6.0 of 10

    A new 2.4 TB benchmark shows no single ML surrogate dominates on shock-driven multiphase flows, and composite losses with SoftAdapt weighting improve interface and spectral fidelity.

  6. Physics-Informed Neural Quantum Control for Rovibrational Photoassociation in a Morse Molecular System

    quant-ph 2026-06 unverdicted novelty 6.0 of 10

    PINQC optimizes neural-network laser fields via differentiable Schrödinger propagation to photoassociate a Morse molecule, stably reaching l_max=6 versus prior TBQCP limits near l_max=4.

  7. Generative Neural Operators through Diffusion Last Layer

    cs.LG 2026-02 conditional novelty 6.0 of 10

    Diffusion in a learned Karhunen-Loève coefficient space turns deterministic neural operators into generative surrogates with predictive uncertainty and improved rollout stability.

  8. Probabilistic operator learning: generative modeling and uncertainty quantification for foundation models of differential equations

    stat.ML 2025-09 conditional novelty 6.0 of 10

    ICON is shown to compute the posterior predictive mean of differential equation solutions, and a generative extension, GenICON, provides samples from this distribution for uncertainty quantification.

  9. Eigenspectrum Analysis of Neural Networks without Aspect Ratio Bias

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Aspect ratio of weight matrices biases heavy-tail spectral metrics; the new FARMS subsampling method removes this bias and improves downstream layer-wise tuning.

  10. MATEY: multiscale adaptive foundation models for spatiotemporal physical systems

    cs.LG 2024-12 conditional novelty 6.0 of 10

    Adaptive patch sizes chosen by local variance let spatiotemporal transformers match uniform high-resolution accuracy with about half the token sequence length.

  11. BCAT: A Block Causal Transformer for PDE Foundation Models for Fluid Dynamics

    cs.LG 2025-01 conditional novelty 5.5 of 10

    BCAT, a block causal transformer for next-frame prediction, achieves state-of-the-art accuracy on 2D fluid dynamics PDE benchmarks, beating larger foundation models with fewer parameters.

  12. FourierFlow: Frequency-aware Flow Matching for Generative Turbulence Modeling

    cs.LG 2025-06 conditional novelty 5.0 of 10

    FourierFlow improves multi-step generative turbulence modeling by adding a differential-attention branch, a frequency-weighted Fourier mixing branch, and MAE-based feature alignment to a flow matching model.

  13. Latent Mamba Operator for Partial Differential Equations

    cs.LG 2025-05 conditional novelty 5.0 of 10

    LaMO replaces attention in latent-token neural operators with bidirectional state-space models and reports consistent accuracy gains on six PDE benchmarks.

  14. Predicting Change, Not States: An Alternate Framework for Neural PDE Surrogates

    cs.LG 2024-12 conditional novelty 5.0 of 10

    Predicting the temporal derivative and integrating it with an ODE solver improves accuracy and stability of neural PDE surrogates compared with direct next-state prediction.

  15. Machine learning for modelling unstructured grid data in computational physics: a review

    cs.LG 2025-02 conditional novelty 2.0 of 10

    A broad review of machine learning techniques for modeling unstructured mesh data in computational physics, with a taxonomy, a qualitative comparison, and a list of public benchmarks.

Pith tools