Pith. sign in

REVIEW 2 major objections 5 minor 4 references

GENERIC-FNO embeds full thermodynamic structure into neural operators so energy is conserved and entropy produced by construction in function space.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-12 14:43 UTC pith:ETUKK62W

load-bearing objection First full GENERIC neural operator in function space with by-construction degeneracy; algebra and experiments hold, expressivity limits are disclosed. the 2 major comments →

arxiv 2606.08343 v3 pith:ETUKK62W submitted 2026-06-06 cs.LG

GENERIC-FNO: Embedding Energy Conservation and Entropy Production into Fourier Neural Operators

classification cs.LG
keywords GENERICmetriplectic systemsFourier neural operatorsstructure-preserving learningenergy conservationentropy productiondegeneracy conditionsneural operators
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Neural operators that roll out PDE trajectories often drift off the physical manifold: energy wanders and quantities that should decay can grow. This paper claims that the full GENERIC (metriplectic) structure of nonequilibrium thermodynamics can be built directly into a neural operator in function space, not just into finite-dimensional models. The construction learns energy and entropy as neural operators and realizes the reversible and irreversible generators as diagonal Fourier multipliers sandwiched between rank-one projections that force the two degeneracy conditions to hold identically. Continuous-time dynamics therefore conserve the learned energy and produce entropy to machine precision, with only the usual O(dt^2) drift from explicit time stepping, and the guarantees transfer zero-shot across a 4x super-resolution range. A gauge-invariant diagnostic, built from a fixed quadratic energy rather than the learned functionals, recovers which systems are reversible and which dissipate more. On four PDEs spanning reversible, dissipative, and mixed regimes the method stays competitive with unconstrained and energy-penalized baselines and improves accuracy on several dissipative and mixed problems at comparable or fewer parameters.

Core claim

GENERIC-FNO is the first neural operator to embed the full GENERIC structure in function space: energy and entropy functionals are learned as neural operators, while the Poisson and friction operators are diagonal Fourier multipliers sandwiched between rank-one projections that enforce both degeneracy conditions exactly by construction (residuals ~10^{-13}). Continuous-time dynamics therefore conserve the learned energy and produce entropy exactly, and those structural guarantees, not only accuracy, transfer zero-shot across a 4x super-resolution range.

What carries the argument

Projection-sandwiched diagonal Fourier multipliers: L = (I - P_S) D_L (I - P_S) and M = (I - P_E) D_M (I - P_E), where D_L is a skew multiplier ia(k) and D_M a positive semi-definite multiplier |b(k)|^2. The sandwiches remove the conjugate gradient directions so L δS/δu = 0 and M δE/δu = 0 hold identically, with no penalty or residual.

Load-bearing premise

The reversible and irreversible generators of the target flows can be well approximated by diagonal Fourier multipliers on a fixed low-frequency band after a rank-one projection sandwich; if richer structure is required, accuracy and dissipation ranking can suffer even while algebraic degeneracy still holds.

What would settle it

Train on a reversible linear-advection flow and check whether the gauge-invariant dissipation rate r_mech stays near zero while rollout accuracy remains competitive with an unconstrained operator of equal capacity; a large positive r_mech or a large accuracy collapse that vanishes only when the irreversible channel is unlocked would falsify the claim that the construction is both exact and useful.

Watch this falsifier — get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. GENERIC-FNO is presented as the first neural operator that embeds the full GENERIC (metriplectic) structure in function space: energy and entropy functionals Eθ, Sϕ are learned as neural operators, while the Poisson and friction operators L and M are diagonal Fourier multipliers sandwiched between rank-one L2 projections that remove the conjugate gradient directions. The construction enforces both degeneracy conditions L δS/δu = 0 and M δE/δu = 0 identically (no soft penalty or residual), so continuous-time dynamics conserve the learned energy and produce entropy exactly; explicit Euler adds only the standard O(Δt²) drift. The paper acknowledges the classical gauge non-uniqueness of (E,S,L,M) and introduces a gauge-invariant dissipation diagnostic r_mech based on the fixed quadratic energy Q = ½∥u∥². Experiments on three backbones (1D/2D FNO, DeepONet) and four PDEs (heat, advection, Burgers, damped wave) report machine-precision structural residuals, zero-shot super-resolution of both accuracy and guarantees (64→256), recovery of the ground-truth dissipation ordering, and competitive or better predictive accuracy on dissipative/mixed problems at comparable or fewer parameters.

Significance. If the claims hold, the work closes a genuine gap between finite-dimensional metriplectic networks and resolution-independent neural operators. The architectural contribution is load-bearing and cleanly proved: Appendix A establishes skew-adjointness/PSD preservation under the sandwich, exact degeneracy (Prop. 2), continuous-time first and second laws (Thm. 1), O(Δt²) discrete drift, gauge invariance of r_mech, and resolution independence on band-limited inputs (Prop. 5). Machine-precision residuals at random initialization, parameter-matched DeepONet ablations, long-horizon energy plots, and an honest treatment of gauge freedom and of the reversible-transport cost are strengths that go beyond typical soft-constraint papers. The result supplies a transferable inductive bias for high-dimensional field surrogates where leaving the physical manifold is costly (closures, complex fluids).

major comments (2)
  1. §6 and Table 6 correctly diagnose that explicit Euler on the skew operator produces coarse-grid amplitude growth on pure advection (aliasing of high-frequency harmonics from the E/S networks, undamped when M≈0). The authors retain Euler to keep discrete r_E ~10^{-7} rather than ~10^{-3} under RK4. This is a legitimate design choice, but the central claim of “exact structural guarantees” for deployed rollouts is then only semi-discrete; the manuscript should state more prominently (abstract/intro and §5.3) that the machine-precision energy residual is continuous-time, while discrete rollouts inherit the standard O(Δt²) integrator error, and that the reversible-transport accuracy cost is an integrator–expressivity interaction rather than a failure of degeneracy.
  2. §4.2–4.5 and Prop. 5: the operators are deliberately restricted to diagonal Fourier multipliers on a fixed low-frequency band. Algebraic degeneracy holds regardless of expressivity, so the structural claim is not threatened; however, the paper’s claim that the construction is “genuinely a function-space map” that recovers physical dissipation ordering rests on this restricted class being rich enough for the target flows. On pure transport the reversible channel is under-expressive (Table 1, Table 6). A short controlled experiment or discussion quantifying how much of the accuracy gap closes under multi-rank or block-banded structure-preserving multipliers (already flagged as future work) would strengthen the expressivity claim without altering the algebraic guarantees.
minor comments (5)
  1. Table 1 vs. Table 2: the DeepONet comparison is not parameter-controlled in the main table (175K vs 58K). Appendix C supplies the matched-budget ablation and should be referenced more visibly in the main text so readers do not misread capacity as structure.
  2. §5.2 damped-wave paragraph: the partial-observation collapse in 1D is important; a one-line note in the table caption or a footnote would prevent readers from wondering why wave is omitted from the headline tables.
  3. Figure 2 is described as log-scaled bars; ensure the published figure caption states the scale explicitly so the large relative gains on Burgers/heat are not misread.
  4. Notation: δE/δu is used both for the continuous variational derivative and for the autodiff Euclidean gradient (App. A.1). A brief reminder at first use in §4.1 would help non-specialist readers.
  5. Related work: the concurrent Baheri & Lindemann and Oprisa & Toth papers are carefully distinguished; a single sentence on how the present spectral, resolution-independent parameterization differs from node-local graph metriplectic models (Tierz et al.) could be tightened for clarity.

Circularity Check

0 steps flagged

No significant circularity: degeneracy and continuous-time laws are algebraic identities of the architecture, not fitted or self-cited predictions.

full rationale

GENERIC-FNO’s load-bearing claims are architectural and algebraic, not circular. The operators are defined as L=(I−P_S)D_L(I−P_S) and M=(I−P_E)D_M(I−P_E) so that L δS/δu=0 and M δE/δu=0 hold identically (Prop. 2); continuous-time energy conservation and entropy production then follow from skew-adjointness and positive semi-definiteness (Thm. 1). The paper states this as “by construction,” verifies residuals ~10^{-13} at random initialization, and does not present the identities as empirical discoveries. E_θ and S_ϕ receive no direct supervision—only the data-fitting loss through the dynamics—so they are not fitted inputs renamed as predictions. Gauge freedom is acknowledged (Öttinger 2005; §4.4), and thermodynamic claims are restricted to the gauge-invariant diagnostic r_mech (fixed quadratic energy Q=½∥u∥²) and structural residuals; per-channel attribution is explicitly demoted to appendix color. Citations for GENERIC and finite-dimensional metriplectic networks are external (Grmela & Öttinger, Lee et al., Zhang et al., Gruber et al.), not self-citations of Sulskis/Ravi. Empirical checks (zero-shot super-resolution of guarantees, recovery of ground-truth dissipation ordering, long-horizon energy boundedness, parameter-matched DeepONet ablations) are independent of the algebraic construction. Expressivity limits of diagonal Fourier multipliers are stated as limitations (§6), not smuggled uniqueness. The derivation chain is therefore self-contained: architecture ⇒ exact degeneracy ⇒ continuous-time laws; data tests accuracy and gauge-invariant diagnostics without forcing them by fit or self-reference.

Axiom & Free-Parameter Ledger

4 free parameters · 6 axioms · 2 invented entities

The structural guarantees rest on standard GENERIC algebra plus the projection sandwich; they do not depend on fitted physical constants. Free parameters are architectural and training choices that affect accuracy, not the algebraic identities. Invented pieces are the architecture and the diagnostic; both have falsifiable handles (machine-precision residuals, ordering of r_mech vs Π⋆).

free parameters (4)
  • Fourier multiplier band size m
    Number of low-frequency modes on which a(k) and b(k) act; chosen by hand and fixed across resolutions; affects expressivity of L and M.
  • Neural-operator backbone widths/depths and MLP heads for Eθ, Sϕ
    Architecture hyperparameters that determine capacity of the learned functionals; fitted via training loss, not derived.
  • Training horizon, learning rate, epoch count, relative L2 loss weights
    Optimization choices that shape the learned flow; structure holds independently but accuracy and long-horizon behavior depend on them.
  • Explicit Euler step size Δt
    Controls O(Δt²) discrete energy drift; retained for exact discrete degeneracy inheritance rather than optimized for accuracy.
axioms (6)
  • domain assumption GENERIC dynamics ∂tu = L δE/δu + M δS/δu with L skew-adjoint, M self-adjoint PSD, and mutual degeneracy LδS/δu=0, MδE/δu=0 imply continuous-time energy conservation and entropy production.
    Taken from Grmela & Öttinger / Öttinger; restated in §3 and proved for the construction in Appendix A.6.
  • domain assumption Fields live on a periodic torus; functionals are Fréchet differentiable; variational derivatives are obtained by reverse-mode autodiff through mean-pooled neural operators.
    Function-space setting of §3 and §4.1; required for L2 projections and resolution independence.
  • standard math Diagonal Fourier multipliers with purely imaginary odd symbol (L) and non-negative even symbol (M), after real-FFT Hermitian enforcement, are skew-adjoint and PSD respectively for any parameter values.
    Proposition 1, Appendix A.3; standard spectral fact used to avoid dense dimension-scaling matrices.
  • standard math Rank-one L2 projections are self-adjoint idempotents; sandwiching preserves skew-adjointness / PSD (Lemma 2).
    Appendix A.2–A.4; the algebraic engine of exact degeneracy.
  • domain assumption For the scalar PDEs studied, the fixed quadratic Q[u]=½∥u∥² is the relevant mechanical energy whose decay measures physical dissipation.
    Stated in §4.4 and §5.4; r_mech is defined with this Q; authors note redefinition would be needed if Q is not the invariant.
  • ad hoc to paper Band-limited initial data and fixed maximum frequency make continuous dynamics identical across grids used in zero-shot super-resolution.
    §5.3 experimental design; required for the claim that structure (not only accuracy) transfers.
invented entities (2)
  • GENERIC-FNO architecture (learned E/S neural operators + projection-sandwiched diagonal Fourier L/M) independent evidence
    purpose: Realize full metriplectic dynamics as a resolution-independent neural operator with exact degeneracy by construction.
    Core contribution; independent evidence via machine-precision residual tables and zero-shot residual transfer.
  • Gauge-invariant dissipation diagnostic r_mech = −⟨u,∂tu⟩/(∥u∥∥∂tu∥) using fixed Q=½∥u∥² independent evidence
    purpose: Separate reversible from dissipative dynamics without reference to the non-unique learned (E,S,L,M).
    Introduced in §4.4; falsifiable by comparison to ground-truth Π⋆ ordering and vanishing on advection.

pith-pipeline@v1.1.0-grok45 · 26671 in / 3838 out tokens · 33085 ms · 2026-07-12T14:43:40.269963+00:00 · methodology

0 comments
read the original abstract

We introduce GENERIC-FNO, the first neural operator to embed the full GENERIC (metriplectic) structure of nonequilibrium thermodynamics -- reversible, energy-conserving dynamics and irreversible, entropy-producing dynamics coupled through the degeneracy conditions -- directly in function space. Existing structure-preserving neural operators enforce at most a single conservation law or reversible (Hamiltonian) structure, while thermodynamically consistent learning has been confined to finite-dimensional, graph, or particle systems. GENERIC-FNO closes this gap: it learns the energy and entropy functionals as neural operators and parameterizes the Poisson and friction operators as diagonal Fourier multipliers sandwiched between rank-one projections that enforce the degeneracy conditions exactly, by construction, with no penalty term, update projection, or residual. The degeneracy identities hold to machine precision (residuals ~10^-13) for any initialization, dimension, or resolution, so the continuous-time dynamics conserve the learned energy and produce entropy exactly; the explicit time stepping adds only a small O(dt^2) drift (per-step residual ~10^-6). We further note that the (E,S,L,M) decomposition of a given flow is not unique, and introduce a gauge-invariant dissipation diagnostic separating reversible from dissipative dynamics independently of the learned functionals. Across three operator backbones (1D/2D FNOs and DeepONet) and four PDEs spanning reversible, dissipative, and mixed regimes, GENERIC-FNO preserves its exact structural guarantees zero-shot across a 4x super-resolution range (64 to 256), recovers the ground-truth ordering of physical dissipation, and is competitive with strong unconstrained and energy-penalized baselines, outperforming them on several dissipative and mixed problems at comparable or fewer parameters.

Figures

Figures reproduced from arXiv: 2606.08343 by Jason Sulskis, Sathya Ravi.

Figure 1
Figure 1. Figure 1: GENERIC-FNO forward map. Two neural operators produce scalar functionals Eθ[u] and Sϕ[u], whose variational derivatives are taken by autodiff. Each diagonal Fourier multiplier—skew i a(k) for the reversible operator L, positive semi-definite |b(k)| 2 for the irreversible operator M—is sandwiched between the rank-one projection that removes the conjugate gradient (dashed: δS/δu feeds the L branch, δE/δu fee… view at source ↗
Figure 2
Figure 2. Figure 2: Rollout L 2 error with three-seed error bars, one panel per backbone (2D FNO, 1D FNO, 1D DeepONet), wave excluded. GENERIC (red) versus unconstrained and energy-penalized baselines. Visual companion to Tables 1–2; bars are log-scaled. On the 2D FNO backbone ( [PITH_FULL_IMAGE:figures/full_fig_p010_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Zero-shot super-resolution (trained at 64). Top: rollout L 2 versus evaluation resolution for FNO and GENERIC-FNO on heat and Burgers (flat-to-improving across the 4× range). Bottom: the structural residual rE and energy drift |dE/step| of GENERIC-FNO versus resolution (log scale), confirming the guarantees, not only the accuracy, transfer zero-shot. Dotted line marks the training grid. 5.4 Gauge-Invariant… view at source ↗
Figure 4
Figure 4. Figure 4: Gauge-invariant dissipation rmech (three-seed mean ± std) per PDE, grouped by backbone. Ad￾vection sits at ≈ 0 (the model adds no spurious dissipation to reversible dynamics) while heat and Burgers are positive; the across-PDE ordering matches the ground-truth Π⋆ in every backbone. Computed from a fixed quadratic energy, so it is independent of the learned (gauge-dependent) E, S. 5.5 Long-Horizon Stability… view at source ↗
Figure 5
Figure 5. Figure 5: Long-horizon rollouts (200 steps; dotted line marks the 15-step training horizon). Left: relative error vs. step (log)—the unconstrained and penalized baselines diverge on heat and Burgers while GENERIC￾FNO stays bounded. Right: physical energy Q = 1 2 ∥u∥ 2 relative to its initial value, against the ground truth (dashed). FNO’s energy blows up on heat (≈ 9×) and Burgers (≈ 5×); GENERIC-FNO stays bounded o… view at source ↗
Figure 6
Figure 6. Figure 6: The structural identities hold to machine precision. Magnitude (log scale) of the GENERIC residuals at random initialization on a 16 × 16 grid: reversible skewness ⟨δE, L δE⟩, the two degeneracy conditions ⟨δS, L δE⟩ and ⟨δE, M δS⟩, energy conservation ⟨δE, ∂tu⟩, and the trained-model energy residual rE. All sit at or near the float64 machine epsilon (dashed line), confirming that degeneracy is enforced by… view at source ↗
Figure 7
Figure 7. Figure 7: Gauge freedom, made visual. Per-seed values of the gauge-dependent dissipative fraction ρM (left) and the gauge-invariant dissipation rate rmech (right), for the non-reversible PDEs across all three backbones. The channel attribution ρM scatters wildly across seeds—for 2D-FNO Burgers it spans the full range from 0.002 to 1.000, i.e. the model represents the same flow once as almost purely reversible and on… view at source ↗
Figure 8
Figure 8. Figure 8: rmech recovers the ground-truth dissipation ordering. Each point is one (backbone, PDE) pair: the model’s gauge-invariant rate rmech against the ground-truth rate Π⋆ (both log scale; markers denote PDE, colors denote backbone). The reversible advection points cluster at the bottom-left (≈ 0) and the dissipative/mixed points lie up-and-to-the-right. Within each backbone the three PDEs are ranked exactly as … view at source ↗
Figure 9
Figure 9. Figure 9: Why we keep explicit Euler. Advection trained at 64 and evaluated to 256. Left: rollout L 2 for the FNO baseline, GENERIC with explicit Euler, and GENERIC with RK4 + anti-aliasing. RK4 roughly halves the coarse-grid blow-up (Euler 0.61 → RK4 0.34 at the train grid) and the two integrators agree at finer grids. Right: the per-step energy residual rE, which RK4 raises from ∼ 10−7 (Euler, exact discrete degen… view at source ↗
Figure 10
Figure 10. Figure 10: Learned thermodynamics along a rollout (illustrative, single realization). For one trained model per PDE we plot the learned energy Eθ[ut], the learned entropy Sϕ[ut], and the fixed me￾chanical energy Q = 1 2 ∥ut∥ 2 along a trajectory, each affinely rescaled to [0, 1] within its panel for display. In this realization Eθ stays flat (conserved) while Sϕ rises monotonically and tracks the decay of Q for the … view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

4 extracted references · 4 linked inside Pith

  1. [1]

    Miles Cranmer, Sam Greydanus, Stephan Hoyer, Peter Battaglia, David Spergel, and Shirley Ho

    arXiv:2509.19526. Miles Cranmer, Sam Greydanus, Stephan Hoyer, Peter Battaglia, David Spergel, and Shirley Ho. Lagrangian neuralnetworks. InICLR 2020 Workshop on Integration of Deep Neural Models and Differential Equations,

  2. [2]

    Karthik Duraisamy, Gianluca Iaccarino, and Heng Xiao

    arXiv:2003.04630. Karthik Duraisamy, Gianluca Iaccarino, and Heng Xiao. Turbulence modeling in the age of data.Annual Review of Fluid Mechanics, 51:357–377, 2019. doi: 10.1146/annurev-fluid-010518-040547. Sam Greydanus, Misko Dzamba, and Jason Yosinski. Hamiltonian neural networks. InAdvances in Neural Information Processing Systems (NeurIPS), volume 32, ...

  3. [3]

    Anthony Gruber, Kookjin Lee, Haksoo Lim, Noseong Park, and Nathaniel Trask

    arXiv:2305.15616. Anthony Gruber, Kookjin Lee, Haksoo Lim, Noseong Park, and Nathaniel Trask. Efficiently parameterized neural metriplectic systems. InInternational Conference on Learning Representations (ICLR), 2025. Quercus Hernández, Alberto Badías, David González, Francisco Chinesta, and Elías Cueto. Structure- preserving neural networks.Journal of Co...

  4. [4]

    Zongyi Li, Nikola Kovachki, Kamyar Azizzadenesheli, Burigede Liu, Kaushik Bhattacharya, Andrew Stu- art, and Anima Anandkumar

    arXiv:2106.12619. Zongyi Li, Nikola Kovachki, Kamyar Azizzadenesheli, Burigede Liu, Kaushik Bhattacharya, Andrew Stu- art, and Anima Anandkumar. Fourier neural operator for parametric partial differential equations. In International Conference on Learning Representations (ICLR), 2021. arXiv:2010.08895. Lu Lu, Pengzhan Jin, Guofei Pang, Zhongqiang Zhang, a...