Pith. sign in

REVIEW 2 cited by

Scaling physics-informed hard constraints with mixture-of-experts

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.13412 v1 pith:CA3NPZCD submitted 2024-02-20 cs.LG cs.AIcs.NAmath.NAmath.OC

classification cs.LGcs.AIcs.NAmath.NAmath.OC
keywords constraintsoptimizationneuralphysicaltrainingapproachdifferentiableduring
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Imposing known physical constraints, such as conservation laws, during neural network training introduces an inductive bias that can improve accuracy, reliability, convergence, and data efficiency for modeling physical dynamics. While such constraints can be softly imposed via loss function penalties, recent advancements in differentiable physics and optimization improve performance by incorporating PDE-constrained optimization as individual layers in neural networks. This enables a stricter adherence to physical constraints. However, imposing hard constraints significantly increases computational and memory costs, especially for complex dynamical systems. This is because it requires solving an optimization problem over a large number of points in a mesh, representing spatial and temporal discretizations, which greatly increases the complexity of the constraint. To address this challenge, we develop a scalable approach to enforce hard physical constraints using Mixture-of-Experts (MoE), which can be used with any neural network architecture. Our approach imposes the constraint over smaller decomposed domains, each of which is solved by an "expert" through differentiable optimization. During training, each expert independently performs a localized backpropagation step by leveraging the implicit function theorem; the independence of each expert allows for parallelization across multiple GPUs. Compared to standard differentiable optimization, our scalable approach achieves greater accuracy in the neural PDE solver setting for predicting the dynamics of challenging non-linear systems. We also improve training stability and require significantly less computation time during both training and inference stages.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Spatio-temporal, multi-field deep learning of shock propagation in meso-structured media

    cs.LG 2025-09 conditional novelty 6.0 of 10

    A CNN-LSTM surrogate autoregressively predicts seven coupled shock fields in meso-structured materials with 1.4-3.2% RMSE, 94% better than single-field models.

  2. GITO: Graph-Informed Transformer Operator for Learning Complex Partial Differential Equations

    cs.LG 2025-06 conditional novelty 5.0 of 10

    GITO, a graph-informed transformer operator, reports lower relative L2 errors than existing transformer-based neural operators on Navier-Stokes, heat conduction, and airfoil benchmark datasets.

Pith tools