REVIEW 11 cited by
Does equivariance matter at scale?
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Does equivariance matter at scale?
read the original abstract
Given large datasets and sufficient compute, is it beneficial to design neural architectures for the structure and symmetries of each problem? Or is it more efficient to learn them from data? We study empirically how equivariant and non-equivariant networks scale with compute and training samples. Focusing on a benchmark problem of rigid-body interactions and on general-purpose transformer architectures, we perform a series of experiments, varying the model size, training steps, and dataset size. We find evidence for three conclusions. First, equivariance improves data efficiency, but training non-equivariant models with data augmentation can close this gap given sufficient epochs. Second, scaling with compute follows a power law, with equivariant models outperforming non-equivariant ones at each tested compute budget. Finally, the optimal allocation of a compute budget onto model size and training duration differs between equivariant and non-equivariant models.
Forward citations
Cited by 11 Pith papers
-
Characterizing Learning in Deep Neural Networks using Tractable Algorithmic Complexity Analysis
QuBD extends algorithmic complexity estimation to quantized DNN weights, revealing that complexity decreases during learning, increases with overfitting, follows grokking patterns, and correlates with generalization.
-
Discretizing Group-Convolutional Neural Networks for 3D Geometry in Feature Space
Feature-space sampling in GCNNs preserves 3D classification accuracy with coarse discretization, enabling precomputation and faster training of equivariant models.
-
Conformal Orbit-Valid Trust Horizons for Equivariant World Models
Conformal calibration produces orbit-valid trust horizons for equivariant world models, with zero violations in 50 audits and non-vacuous certificates on 2D/3D substrates.
-
Equivariance and Augmentation for Bayesian Neural Networks
Derives exact equivariance conditions for augmented BNNs under variational inference and proposes orbit expansion symmetrization that outperforms baselines on equivariance and accuracy.
-
How the Optimizer Shapes Learned Solutions in Equivariant Neural Networks
Muon optimizer improves performance over Adam in equivariant networks on ModelNet40 and produces solutions with larger Hessian curvature, more regular loss surfaces, and higher stable/effective ranks.
-
Symmetry in the Wild: The Role of Equivariance in Neural Fluid Surrogates
Explicit E(3)-equivariance in neural CFD surrogates improves generalization on diverse-geometry hemodynamics benchmarks but degrades in-distribution performance on strongly aligned aerodynamics data, consistently beat...
-
Physics Priors Offer Useful Accuracy-Carbon Trade-Offs in Spatio-Temporal Forecasting
Stronger physics priors in neural networks for spatio-temporal shear flow forecasting yield substantially lower training carbon footprints than weak or no priors, though inference savings are less consistent.
-
Machine Learning Interatomic Potentials: Advancing Open-Source Software for Efficient and Scalable Molecular Simulation
mlip v2 is a new software release that integrates API redesign, e3j backend, eSEN model, improved charge modeling, and expanded simulation capabilities to support larger-scale molecular modeling.
-
Statistical Properties of Training & Generalization
Neural scaling laws in deep learning interact with physics constraints and inductive biases beyond classical statistics.
-
Six Open Questions in Machine-Learned Interatomic Potential Foundation Models
This perspective article develops a definition of foundational MLIPs and poses six open questions that the authors believe will define future research in machine-learned interatomic potentials.
-
Statistical Properties of Training & Generalization
Review of neural scaling laws and their relation to constraints and inductive biases when applying machine learning to physics problems.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.