Pith. sign in

REVIEW 11 cited by

Does equivariance matter at scale?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.23179 v2 pith:7GWX2CYT submitted 2024-10-30 cs.LG

Does equivariance matter at scale?

classification cs.LG
keywords computenon-equivarianttrainingdataequivariantmodelssizearchitectures
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Given large datasets and sufficient compute, is it beneficial to design neural architectures for the structure and symmetries of each problem? Or is it more efficient to learn them from data? We study empirically how equivariant and non-equivariant networks scale with compute and training samples. Focusing on a benchmark problem of rigid-body interactions and on general-purpose transformer architectures, we perform a series of experiments, varying the model size, training steps, and dataset size. We find evidence for three conclusions. First, equivariance improves data efficiency, but training non-equivariant models with data augmentation can close this gap given sufficient epochs. Second, scaling with compute follows a power law, with equivariant models outperforming non-equivariant ones at each tested compute budget. Finally, the optimal allocation of a compute budget onto model size and training duration differs between equivariant and non-equivariant models.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Characterizing Learning in Deep Neural Networks using Tractable Algorithmic Complexity Analysis

    cs.LG 2026-05 unverdicted novelty 7.0

    QuBD extends algorithmic complexity estimation to quantized DNN weights, revealing that complexity decreases during learning, increases with overfitting, follows grokking patterns, and correlates with generalization.

  2. Discretizing Group-Convolutional Neural Networks for 3D Geometry in Feature Space

    cs.CV 2026-05 unverdicted novelty 7.0

    Feature-space sampling in GCNNs preserves 3D classification accuracy with coarse discretization, enabling precomputation and faster training of equivariant models.

  3. Conformal Orbit-Valid Trust Horizons for Equivariant World Models

    cs.LG 2026-06 unverdicted novelty 6.0

    Conformal calibration produces orbit-valid trust horizons for equivariant world models, with zero violations in 50 audits and non-vacuous certificates on 2D/3D substrates.

  4. Equivariance and Augmentation for Bayesian Neural Networks

    cs.LG 2026-06 unverdicted novelty 5.0

    Derives exact equivariance conditions for augmented BNNs under variational inference and proposes orbit expansion symmetrization that outperforms baselines on equivariance and accuracy.

  5. How the Optimizer Shapes Learned Solutions in Equivariant Neural Networks

    cs.LG 2026-05 unverdicted novelty 5.0

    Muon optimizer improves performance over Adam in equivariant networks on ModelNet40 and produces solutions with larger Hessian curvature, more regular loss surfaces, and higher stable/effective ranks.

  6. Symmetry in the Wild: The Role of Equivariance in Neural Fluid Surrogates

    cs.LG 2026-05 unverdicted novelty 5.0

    Explicit E(3)-equivariance in neural CFD surrogates improves generalization on diverse-geometry hemodynamics benchmarks but degrades in-distribution performance on strongly aligned aerodynamics data, consistently beat...

  7. Physics Priors Offer Useful Accuracy-Carbon Trade-Offs in Spatio-Temporal Forecasting

    cs.LG 2025-09 unverdicted novelty 5.0

    Stronger physics priors in neural networks for spatio-temporal shear flow forecasting yield substantially lower training carbon footprints than weak or no priors, though inference savings are less consistent.

  8. Machine Learning Interatomic Potentials: Advancing Open-Source Software for Efficient and Scalable Molecular Simulation

    physics.chem-ph 2026-05 unverdicted novelty 4.0

    mlip v2 is a new software release that integrates API redesign, e3j backend, eSEN model, improved charge modeling, and expanded simulation capabilities to support larger-scale molecular modeling.

  9. Statistical Properties of Training & Generalization

    stat.ML 2026-06 unverdicted novelty 2.0

    Neural scaling laws in deep learning interact with physics constraints and inductive biases beyond classical statistics.

  10. Six Open Questions in Machine-Learned Interatomic Potential Foundation Models

    cond-mat.mtrl-sci 2026-06 unverdicted novelty 2.0

    This perspective article develops a definition of foundational MLIPs and poses six open questions that the authors believe will define future research in machine-learned interatomic potentials.

  11. Statistical Properties of Training & Generalization

    stat.ML 2026-06 unverdicted novelty 1.0

    Review of neural scaling laws and their relation to constraints and inductive biases when applying machine learning to physics problems.