Gradient flow on a two-layer network trained to compose finite-group elements provably pushes each neuron to a single irreducible representation with rank-one cross-layer alignment; for Abelian groups it yields a uniformly diversified, Haar-phase majority-vote predictor.
There Will Be a Scientific Theory of Deep Learning
11 Pith papers cite this work. Polarity classification is still indexing.
abstract
In this paper, we make the case that a scientific theory of deep learning is emerging. By this we mean a theory which characterizes important properties and statistics of the training process, hidden representations, final weights, and performance of neural networks. We pull together major strands of ongoing research in deep learning theory and identify five growing bodies of work that point toward such a theory: (a) solvable idealized settings that provide intuition for learning dynamics in realistic systems; (b) tractable limits that reveal insights into fundamental learning phenomena; (c) simple mathematical laws that capture important macroscopic observables; (d) theories of hyperparameters that disentangle them from the rest of the training process, leaving simpler systems behind; and (e) universal behaviors shared across systems and settings which clarify which phenomena call for explanation. Taken together, these bodies of work share certain broad traits: they are concerned with the dynamics of the training process; they primarily seek to describe coarse aggregate statistics; and they emphasize falsifiable quantitative predictions. We argue that the emerging theory is best thought of as a mechanics of the learning process, and suggest the name learning mechanics. We discuss the relationship between this mechanics perspective and other approaches for building a theory of deep learning, including the statistical and information-theoretic perspectives. In particular, we anticipate a symbiotic relationship between learning mechanics and mechanistic interpretability. We also review and address common arguments that fundamental theory will not be possible or is not important. We conclude with a portrait of important open directions in learning mechanics and advice for beginners. We host further introductory materials, perspectives, and open questions at learningmechanics.pub.
citation-role summary
citation-polarity summary
years
2026 11roles
background 1polarities
background 1representative citing papers
Under a time-invariant PSD training operator, gradient-flow iterates equal BLUPs in a random-effects model, so REML of the training-time variance component yields asymptotically prediction-optimal early stopping.
BCJR-QAT makes trellis quantization differentiable via BCJR soft decoding at finite temperature, allowing QAT to improve 2-bit LLM perplexity over PTQ with a fused GPU kernel and a drift-budget escape condition.
A coarse-grained Levy-jump model predicts that recurrent networks settle into either a collapsed (exponential-forgetting) or anti-collapsed (power-law-forgetting) regime, with one spectral exponent beta governing both the time-scale spectrum and the envelope.
Critical percolation clusters embedded in high dimensions, combined with taxonomic latent variables, form an analytically tractable synthetic data model whose ground-truth hierarchy can be linearly decoded from network activations.
PMNet uses unitary phasor dynamics and hierarchical anchors to make explicit memory stable for long sequences, matching a 3x larger Mamba model on long-context robustness with a 119M parameter network.
Perturbation probing identifies tiny sets of FFN neurons that control refusal templates and language routing in LLMs, enabling precise ablations and directional interventions that alter behavior on benchmarks while preserving safety.
Deep feedforward ReLU networks generalize two-layer principles via paths, piecewise linear manifolds, and continuity restriction to explain training solutions.
In spiking ResNets, 1FC ensembles defined by pairwise correlations show ReLU-like cofiring-to-response mapping whose gain scales with ensemble size, with reliable class encoding restricted to infrequent high-cofiring events.
Derives an asymptotic equivalent for the Representation Gap in equivariant diffusion models, showing it depends primarily on the intrinsic dimension of the task.
Formalizes emergent intelligence in foundation models as the limit of E(N,P,K) as N,P,K approach infinity, proves existence conditions via nonlinear Lipschitz operators, and derives scaling laws from covering numbers.
citing papers explorer
-
Neural Networks Provably Learn Spectral Representations for Group Composition
Gradient flow on a two-layer network trained to compose finite-group elements provably pushes each neuron to a single irreducible representation with rank-one cross-layer alignment; for Abelian groups it yields a uniformly diversified, Haar-phase majority-vote predictor.
-
Gradient-Flow Optimization as Dynamic Random-Effects Inference: Testing and Early Stopping with Applications to Deep Learning
Under a time-invariant PSD training operator, gradient-flow iterates equal BLUPs in a random-effects model, so REML of the training-time variance component yields asymptotically prediction-optimal early stopping.
-
BCJR-QAT: A Differentiable Relaxation of Trellis-Coded Weight Quantization
BCJR-QAT makes trellis quantization differentiable via BCJR soft decoding at finite temperature, allowing QAT to improve 2-bit LLM perplexity over PTQ with a fused GPU kernel and a drift-budget escape condition.
-
Anti-Collapse Dynamics and the Emergence of Multi-Time-Scale Learning in Recurrent Neural Networks
A coarse-grained Levy-jump model predicts that recurrent networks settle into either a collapsed (exponential-forgetting) or anti-collapsed (power-law-forgetting) regime, with one spectral exponent beta governing both the time-scale spectrum and the envelope.
-
Critical Percolation as a Synthetic Data Model for Interpretability
Critical percolation clusters embedded in high dimensions, combined with taxonomic latent variables, form an analytically tractable synthetic data model whose ground-truth hierarchy can be linearly decoded from network activations.
-
Phasor Memory Networks: Stable Backpropagation Through Time for Scalable Explicit Memory
PMNet uses unitary phasor dynamics and hierarchical anchors to make explicit memory stable for long sequences, matching a 3x larger Mamba model on long-context robustness with a 119M parameter network.
-
Perturbation Probing: A Two-Pass-per-Prompt Diagnostic for FFN Behavioral Circuits in Aligned LLMs
Perturbation probing identifies tiny sets of FFN neurons that control refusal templates and language routing in LLMs, enabling precise ablations and directional interventions that alter behavior on benchmarks while preserving safety.
-
On the Principles of Deep Feedforward ReLU Networks
Deep feedforward ReLU networks generalize two-layer principles via paths, piecewise linear manifolds, and continuity restriction to explain training solutions.
-
Rare Events, Real Signals: Functional Ensembles as Units of Computation in Deep Spiking Networks
In spiking ResNets, 1FC ensembles defined by pairwise correlations show ReLU-like cofiring-to-response mapping whose gain scales with ensemble size, with reliable class encoding restricted to infrequent high-cofiring events.
-
Representation Gap: Explaining the Unreasonable Effectiveness of Neural Networks from a Geometric Perspective
Derives an asymptotic equivalent for the Representation Gap in equivariant diffusion models, showing it depends primarily on the intrinsic dimension of the task.
-
A Limit Theory of Foundation Models: A Mathematical Approach to Understanding Emergent Intelligence and Scaling Laws
Formalizes emergent intelligence in foundation models as the limit of E(N,P,K) as N,P,K approach infinity, proves existence conditions via nonlinear Lipschitz operators, and derives scaling laws from covering numbers.