Pith. sign in

REVIEW 11 cited by

The Principles of Deep Learning Theory

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2106.10165 v2 pith:HQEJJINJ submitted 2021-06-18 cs.LG cs.AIhep-thstat.ML

classification cs.LGcs.AIhep-thstat.ML
keywords networkslearningexplainflownetworkratioaspectdeep
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This book develops an effective theory approach to understanding deep neural networks of practical relevance. Beginning from a first-principles component-level picture of networks, we explain how to determine an accurate description of the output of trained networks by solving layer-to-layer iteration equations and nonlinear learning dynamics. A main result is that the predictions of networks are described by nearly-Gaussian distributions, with the depth-to-width aspect ratio of the network controlling the deviations from the infinite-width Gaussian description. We explain how these effectively-deep networks learn nontrivial representations from training and more broadly analyze the mechanism of representation learning for nonlinear models. From a nearly-kernel-methods perspective, we find that the dependence of such models' predictions on the underlying learning algorithm can be expressed in a simple and universal way. To obtain these results, we develop the notion of representation group flow (RG flow) to characterize the propagation of signals through the network. By tuning networks to criticality, we give a practical solution to the exploding and vanishing gradient problem. We further explain how RG flow leads to near-universal behavior and lets us categorize networks built from different activation functions into universality classes. Altogether, we show that the depth-to-width ratio governs the effective model complexity of the ensemble of trained networks. By using information-theoretic techniques, we estimate the optimal aspect ratio at which we expect the network to be practically most useful and show how residual connections can be used to push this scale to arbitrary depths. With these tools, we can learn in detail about the inductive bias of architectures, hyperparameters, and optimizers.

Discussion (0). Sign in to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Criticality and Saturation in Orthogonal Neural Networks

    cs.LG 2026-05 conditional novelty 7.0 of 10

    Derives layer-wise recursions for finite-width tensors under orthogonal initialization that reproduce the observed large-depth stability of nonlinear networks.

  2. Statistics of correlations in nonlinear recurrent neural networks

    q-bio.NC 2025-10 conditional novelty 7.0 of 10

    A replica path-integral calculation yields analytic large-N formulas for covariance statistics and participation dimension in nonlinear recurrent neural networks with quenched Gaussian disorder, confirmed by simulation.

  3. Interpreting learning dynamics of autoencoders: Transient scaling and emerging concepts of the Ising model

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Autoencoders trained on Ising spin configurations learn large-scale magnetization before small-scale energy features; deep models often arrest before the energy stage, and recursive self-application reveals stable lat...

  4. Interpreting learning dynamics of autoencoders: Transient scaling and emerging concepts of the Ising model

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Unsupervised autoencoders on Ising configurations form magnetization then energy representations in two dynamical regimes, with recursive error flow fields sharing topology across layers.

  5. Statistics of correlations in nonlinear recurrent neural networks

    q-bio.NC 2025-10 unverdicted novelty 6.0 of 10

    Derives exact correlation statistics for nonlinear RNNs in the large-N limit with Gaussian quenched disorder using path integrals, generalizing linear results and adding 1/N corrections.

  6. Optimized Weight Initialization on the Stiefel Manifold for Deep ReLU Neural Networks

    cs.LG 2025-08 conditional novelty 6.0 of 10

    An initialization that maximizes a semi-orthogonal weight matrix's alignment with the all-ones vector prevents dying ReLU and keeps 100-layer ReLU networks trainable.

  7. Pre-Strings Lectures on Artificial Intelligence

    hep-th 2026-07 accept novelty 5.5 of 10

    Lecture notes define neural-network field theory and survey how it recovers known QFT/string results plus applied AI techniques for string problems.

  8. Man, Machine, and Mathematics

    math.OC 2026-04 unverdicted novelty 5.0 of 10

    A high-level outline is given for a unified theory that reduces learning to a small set of ideas from dynamical systems, geometry, and physics via definitions of solvable problems and parametrized methods.

  9. Viability of perturbative expansion for quantum field theories on neurons

    hep-th 2025-08 unverdicted novelty 5.0 of 10

    The work tests perturbative viability of single-layer neural networks for local QFTs at finite neuron number N in phi^4 theory, finding UV-cutoff-sensitive O(1/N) corrections with weak convergence and proposing a modi...

  10. Integrating Out, Twice:The Open-System Case That Neural-Network Ensemble Theory Is Missing

    cs.LG 2026-06 unverdicted novelty 4.0 of 10

    Neural-network ensembles match closed Gaussian systems but lack the open-system non-Hermitian generator and continuous spectrum required by nuclear optical models, yielding a structural negative on applicability.

  11. Lectures on Semiclassical Methods for Composite Operators

    hep-th 2026-06 unverdicted novelty 3.0 of 10

    Lecture notes develop semiclassical methods to compute large-n scaling dimensions of composite operators in CFTs, recovering known results in free theory and deriving one-loop corrections at the Wilson-Fisher fixed point.

Pith tools