Pith. sign in

REVIEW 4 cited by

Tensor Programs III: Neural Matrix Laws

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2009.10685 v3 pith:2LBHIH3A submitted 2020-09-22 cs.NE math.PR

classification cs.NEmath.PR
keywords neuralactivationsmasterasymptoticindependencelawsmatrixnetwork
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

In a neural network (NN), *weight matrices* linearly transform inputs into *preactivations* that are then transformed nonlinearly into *activations*. A typical NN interleaves multitudes of such linear and nonlinear transforms to express complex functions. Thus, the (pre-)activations depend on the weights in an intricate manner. We show that, surprisingly, (pre-)activations of a randomly initialized NN become *independent* from the weights as the NN's widths tend to infinity, in the sense of asymptotic freeness in random matrix theory. We call this the Free Independence Principle (FIP), which has these consequences: 1) It rigorously justifies the calculation of asymptotic Jacobian singular value distribution of an NN in Pennington et al. [36,37], essential for training ultra-deep NNs [48]. 2) It gives a new justification of gradient independence assumption used for calculating the Neural Tangent Kernel of a neural network. FIP and these results hold for any neural architecture. We show FIP by proving a Master Theorem for any Tensor Program, as introduced in Yang [50,51], generalizing the Master Theorems proved in those works. As warmup demonstrations of this new Master Theorem, we give new proofs of the semicircle and Marchenko-Pastur laws, which benchmarks our framework against these fundamental mathematical results.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Precise gradient descent training dynamics for finite-width multi-layer neural networks

    cs.LG 2025-05 conditional novelty 8.0 of 10

    Gradient descent iterates of finite-width multi-layer networks on single-index data obey a state evolution law, giving exact training/test error formulas and a data-driven test error estimator.

  2. Correlation flow governs learning at criticality

    cs.LG 2026-08 conditional novelty 7.0 of 10

    At critical initialization, the infinite-depth neural tangent kernel converges to the fixed-point output correlation matrix divided by an activation-dependent constant, making learning dynamics equivalent to correlati...

  3. SETOL: A Semi-Empirical Theory of (Deep) Learning

    cs.LG 2025-07 conditional novelty 7.0 of 10

    SETOL derives the HTSR layer quality metrics as integrated R-transforms of the layer spectral density, and proposes a determinant condition (ERG) as a marker of ideal learning.

  4. Random weights of DNNs and emergence of fixed points

    cs.LG 2025-01 reject novelty 6.0 of 10

    Heavy-tailed random weights in square feedforward DNNs produce multiple stable fixed point attractors, while Gaussian weights yield a single fixed point, with a non-monotone dependence on depth.

Pith tools