Pith. sign in

REVIEW 4 cited by

Quantitative CLTs in Deep Neural Networks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.06092 v5 pith:PT25LPFS submitted 2023-07-12 cs.LG cs.AImath.PRstat.ML

classification cs.LGcs.AImath.PRstat.ML
keywords networkboundsconnectedfullygammagaussianlargeneural
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

We study the distribution of a fully connected neural network with random Gaussian weights and biases in which the hidden layer widths are proportional to a large constant $n$. Under mild assumptions on the non-linearity, we obtain quantitative bounds on normal approximations valid at large but finite $n$ and any fixed network depth. Our theorems show both for the finite-dimensional distributions and the entire process, that the distance between a random fully connected network (and its derivatives) to the corresponding infinite width Gaussian process scales like $n^{-\gamma}$ for $\gamma>0$, with the exponent depending on the metric used to measure discrepancy. Our bounds are strictly stronger in terms of their dependence on network width than any previously available in the literature; in the one-dimensional case, we also prove that they are optimal, i.e., we establish matching lower bounds.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Geometric Dyson Brownian Motions and the Free Log-Normal Limit for a Non-Square Gaussian Matrix Product

    math.PR 2026-06 unverdicted novelty 7.0 of 10

    In double asymptotic limits, the squared singular value process of non-square matrix products obeys geometric Dyson Brownian motion whose T-transform solves a Burgers equation, producing the free log-normal law via fr...

  2. Proportional infinite-width infinite-depth limit for deep linear neural networks

    stat.ML 2024-11 accept novelty 7.0 of 10

    Deep linear neural networks in the proportional depth-width limit converge to a nontrivial mixture of Gaussians, with posterior output correlations that depend on the observed labels.

  3. Effective Non-Random Extreme Learning Machine

    stat.ML 2024-11 conditional novelty 6.0 of 10

    ENR-ELM constructs deterministic hidden-layer features from the eigenvectors of the NNGP kernel and fits the output layer by projection or incremental regression, matching ELM accuracy with less randomness and faster ...

  4. Kernel shape renormalization explains output-output correlations in finite Bayesian one-hidden-layer networks

    cond-mat.dis-nn 2024-12 conditional novelty 4.0 of 10

    Output-output correlations in finite Bayesian one-hidden-layer networks follow the kernel shape renormalization order parameter, with readout weight overlap equal to Q*_ab/λ1.

Pith tools