In a sequential d,n→∞ then m→∞ limit with dm/n fixed, the spectrum of a deep linear Gaussian network's feature covariance converges to the free log-normal law, whose T-transform solves a Burgers equation.
Trevisan
5 Pith papers cite this work. Polarity classification is still indexing.
representative citing papers
Transformers converge pathwise to a stochastic particle system and SPDE in the scaling limit, exhibiting synchronization by noise and exponential energy dissipation when common noise is coercive relative to self-attention drift.
In the wide-width limit under Gaussian likelihood, the posterior of the network output is identified when the random covariance matrix is positive definite, with mild conditions ensuring invertibility and order-independent sequential limits.
Quantitative 2-Wasserstein bounds are established between finite-width deep neural networks and their infinite-width Gaussian limits using a Lindeberg principle for successive Gaussian replacement of weights.
In the LP/N = Θ(1) regime, Bayesian predictive posteriors for deep MLPs equal those of data-dependent kernels to first order, with a criterion identifying data processes that benefit from larger effective depth.
citing papers explorer
-
Geometric Dyson Brownian Motions and the Free Log-Normal Limit for a Non-Square Gaussian Matrix Product
In a sequential d,n→∞ then m→∞ limit with dm/n fixed, the spectrum of a deep linear Gaussian network's feature covariance converges to the free log-normal law, whose T-transform solves a Burgers equation.
-
Stochastic Scaling Limits and Synchronization by Noise in Deep Transformer Models
Transformers converge pathwise to a stochastic particle system and SPDE in the scaling limit, exhibiting synchronization by noise and exponential energy dissipation when common noise is coercive relative to self-attention drift.
-
Posterior Bayesian Neural Networks with Dependent Weights
In the wide-width limit under Gaussian likelihood, the posterior of the network output is identified when the random covariance matrix is positive definite, with mild conditions ensuring invertibility and order-independent sequential limits.
-
Universality in Deep Neural Networks: An approach via the Lindeberg exchange principle
Quantitative 2-Wasserstein bounds are established between finite-width deep neural networks and their infinite-width Gaussian limits using a Lindeberg principle for successive Gaussian replacement of weights.
-
Bayesian Inference with Shaped Deep Non-linear MLPs
In the LP/N = Θ(1) regime, Bayesian predictive posteriors for deep MLPs equal those of data-dependent kernels to first order, with a criterion identifying data processes that benefit from larger effective depth.