REVIEW 46 cited by
Deep Neural Networks as Gaussian Processes
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
It has long been known that a single-layer fully-connected neural network with an i.i.d. prior over its parameters is equivalent to a Gaussian process (GP), in the limit of infinite network width. This correspondence enables exact Bayesian inference for infinite width neural networks on regression tasks by means of evaluating the corresponding GP. Recently, kernel functions which mimic multi-layer random neural networks have been developed, but only outside of a Bayesian framework. As such, previous work has not identified that these kernels can be used as covariance functions for GPs and allow fully Bayesian prediction with a deep neural network. In this work, we derive the exact equivalence between infinitely wide deep networks and GPs. We further develop a computationally efficient pipeline to compute the covariance function for these GPs. We then use the resulting GPs to perform Bayesian inference for wide deep neural networks on MNIST and CIFAR-10. We observe that trained neural network accuracy approaches that of the corresponding GP with increasing layer width, and that the GP uncertainty is strongly correlated with trained network prediction error. We further find that test performance increases as finite-width trained networks are made wider and more similar to a GP, and thus that GP predictions typically outperform those of finite-width networks. Finally we connect the performance of these GPs to the recent theory of signal propagation in random neural networks.
Forward citations
Cited by 46 Pith papers
-
Querying Kernel Methods Suffices for Reconstructing their Training Data
Query-only access to kernel regression, SVM and KDE models suffices to reconstruct their exact training points, via a measure-theoretic proof and image experiments.
-
Effect of Activation Functions on the Training of Overparametrized Neural Nets
In overparametrized 2-layer networks, non-smooth activations provably yield large NTK minimum eigenvalues, while smooth activations can have zero or exponentially small eigenvalues on low-dimensional data, predicting ...
-
The generalization error of random features regression: Precise asymptotics and double descent curve
Mei and Montanari derive the exact asymptotic test error of random features ridge regression and show it reproduces the full double descent phenomenon without any misspecified structure.
-
The Cost of Discretization in Functional Linear Regression: Minimax Rates and Adaptation
Matching minimax prediction rates for discretely observed functional linear regression are n^{-ν/(ν+1)}+(nm)^{-ν/κ} under independent design, and those two terms plus m^{-ν}+m^{-4α} under common design.
-
Spectral Anatomy of Quantum Gaussian Process Kernels
Normalized spectral entropy S(K)/log n of the kernel Gram matrix governs both dequantization in QGP regression and posterior pathologies, supported by proved bounds, a variance identity, and hardware transfer experiments.
-
Gauge-covariant stochastic neural fields: Stability and finite-width effects
A gauge-covariant stochastic neural field theory is introduced that derives the maximal Lyapunov exponent and amplification factor, showing finite-width effects as perturbative corrections to dressed kernels that leav...
-
Geometry of Neural Reinforcement Learning in Continuous State and Action Spaces
For wide two-layer linearized neural policies in deterministic continuous RL, the locally attainable states concentrate on a manifold of dimension at most 2da+1, independent of the state dimension.
-
SETOL: A Semi-Empirical Theory of (Deep) Learning
SETOL derives the HTSR layer quality metrics as integrated R-transforms of the layer spectral density, and proposes a determinant condition (ERG) as a marker of ideal learning.
-
The Importance of Being Lazy: Scaling Limits of Continual Learning
Increasing network width reduces forgetting only in lazy training; the best continual-learning performance occurs at a small critical level of feature learning that transfers across model sizes.
-
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer
Freezing query and key attention weights still lets transformers form induction heads and stay close to standard performance on language modeling, while random static attention (MixiT) fails on in-context tasks but su...
-
Adaptive kernel predictors from feature-learning infinite limits of neural networks
Feature-learning infinite-width neural networks are kernel machines with data-dependent kernels, defined by a min-max saddle point (Bayesian/Langevin) or a DMFT fixed point (gradient flow with weight decay).
-
The S-matrix bootstrap with neural optimizers I: zero double discontinuity
A neural optimizer solves the Atkinson-Mandelstam unitarity equations over the full space of amplitudes and, together with standard bootstrap methods, maps the allowed region of zero-double-discontinuity S-matrices, t...
-
A Gaussian Process framework for constraining the nuclear equation of state from microscopic calculations with correlated uncertainties
GPDiff fits a hierarchical Gaussian process to microscopic asymmetric-matter energies and propagates correlated uncertainties to EOS parameters and neutron-star matter properties.
-
Why Post-Norm Transformers Collapse: Attention Amplification and Gradient Repair Failure
Token similarity in Post-Norm decoders is amplified by causal attention at initialization, and RMSNorm backward contraction prevents gradients from repairing the resulting collapse.
-
A Statistical Difference between Single-Layer Learning and Hierarchical Learning in Wide Neural Networks
Even for wide three-layer nets, hierarchical learning of both layers yields a smaller asymptotic generalization error than single-layer learning with a fixed random kernel, because of singularities.
-
Spectral phase transitions and trainability in neural network learning dynamics
SGD on neural network weights induces a BBP phase transition that detaches signal eigenvalues from the random bulk, yielding an analytically solvable phase diagram for trainability in a linear teacher-student model.
-
Radial Basis Function Networks as Projection Heads in Self-Supervised Learning
RBFN projection heads serve as competitive replacements for MLP heads in SSL and enable SNS, a label-free metric from RBF parameters that correlates strongly with logistic regression evaluation.
-
Spontaneous symmetry breaking and Goldstone modes for deep information propagation
Equivariant neural networks support Goldstone-like modes enabling coherent information propagation across depth and recurrent iterations.
-
An Explainable Gaussian Process Auto-encoder for Tabular Data
A Gaussian-process autoencoder with a latent-space density estimator generates counterfactual examples for tabular data, with competitive or better scores on several evaluation metrics.
-
Opening the Black Box: Interpretable Remedies for Popularity Bias in Recommender Systems
A sparse autoencoder with synthetic user probes identifies popularity-encoding neurons in a recommender model, and steering those neurons improves exposure fairness with limited accuracy loss.
-
Bayesian Neural Network Surrogates for Bayesian Optimization of Carbon Capture and Storage Operations
Comparative evaluation of Bayesian Neural Network surrogates versus Gaussian Processes in Bayesian Optimization applied to Carbon Capture and Storage operations, presented as the first such application in reservoir en...
-
Controllable Feature Whitening for Hyperparameter-Free Bias Mitigation
Controllable Feature Whitening decorrelates target and bias features via a covariance-based whitening transform, reducing spurious-correlation reliance without adversarial training.
-
Interpretable Bayesian Tensor Network Kernel Machines with Automatic Rank and Feature Selection
A fully Bayesian tensor-network kernel machine with automatic rank and feature dimension selection, trained by mean-field variational inference.
-
Phi: Leveraging Pattern-based Hierarchical Sparsity for High-Efficiency Spiking Neural Networks
Phi decomposes SNN activations into pre-computed pattern rows plus sparse +/-1 corrections, yielding a 3.45x speedup and 4.93x energy savings over the Stellar accelerator.
-
Statistical Physics of Deep Neural Networks: Generalization Capability, Beyond the Infinite Width, and Feature Learning
A statistical-physics thesis derives a last-layer-only generalization bound, finite-width generalization formulas, and a Student's t-process equivalence for deep networks.
-
Distributionally Robust Coreset Selection under Covariate Shift
DRCS derives an upper bound on worst-case validation error under covariate shift and greedily chooses a coreset that minimizes this bound.
-
Efficient Bayesian Deep Ensembles via Analytic Predictive Inference
Independent neural predictors form a low-rank feature map that is aggregated by exact Bayesian linear regression, yielding a finite-rank Student-t process with competitive UCI regression performance.
-
Criticality analysis of nuclear binding energy neural networks
On a two-input nuclear binding energy network, the paper validates ANNFT predictions for variance, kurtosis, and an optimal depth-to-width ratio r*=0.034 under SGD, while adaptive optimizers obscure criticality.
-
Improved Physics-informed neural networks loss function regularization with a variance-based term
Adding a standard-deviation term to the mean loss in physics-informed neural networks reduces peak errors on Poisson, Burgers, elasticity, and Navier-Stokes benchmarks.
-
'In-Between' Uncertainty in Bayesian Neural Networks
MFVI in BNNs underestimates uncertainty between data regions, leading to overconfident OOD predictions, while linearised Laplace approximation performs better.
-
Discrete signaling mediates chaotic regularization in recurrent neural networks
Chaotic dynamics in RNNs induce local roughness but preserve global smoothness in representations, acting as an intrinsic regularizer and generating power-law spectral signatures.
-
Feature learning is decoupled from generalization in high capacity neural networks
Current feature learning measures quantify the magnitude of representation change, which the authors argue is decoupled from the generalization benefit that neural networks show over their neural tangent kernel.
-
Semantic-Aware Gaussian Process Calibration with Structured Layerwise Kernels for Deep Neural Networks
SAL-GP couples layerwise GPs with an additive kernel to calibrate classifier confidence, but experimental evidence is mixed, with one variant often no better than a single-layer GP.
-
Simplifying Graph Kernels for Efficient
SGTK and SGNK perform K-step graph aggregation before a single NTK or Gaussian process kernel update, yielding large speedups over GNTK with roughly competitive accuracy.
-
Providing Machine Learning Potentials with High Quality Uncertainty Estimates
A Bayesian neural network version of the ANI-1x potential provides uncertainty estimates that, in the tested cases, cover the model's errors at least as well as a nine-model ensemble.
-
Uncertainty separation via ensemble quantile regression
An ensemble of quantile regressors plus an iterative data-augmentation algorithm separates aleatoric from epistemic uncertainty in synthetic tasks.
-
Inherently Interpretable and Uncertainty-Aware Models for Online Learning in Cyber-Security Problems
Rolling-buffer Additive Gaussian Processes deliver competitive, interpretable, uncertainty-aware URL phishing classification in an online setting.
-
Finite size corrections for neural network Gaussian processes
For finite single-hidden-layer networks with symmetric weight initialization, the output distribution is a Gaussian with an O(1/N) fourth-Hermite correction, an instance of the classical Edgeworth expansion.
-
Lectures on Semiclassical Methods for Composite Operators
Lecture notes develop semiclassical methods to compute large-n scaling dimensions of composite operators in CFTs, recovering known results in free theory and deriving one-loop corrections at the Wilson-Fisher fixed point.
-
SoK: A Comprehensive Analysis of the Current Status of Neural Tangent Generalization Attacks with Research Directions
NTGA is the first clean-label generalization attack under black-box settings but is vulnerable to adversarial training and image transformations, with newer attacks outperforming it.
-
Bulk-boundary decomposition of neural networks
The paper reframes SGD training of deep networks as a local Lagrangian with data confined to the boundaries, but the advertised energy continuity equation is absent from the body.
-
Model-free Reinforcement Learning for Model-based Control: Towards Safe, Interpretable and Sample-efficient Agents
A perspective paper argues that model predictive control can be used as a learned policy in model-free reinforcement learning and reviews the methods and open problems.
-
Issues with Neural Tangent Kernel Approach to Neural Networks
The paper reports numerical evidence that the NTK equivalence theorem does not hold for finite-width neural networks trained with SGD, because trained networks and NTK kernel regressors behave differently when a layer...
-
Deep Learning Theory Review: An Optimal Control and Dynamical Systems Perspective
A review that frames neural networks as dynamical systems, SGD as stochastic dynamics, and training as mean-field optimal control to unify deep learning theory.
-
Deep learning applied to computational mechanics: A comprehensive review, state of the art, and the classics
A comprehensive review of deep learning techniques for computational mechanics, including LSTM for constitutive modeling, PINNs for PDE solving, optimizers, and kernel methods.
-
Physics-Driven Learning for Inverse Problems in Quantum Chromodynamics
A perspective article reviewing physics-driven machine learning for inverse problems in QCD, without introducing new data, derivations, or quantitative results.
Discussion (0). Continue with ORCID to comment.