Pith. sign in

REVIEW 1 cited by

Information-theoretic reduction of deep neural networks to linear models in the overparametrized proportional regime

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2505.03577 v1 pith:LJ4ZJJ53 submitted 2025-05-06 math.ST cond-mat.dis-nncs.ITmath-phmath.ITmath.MPstat.MLstat.TH

classification math.STcond-mat.dis-nncs.ITmath-phmath.ITmath.MPstat.MLstat.TH
keywords deepneuralmodelnetworksregimeequivalencelinearoptimal
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We rigorously analyse fully-trained neural networks of arbitrary depth in the Bayesian optimal setting in the so-called proportional scaling regime where the number of training samples and width of the input and all inner layers diverge proportionally. We prove an information-theoretic equivalence between the Bayesian deep neural network model trained from data generated by a teacher with matching architecture, and a simpler model of optimal inference in a generalized linear model. This equivalence enables us to compute the optimal generalization error for deep neural networks in this regime. We thus prove the "deep Gaussian equivalence principle" conjectured in Cui et al. (2023) (arXiv:2302.00375). Our result highlights that in order to escape this "trivialisation" of deep neural networks (in the sense of reduction to a linear model) happening in the strongly overparametrized proportional regime, models trained from much more data have to be considered.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Microscopic and collective signatures of feature learning in neural networks

    cond-mat.dis-nn 2025-08 conditional novelty 6.0 of 10

    In over-parameterized Bayesian one-hidden-layer networks, class-manifold separation becomes nonmonotonic in temperature and hidden weights develop data-dependent correlations, signatures of feature learning despite Ga...

Pith tools