Pith. sign in

REVIEW 1 cited by

Optimal generalisation and learning transition in extensive-width shallow neural networks near interpolation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.18530 v3 pith:UM24OJY7 submitted 2025-01-30 stat.ML cond-mat.dis-nncond-mat.stat-mechcs.ITcs.LGmath.IT

classification stat.MLcond-mat.dis-nncond-mat.stat-mechcs.ITcs.LGmath.IT
keywords errorgeneralisationinterpolationlearningnetworkphaseweightdecays
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

We consider a teacher-student model of supervised learning with a fully-trained two-layer neural network whose width $k$ and input dimension $d$ are large and proportional. We provide an effective theory for approximating the Bayes-optimal generalisation error of the network for any activation function in the regime of sample size $n$ scaling quadratically with the input dimension, i.e., around the interpolation threshold where the number of trainable parameters $kd+k$ and of data $n$ are comparable. Our analysis tackles generic weight distributions. We uncover a discontinuous phase transition separating a "universal" phase from a "specialisation" phase. In the first, the generalisation error is independent of the weight distribution and decays slowly with the sampling rate $n/d^2$, with the student learning only some non-linear combinations of the teacher weights. In the latter, the error is weight distribution-dependent and decays faster due to the alignment of the student towards the teacher network. We thus unveil the existence of a highly predictive solution near interpolation, which is however potentially hard to find by practical algorithms.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Microscopic and collective signatures of feature learning in neural networks

    cond-mat.dis-nn 2025-08 conditional novelty 6.0 of 10

    In over-parameterized Bayesian one-hidden-layer networks, class-manifold separation becomes nonmonotonic in temperature and hidden weights develop data-dependent correlations, signatures of feature learning despite Ga...

Pith tools