REVIEW 5 cited by
Surprises in High-Dimensional Ridgeless Least Squares Interpolation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
abstract
Interpolators -- estimators that achieve zero training error -- have attracted growing attention in machine learning, mainly because state-of-the art neural networks appear to be models of this type. In this paper, we study minimum $\ell_2$ norm ("ridgeless") interpolation in high-dimensional least squares regression. We consider two different models for the feature distribution: a linear model, where the feature vectors $x_i \in {\mathbb R}^p$ are obtained by applying a linear transform to a vector of i.i.d. entries, $x_i = \Sigma^{1/2} z_i$ (with $z_i \in {\mathbb R}^p$); and a nonlinear model, where the feature vectors are obtained by passing the input through a random one-layer neural network, $x_i = \varphi(W z_i)$ (with $z_i \in {\mathbb R}^d$, $W \in {\mathbb R}^{p \times d}$ a matrix of i.i.d. entries, and $\varphi$ an activation function acting componentwise on $W z_i$). We recover -- in a precise quantitative way -- several phenomena that have been observed in large-scale neural networks and kernel machines, including the "double descent" behavior of the prediction risk, and the potential benefits of overparametrization.
Forward citations
Cited by 5 Pith papers
-
On the Multiple Descent of Minimum-Norm Interpolants and Restricted Lower Isometry of Kernels
Minimum-norm interpolants in reproducing kernel Hilbert spaces have risk that can exhibit multiple peaks and valleys as the sample size grows, with peak locations predicted by the scaling d = n^α.
-
The generalization error of random features regression: Precise asymptotics and double descent curve
Mei and Montanari derive the exact asymptotic test error of random features ridge regression and show it reproduces the full double descent phenomenon without any misspecified structure.
-
Impact of Bottleneck Layers and Skip Connections on the Generalization of Linear Denoising Autoencoders
Two-layer linear denoising autoencoders show a bias-variance trade-off in bottleneck width, and skip connections reduce variance near the interpolation peak.
-
Eigenvalue Calibration for Semantic Embeddings of Large Language Models
Temperature scaling of density-matrix eigenvalues from LLM semantic embeddings optimizes proper-score calibration and corrects systematic overconfidence so entropy equals risk.
-
Quantifying the Prediction Uncertainty of Machine Learning Models for Individual Data
A per-sample confidence score derived from the pNML min-max regret is applied to linear regression and neural networks, and improves OOD detection, adversarial robustness, and active learning.
Discussion (0). Continue with ORCID to comment.