Pith. sign in

REVIEW 3 cited by

More Data Can Hurt for Linear Regression: Sample-wise Double Descent

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1912.07242 v1 pith:KK2VDR3I submitted 2019-12-16 stat.ML cs.LGcs.NEmath.STstat.TH

classification stat.MLcs.LGcs.NEmath.STstat.TH
keywords linearregressionsamplesbehaviordatadescentestimatorincreases
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this expository note we describe a surprising phenomenon in overparameterized linear regression, where the dimension exceeds the number of samples: there is a regime where the test risk of the estimator found by gradient descent increases with additional samples. In other words, more data actually hurts the estimator. This behavior is implicit in a recent line of theoretical works analyzing "double-descent" phenomenon in linear models. In this note, we isolate and understand this behavior in an extremely simple setting: linear regression with isotropic Gaussian covariates. In particular, this occurs due to an unconventional type of bias-variance tradeoff in the overparameterized regime: the bias decreases with more samples, but variance increases.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Influence Diagnostics in High-dimensional M-estimation: Precise Asymptotics

    stat.ML 2026-07 accept novelty 7.0 of 10

    Under Gaussian design with n ≍ d, the empirical distribution of leave-one-out influences for convex M-estimators converges to the pushforward of a four-dimensional Gaussian through an explicit nonlinear map built from...

  2. The Geometry of Saturation: Effective Rank Predicts When Labels Stop Helping in Few-Shot Classification

    cs.LG 2026-06 unverdicted novelty 6.0 of 10

    The saturation index S(K)=erank(pooled within-class covariance)/K tracks when few-shot labels become redundant for fixed linear probes, with within-task median Spearman ρ≈0.81 on 17 binary tasks.

  3. Transformers Don't In-Context Learn Least Squares Regression

    cs.LG 2025-07 conditional novelty 6.0 of 10

    In-context regression transformers do not approximate OLS: they underperform it even in-distribution, fail on out-of-subspace prompts, and their failures correlate with a low-rank spectral signature in the residual stream.

Pith tools