Pith. sign in

REVIEW 4 cited by

Risk Comparisons in Linear Regression: Implicit Regularization Dominates Explicit Regularization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2509.17251 v2 pith:B3N35MC7 submitted 2025-09-21 stat.ML cs.LG

classification stat.MLcs.LG
keywords problemsregressionpolynomiallyridgedominateslinearregularizationtheory
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Existing theory suggests that for linear regression problems categorized by capacity and source conditions, gradient descent (GD) is always minimax optimal, while both ridge regression and online stochastic gradient descent (SGD) are polynomially suboptimal for certain categories of such problems. Moving beyond minimax theory, this work provides instance-wise comparisons of the finite-sample risks for these algorithms on any well-specified linear regression problem. Our analysis yields three key findings. First, GD dominates ridge regression: with comparable regularization, the excess risk of GD is always within a constant factor of that of ridge, but ridge can be polynomially worse even when tuned optimally. Second, GD is incomparable with SGD. While it is known that for certain problems GD can be polynomially better than SGD, the reverse is also true: we construct problems, inspired by benign overfitting theory, where optimally stopped GD is polynomially worse. Finally, GD dominates SGD for a significant subclass of problems -- those with fast and continuously decaying covariance spectra -- which includes all problems satisfying the standard capacity condition.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Minimax Optimal Early-Stopped Gradient Descent for Gaussian Mixture Classification

    stat.ML 2026-08 accept novelty 7.0 of 10

    Early-stopped gradient descent achieves the minimax-optimal classification error for Gaussian mixtures with label noise under fast-decaying covariance spectra, while interpolating classifiers can be exponentially worse.

  2. Beyond Negative-Ridge Endpoints: Mixed-Sign Spectral Regularization via Negative-Shifted Gradient Descent

    cs.LG 2026-07 accept novelty 7.0 of 10

    Early-stopped negative-shifted gradient descent beats every stable negative-ridge endpoint and every positive shrinker in gapped high-dimensional linear models, by polynomial risk factors.

  3. A Defense of the Quadratic Model

    cs.LG 2026-07 conditional novelty 7.0 of 10

    Local Taylor-expanded quadratic models reproduce a 150M-parameter LLM's validation loss for up to 10% of training late in the run, and LLM pretraining operates within a factor of 2 of a stochastic or deterministic edg...

  4. Statistical Inference on Gradient Flows

    math.ST 2026-05 unverdicted novelty 7.0 of 10

    Proves uniform CLT for gradient flows in ERM and constructs an algorithm-aware, inversion-free covariance estimator for asymptotically valid time-uniform confidence intervals.

Pith tools