Pith. sign in

REVIEW 4 minor 23 references

Early-stopped negative-shifted gradient descent escapes the pole barrier that limits stable negative ridge in overparameterized regression.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 04:39 UTC pith:IHL2WBZ6

load-bearing objection A genuinely new and honestly-scoped theory paper: the finite-time negative-shifted path escapes the endpoint pole barrier, but the headline separation is conditional on exact head support and gapped spectra — still deserves a serious referee.

arxiv 2607.22474 v1 pith:IHL2WBZ6 submitted 2026-07-24 cs.LG math.STstat.MLstat.TH

Beyond Negative-Ridge Endpoints: Mixed-Sign Spectral Regularization via Negative-Shifted Gradient Descent

classification cs.LG math.STstat.MLstat.TH MSC 62J0768T0560B20
keywords negative-shifted gradient descentmixed-sign spectral regularizationearly stoppingnegative ridgeoverparameterized linear regressionbenign overfittingimplicit regularizationvalidation oracle
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper claims that negative ridge is limited not by its sign but by its endpoint: a stable negative-ridge filter must keep its pole below the smallest empirical eigenvalue, so it cannot reach the shift that cancels the implicit ridge created by weak tail directions. Early-stopped negative-shifted gradient descent escapes because its finite-time filter is smooth at that would-be pole and crosses the ridgeless level on a leading spectral prefix. In Gaussian spike-plus-flat models, this path is shown to achieve risk of order the tail-to-head ratio while every admissible negative-ridge endpoint pays the square root of that ratio and positive shrinkage keeps a constant head-bias floor. A general theorem lets the tail trace set the implicit floor and the squared tail spectrum control exposure, with the floor-critical path recovering all head scales at once and beating uniform rescalings of ridgeless regression once scales separate. If correct, NS-GD is a mixed-sign implicit regularizer whose sign change happens once, on a controlled prefix, with the stopping time setting the crossover.

Core claim

The central discovery is that the finite-time filter of negative-shifted gradient flow, f(ν,t)(μ) = μ ∫₀ᵗ e^{-(μ−ν)s} ds, has a removable singularity where the stable negative-ridge endpoint has a pole: at μ = ν the value is νt rather than infinity. This makes the displacement from ridgeless positive exactly on a leading spectral prefix and negative below a single crossover set by the stopping time. Placing the shift at the tail floor plus a head scale cancels the implicit shrinkage the tail induces on the head, while exposing the tail for only a bounded time. In the common-spike model this gives explicit head recovery with risk O(λ_T/λ_h), whereas every admissible stable endpoint pays Ω(sqr

What carries the argument

The removable-pole filter f(ν,t)(μ) = μ∫₀ᵗ e^{-(μ−ν)s} ds, the finite-time analogue of the negative-ridge endpoint filter μ/(μ−ν). Its smooth value at the would-be pole and its single spectral crossover are what let the algorithm use supercritical shifts outside the stable endpoint range. The other load-bearing part is the localized semigroup perturbation control: because the shifted generator is not a contraction, the analysis restricts perturbation integrals to the positive head block, so off-diagonal head-tail coupling contributes only quadratically through leave-and-return terms.

Load-bearing premise

The signal is exactly supported on the leading k coordinates, with a strictly gapped spectrum and a tail that behaves like pure noise floor; the whole exposure control breaks if the tail carries any signal mass that leaks into the supercritical shift.

What would settle it

Simulate a Gaussian design with a gapped spectrum, a head-supported signal, and deliberately add tail signal mass of size comparable to the head mass. If the floor-critical NS-GD path does not degrade by at least the predicted leakage scale, or if the risk separation from the best admissible negative-ridge endpoint persists, the trace-square-mass exposure control is wrong.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • NS-GD can legally use negative shifts larger than the smallest empirical eigenvalue, something no stable negative-ridge endpoint can do, and still stay bounded by early stopping.
  • The finite-time filter realizes anti-shrinkage on a leading prefix and shrinkage or exposure control on the lower spectrum, so it can correct implicit over-shrinkage without inflating low-mode variance.
  • In the common-spike model the stopped path beats every admissible negative-ridge endpoint by a polynomial factor in risk, and beats positive ridge and early stopping by a larger factor.
  • For heterogeneous heads with a high-effective-rank tail, one shift chosen from the tail trace recovers all resolved head scales at once, out-performing uniform rescaling of ridgeless once the head scales separate.
  • A validation-selected finite grid inherits these separations under an explicit validation-size condition, so the improvement is not limited to oracle-tuned paths.
  • In no-gap power-law spectra the theory's separation closes; the signed path still helps, but only as a shape effect rather than a sign effect.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The removable-pole mechanism suggests a practical diagnostic: if a fitted finite-time negative-shift filter shows anti-shrinkage on a leading prefix that stops well before the tail, the model is in the gapped head-tail regime; in a no-gap spectrum the prefix should shrink toward a shape-only effect.
  • The same filter calculus should transfer to kernel ridge regression, but the trace-class Mercer tail makes the implicit floor vanish at rate 1/n there, so the practical payoff would be concentrated in high-dimensional kernels with divergent tail trace rather than fixed-domain kernels.
  • A testable extension of the theorem would replace head-supported signals with source-condition tail decay: the paper's robustness experiments suggest the negative-sign advantage degrades into a shape effect, so a quantitative threshold on tail signal mass could tell practitioners when NS-GD is safe to prefer over positive shrinkage.
  • Because the discrete one-step filter is independent of the shift, the paper leaves open whether large-step ordinary gradient descent could emulate the same effect; checking that would clarify whether the mixed-sign prefix geometry is essential or merely convenient.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

0 major / 4 minor

Summary. The paper introduces negative-shifted gradient flow/descent (NS-GF/NS-GD), a two-parameter spectral regularization path for overparameterized linear regression. Its central observation is that the finite-time filter f_{ν,t}(μ)=μ∫₀ᵗ e^{-(μ-ν)s}ds is smooth at μ=ν (with removable value νt), whereas the stable negative-ridge endpoint A_ν(μ)=μ/(μ-ν) has a pole below the smallest empirical eigenvalue. Corollaries 3–4 show that the finite-time filter exceeds the ridgeless level exactly on a leading spectral prefix. In Gaussian spike-plus-flat and heterogeneous head–tail models, under the explicit assumptions of a head-supported signal, a gapped spectrum, and high-effective-rank tail, the paper proves risk separations: the path attains O(λ_T/λ_h) while admissible negative endpoints pay Ω(√(λ_T/λ_h)) and positive shrinkage retains Ω(1) (Theorem 9); with heterogeneous heads the trace T_k sets the floor and S_k controls exposure, with a shape-defect lower bound on uniform rescalings of ridgeless (Theorem 11). A finite-grid hold-out inequality (Theorem 15) transfers the separation to validation-selected Algorithm 1. The paper also reports extensive paired simulations confirming the predicted endpoint wall and the finite-time prefix geometry.

Significance. If correct, the paper establishes a genuinely new object in spectral regularization: a mixed-sign filter that is not available to positive shrinkage or to stable negative-ridge endpoints. The endpoint–path asymmetry is conceptually interesting and is supported by careful random-matrix analysis, not by heuristic arguments. The paper is strong methodologically: comparators are oracle-tuned over their full classes, the path parameters are computed from population model quantities rather than fitted to data, the main structural claims (removable-pole geometry, head identity λ_h h(a+λ_h)=1, tail leakage bounds, and the polynomial common-spike example) are independently checkable and appear sound, and the claim of reproducible code is credible. The validation oracle inequality is a useful transfer tool. The principal weakness is that the separation theorems are certified only for exactly head-supported signals; the paper's own robustness experiments show that the advantage degrades outside this regime. This is a real scope limitation, but it is explicitly stated and does not invalidate the conditional theorems.

minor comments (4)
  1. [Section 3, shared framework; Lemma 24] The exact head-support assumption β*_j=0 for j>k is load-bearing for Theorems 9 and 11, and Lemma 24 labels this as 'head support removes true tail-signal bias' rather than giving a quantitative degradation bound. The paper should add an explicit statement in Section 5 that no continuous risk bound is proved as a function of tail-signal mass, and that the polynomial separation is certified only at this support restriction. This is not a correctness issue for the stated theorems, but it would prevent readers from overinterpreting the 'general high-effective-rank tail' language in the abstract.
  2. [Figures 2–4 and Section 3 display equations] Several displayed equations and figure captions contain encoding artifacts, e.g., '/uni00000017/uni00000013/...' sequences in the Figure 2/3 captions and in the shared-path-and-comparator display. These need to be regenerated with the correct mathematical glyphs.
  3. [Section 3.3, validation size paragraph] The validation-size condition nval ≫ r_n^{-2} log(|G|/δ) is terse; its instantiation in the polynomial example (nval ≫ n^{2ζ} to retain the full rate, nval ≫ n^ζ to retain the endpoint separation) would be clearer if written as explicit requirements on the validation sample size.
  4. [Section 3, shared risk notation] The definition R⋆_scale = inf_{c∈R} Rpop(μ↦ c·1{μ>0}|X) should be explained explicitly as a uniform rescaling of the ridgeless row-space filter; the notation 1{μ>0} alone is easy to confuse with a population-level indicator.

Circularity Check

0 steps flagged

No significant circularity: theorems are proved from explicit population-model parameters, with no self-citations or fitted-input predictions.

full rationale

The derivation chain is self-contained. The filter f_{ν,t}(µ)=µ∫_0^t e^{-(µ-ν)s} ds and its removable value at µ=ν follow directly from integrating the negative-shifted gradient flow; Corollaries 3–4 derive the leading-prefix geometry from the monotonicity of g_ν, not from any fitted equivalence or cited result. The separation theorems (Theorem 9; Theorem 11; Corollaries 12–14) choose (ν*,t*)=(a+λ_h,1/λ_h) and (a,t_fc) from explicit population quantities (λ_h, a=T_k/n, λ_-, S_k, Θ_H) and prove the upper bounds through Gaussian concentration and localized Duhamel estimates; no parameter is fitted to the data whose risk is then called a prediction. Comparator lower bounds R*_neg, R*_+, R*_scale are proven for oracle-tuned classes, so the comparison is conservative, not circular. The reference list contains no self-citations and no uniqueness theorem imported from the author's prior work; cited results (Marchenko–Pastur, Tsigler–Bartlett, Laurent–Massart, Duhamel) are external mathematical tools. The manuscript explicitly acknowledges its scope limits—exact head support, gapped spectra, and the m=1 reduction to ordinary GD—and these are robustness/extension concerns, not circularity. No load-bearing step reduces by construction to its own inputs.

Axiom & Free-Parameter Ledger

4 free parameters · 15 axioms · 1 invented entities

The central claim rests on standard random-matrix and perturbation machinery (listed as standard_math), plus a clearly stated Gaussian gapped head-tail regime with head-supported signal. The only hand-chosen quantities are the signed level nu, stopping time t, step size eta, and iteration count m - all derived analytically from population model parameters (T_k, lambda_h, S_k), not fitted to data. In Algorithm 1 the regularization coordinates are selected by validation, which is the legitimate, disclosed protocol. No invented physical entities or fitted constants are smuggled in; the power-law and flat-tail conditions are explicit regime windows.

free parameters (4)
  • nu (signed level) = nu_* = a + lambda_h (Case I); nu_fc = a = T_k/n (Case II)
    Theoretically chosen 'floor-critical' shifts at which the separation theorems are proved; derived from population model quantities (tail trace T_k, head scale lambda_h), not fitted to data. In Algorithm 1, nu is selected by validation over a finite grid.
  • t (finite horizon / stopping time) = t_* = 1/lambda_h (Case I); t_fc = L_k/lambda_-, L_k = 1/2 log(1 + Theta_H/V_{T,k}) (Case II)
    Chosen so the head exponent vanishes exactly (Case I) or head recovery balances tail exposure (Case II). Proofs require w_k t = o(1); Theorem 9 and Theorem 11 are proved at these specific pairs.
  • eta (step size) = eta_* = (m lambda_h)^{-1} (Case I); eta_m = t_fc/m (Corollary 14); general small-step condition eta_nu(muhat_1 + nu) <=
    Stability and discretization parameter; central to the discrete transfer in Corollary 14 and to Algorithm 1's small-step family.
  • m (iteration count) = free integer >= 1; Corollary 14 requires m >= (lambda_-/w_k)(1 + L_k^2)
    Computational/trade-off coordinate. m = 1 is degenerate (filter independent of the signed level; reduces to a single large-step ordinary GD update, as admitted in Section 3.1).
axioms (15)
  • standard math Marchenko-Pastur law / semicircle law for sample covariance matrices in the regime d_T/n -> infinity (Bai-Yin 1988; Marchenko-Pastur 1967)
    Proposition 7 and Lemma 22 use the bulk-width order a - muhat^+_min = Theta(w_T), with the tail Gram KT = (lambda_T/n) Z_T Z_T^T.
  • standard math Gaussian singular-value and operator-norm concentration (Vershynin 2018, Chapter 4)
    Lemma 22 (head Wishart concentration, tail Gram concentration), Lemma 26 (weighted Gaussian Gram bound), Lemma 27 (localization event).
  • standard math Laurent-Massart chi-square tail bounds (Laurent and Massart 2000)
    Lemma 26 and the proof of Theorem 15 (validation residuals are chi-square distributed).
  • standard math Duhamel semigroup perturbation formula and matrix functional calculus (Higham 2008)
    Lemmas 22 and 28 control h_t(K) - h_t(K_0) for the noncontractive shifted semigroup.
  • standard math Matrix inversion lemma, Weyl and Cauchy interlacing inequalities (Horn and Johnson 2012)
    Used throughout Lemmas 23-27 for head maps, endpoint resolvents, and eigenvalue interlacing.
  • standard math Invariant graph subspace / operator Riccati equation existence (Kostrykin, Makarov, Motovilov 2005)
    Lemma 28's graph-angle and quadratic block-relocation estimates are the core new technical tool for the heterogeneous-head theorem.
  • standard math Conditional Gaussian risk identities (Proposition 1; cf. Hastie et al. 2022; Wu and Xu 2020)
    The exact bias-variance identity and the Q-metric oracle (Proposition 6, completing the square) are foundational for all risk comparisons.
  • domain assumption Gaussian design X = Z Sigma^{1/2} with diagonal Sigma = diag(Lambda_H, Lambda_T)
    Shared framework in Section 3; all concentration events and the MP barrier are derived under it.
  • domain assumption Head-supported signal: beta*_j = 0 for all j > k
    Section 3, 'The signal is head-supported'. The floor-canceling mechanism assumes the tail carries no signal; Lemma 24 bounds leakage only via head support.
  • domain assumption Gapped spectrum: k = o(n), lambda_{k+1}/lambda_- -> 0, lambda_+/lambda_- = O(1), a = T_k/n = Theta(lambda_-)
    Assumptions 8 and 10; the head/tail split is load-bearing for both theorems and for the separations in Corollaries 12-13.
  • domain assumption High-effective-rank tail: r_k -> 0 and k w_k^2/S_k -> 0 (projected tail-square-mass requirement)
    Assumption 10; ensures the tail Gram concentrates around a I_n and that head-frame perturbations are small.
  • domain assumption Bounded signal and noise scales: Theta_H, sigma_epsilon^2 and their reciprocals bounded; sigma_epsilon^2 > 0
    Shared framework and Theorem 15; the risk separations are proved in this interior regime.
  • domain assumption Independent Gaussian validation sample of size n_val, independent of training data
    Theorem 15 requires x_val ~ N(0, Sigma), y_val = x_val^T beta* + epsilon_val, validation independent of D_tr.
  • ad hoc to paper Power-law dimension window max{1/(1-alpha), 3/(3-4alpha)} < q < 3 (Corollary 13)
    Regime chosen so that d_n = n^q gives a still-gapped tail, vanishing leakage, and a vanishing head-noise term; the theorem is proved inside this window.
  • ad hoc to paper Flat-tail separation conditions w_k L_k^2/lambda_- -> 0 and k lambda_-/(n w_k) -> 0 (Corollary 12)
    Additional conditions needed only for the ratio-to-zero claims (R_sign/R*_neg -> 0 and R_sign/(R*_scale ^ R*_+) -> 0), beyond the bound-level assumptions of Theorem 11.
invented entities (1)
  • Negative-shifted gradient descent (NS-GD / NS-GF), the removable-pole mixed-sign filter independent evidence
    purpose: New spectral regularization path that realizes head anti-shrinkage with lower-spectrum shrinkage or exposure control, inaccessible to positive shrinkage and to the stable negative-ridge endpoint.
    Not an unexplained postulate: the paper proves risk bounds at theoretically chosen (nu, t), validates the predicted separations in simulations (Figures 3-4), and Theorem 15 supplies a hold-out guarantee for the deployed validation-selected version.

pith-pipeline@v1.3.0-alltime-deepseek · 41572 in / 38918 out tokens · 388396 ms · 2026-08-01T04:39:46.050528+00:00 · methodology

0 comments
read the original abstract

In overparameterized linear regression, many weak spectral directions act like a ridge penalty on the signal-bearing spectrum; negative ridge is the natural correction, pushing filters above one. The stable negative-ridge endpoint, however, is structurally limited: its pole must stay below the smallest nonzero empirical eigenvalue, and it anti-shrinks smaller eigenvalues more than larger ones. Early-stopped negative-shifted gradient descent escapes this constraint. Its filter is smooth at the would-be pole and mixed-sign-capable: above-ridgeless directions form a leading prefix, with lower directions shrunk or exposure-controlled while stopping sets the crossover. In a Gaussian spike-plus-flat model we discover a Marchenko-Pastur barrier: the shift that cancels the implicit penalty lies a bulk width above the smallest empirical eigenvalue, and the stopped path improves on every admissible endpoint by a polynomial factor in risk under explicit conditions. Our main theorem permits a general high-effective-rank tail: its trace sets the implicit floor, its squared spectrum controls exposure, and the floor-critical path recovers all head scales at once, beyond positive shrinkage and, once scales separate, every uniform rescaling of ridgeless. Handling the noncontractive shifted dynamics is the central technical challenge; localized Duhamel integrals control them. A finite-grid hold-out inequality transfers the separations to the validation-selected algorithm.

Figures

Figures reproduced from arXiv: 2607.22474 by Peng Zhao.

Figure 1
Figure 1. Figure 1: Endpoint–path asymmetry for mixed-sign spectral regularization. A stable negative-ridge [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Numerical confirmation of the Marchenko–Pastur endpoint barrier on Gaussian common [PITH_FULL_IMAGE:figures/full_fig_p010_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Conditional-risk separation in exact Gaussian spike-plus-flat designs: common spikes in [PITH_FULL_IMAGE:figures/full_fig_p017_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Separation across tail geometry at fixed sample size ( [PITH_FULL_IMAGE:figures/full_fig_p018_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Spiked+Flat representative run: final spectral filters and cumulative population RMSE [PITH_FULL_IMAGE:figures/full_fig_p020_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Power-law αPL = 1 diagnostic (no-gap spectrum), all four methods validation-selected. Left: learned filters fi(T). Right: cumulative top-k mode inclusion, reporting exact conditional population RMSE. The negative-ridge endpoint collapses to ridgeless (its sign gives no gain), while NS-GD adds only modest head anti-shrinkage and shrinks the tail; all methods cluster, so the advantage here is a shape effect … view at source ↗
Figure 7
Figure 7. Figure 7: Spiked+Flat representative run: NS-GD versus baselines over training. Left: head-mode [PITH_FULL_IMAGE:figures/full_fig_p022_7.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

23 extracted references · 2 canonical work pages

  1. [5]

    Chen Cheng and Andrea Montanari

    doi: 10.1007/s10208-006-0196-8. Chen Cheng and Andrea Montanari. Dimension free ridge regression.The Annals of Statistics, 52 (6):2879–2912,

  2. [13]

    Dmitry Kobak, Jonathan Lomond, and Benoit Sanchez

    doi: 10.1214/26-EJS2497. Dmitry Kobak, Jonathan Lomond, and Benoit Sanchez. The optimal ridge penalty for real-world high-dimensional data can be zero or negative due to the implicit ridge regularization.Journal of Machine Learning Research, 21(169):1–16,

  3. [16]

    Laura Lo Gerfo, Lorenzo Rosasco, Francesca Odone, Ernesto De Vito, and Alessandro Verri

    doi: 10.1214/19-AOS1849. Laura Lo Gerfo, Lorenzo Rosasco, Francesca Odone, Ernesto De Vito, and Alessandro Verri. Spectral algorithms for supervised learning.Neural Computation, 20(7):1873–1897,

  4. [17]

    Vladimir A

    doi: 10.1162/neco.2008.05-07-517. Vladimir A. Marchenko and Leonid A. Pastur. Distribution of eigenvalues for some sets of random matrices.Mathematics of the USSR-Sbornik, 1(4):457–483,

  5. [21]

    Alexander Tsigler and Peter L

    doi: 10.48550/arXiv.2406.04425. Alexander Tsigler and Peter L. Bartlett. Benign overfitting in ridge regression.Journal of Machine Learning Research, 24(123):1–76,

  6. [22]

    Full version available as arXiv:2509.17251

    URL https://proceedings.mlr.press/ v336/wu26a.html. Full version available as arXiv:2509.17251. Yuan Yao, Lorenzo Rosasco, and Andrea Caponnetto. On early stopping in gradient descent learning. Constructive Approximation, 26:289–315,

  7. [23]

    Difan Zou, Jingfeng Wu, Vladimir Braverman, Quanquan Gu, Dylan P

    doi: 10.1007/s00365-006-0663-2. Difan Zou, Jingfeng Wu, Vladimir Braverman, Quanquan Gu, Dylan P. Foster, and Sham M. Kakade. The benefits of implicit regularization from SGD in least squares problems. InAdvances in Neural Information Processing Systems, volume 34, pages 21169–21181,

  8. [1970]

    doi: 10.1080/00401706.1970.10488634. Roger A. Horn and Charles R. Johnson.Matrix Analysis. Cambridge University Press, 2 edition,

  9. [1988]

    doi: 10.1214/aop/1176991792. Peter L. Bartlett, Philip M. Long, G´ abor Lugosi, and Alexander Tsigler. Benign overfitting in linear regression.Proceedings of the National Academy of Sciences, 117(48):30063–30070,

  10. [1996]

    Trevor Hastie, Andrea Montanari, Saharon Rosset, and Ryan J

    doi: 10.1007/978-94-009-1740-8. Trevor Hastie, Andrea Montanari, Saharon Rosset, and Ryan J. Tibshirani. Surprises in high- dimensional ridgeless least squares interpolation.The Annals of Statistics, 50(2):949–986,

  11. [2000]

    ridgeless

    doi: 10.1214/aos/1015957395. Tengyuan Liang and Alexander Rakhlin. Just interpolate: Kernel “ridgeless” regression can generalize.The Annals of Statistics, 48(3):1329–1347,

  12. [2005]

    B´ eatrice Laurent and Pascal Massart

    doi: 10.1007/s00020-003-1248-6. B´ eatrice Laurent and Pascal Massart. Adaptive estimation of a quadratic functional by model selection.The Annals of Statistics, 28(5):1302–1338,

  13. [2007]

    Andrea Caponnetto and Ernesto De Vito

    doi: 10.1016/j.jco.2006.07.001. Andrea Caponnetto and Ernesto De Vito. Optimal rates for the regularized least-squares algorithm. Foundations of Computational Mathematics, 7(3):331–368,

  14. [2008]

    Arthur E Hoerl and Robert W Kennard

    doi: 10.1137/1.9780898717778. Arthur E Hoerl and Robert W Kennard. Ridge regression: Biased estimation for nonorthogonal problems.Technometrics, 12(1):55–67,

  15. [2011]

    Dominic Richards, Jaouad Mourtada, and Lorenzo Rosasco

    doi: 10.1109/Allerton.2011.6120320. Dominic Richards, Jaouad Mourtada, and Lorenzo Rosasco. Asymptotics of ridge(less) regression under general source condition. InProceedings of The 24th International Conference on Artificial Intelligence and Statistics, volume 130 ofProceedings of Machine Learning Research, pages 3889–3897. PMLR,

  16. [2012]

    Ting Hu and Yunwen Lei

    doi: 10.1017/CBO9781139020411. Ting Hu and Yunwen Lei. Early stopping for iterative regularization with general loss functions. Journal of Machine Learning Research, 23(339):1–36,

  17. [2018]

    doi: 10.1214/17-AOS1549. Heinz W. Engl, Martin Hanke, and Andreas Neubauer.Regularization of Inverse Problems. Kluwer Academic Publishers, Dordrecht,

  18. [2020]

    Frank Bauer, Sergei Pereverzev, and Lorenzo Rosasco

    doi: 10.1073/pnas.1907378117. Frank Bauer, Sergei Pereverzev, and Lorenzo Rosasco. On regularization algorithms in learning theory.Journal of Complexity, 23(1):52–72,

  19. [2021]

    On regularization via early stopping for least squares regression.arXiv preprint arXiv:2406.04425,

    Rishi Sonthalia, Jackie Lok, and Elizaveta Rebrova. On regularization via early stopping for least squares regression.arXiv preprint arXiv:2406.04425,

  20. [2022]

    Nicholas J

    doi: 10.1214/21-AOS2133. Nicholas J. Higham.Functions of Matrices: Theory and Computation. Society for Industrial and Applied Mathematics, Philadelphia, PA,

  21. [2024]

    Edgar Dobriban and Stefan Wager

    doi: 10.1214/24-AOS2449. Edgar Dobriban and Stefan Wager. High-dimensional asymptotics of prediction: Ridge regression and classification.The Annals of Statistics, 46(1):247–279,

  22. [2025]

    Pratik Patil and Jin-Hong Du

    doi: 10.1007/s10107-024-02171-3. Pratik Patil and Jin-Hong Du. Generalized equivalences between subsampling and ridge regulariza- tion. InAdvances in Neural Information Processing Systems, volume 36,

  23. [2026]

    260412235902

    doi: 10.4310/ATMP. 260412235902. Zhidong D. Bai and Y. Q. Yin. Convergence to the semicircle law.The Annals of Probability, 16(2): 863–875,