Pith. sign in

REVIEW 2 minor 46 references

Resetting gradient flow to the origin at Poisson rate r produces the ridge estimator exactly.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.3

2026-06-29 09:09 UTC pith:MLCFA2CF

load-bearing objection The paper derives ridge as the stationary mean under Poisson resetting in linear gradient flow on quadratics and shows exponential is the unique renewal law for exact ridge filters across eigenvalues.

arxiv 2605.30059 v1 pith:MLCFA2CF submitted 2026-05-28 cs.LG cond-mat.stat-mechstat.ML

Ridge Regression from Poisson Resetting: A Renewal Perspective on Spectral Regularization

classification cs.LG cond-mat.stat-mechstat.ML
keywords ridge regressionstochastic resettingrenewal theoryspectral regularizationgradient flowPoisson processstationary age
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper connects stochastic resetting from physics to ridge regularization by showing that resetting continuous-time linear gradient flow to the origin at constant rate r yields a stationary mean identical to the ridge solution with penalty lambda equal to r. This identity follows from the Laplace-transform link between ridge and exponential-time averaging of the flow trajectory, now reinterpreted as the stationary age under Poisson resetting. The exponential distribution is proven to be the only renewal law that induces the exact ridge spectral filter in every eigendirection for arbitrary positive curvature. Non-exponential renewal laws generate other filters, while an additive Ornstein-Uhlenbeck extension matches the ridge mean but carries extra stationary covariance from noise and reset timing.

Core claim

For linear gradient flow with Poisson resetting to the origin at rate r, the stationary mean is exactly the ridge estimator (X^T X + r I)^{-1} X^T y. The exponential reset-time distribution is the unique renewal law whose stationary mean reproduces scalar ridge in every eigendirection as an exact filter identity for every positive curvature, while non-exponential renewal laws generate alternative spectral filters.

What carries the argument

The stationary age distribution of the renewal reset process, which supplies the Laplace-transform weighting that converts the gradient-flow trajectory into the ridge spectral filter when the reset law is exponential.

Load-bearing premise

The results require continuous-time gradient flow on quadratic objectives with isotropic resetting.

What would settle it

Compute the stationary mean of the reset gradient flow on a non-quadratic loss and verify whether it deviates from any ridge solution of the form (X^T X + lambda I)^{-1} X^T y.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Non-exponential renewal reset laws induce alternative spectral regularizers whose predictive behavior can differ from ridge.
  • In the Ornstein-Uhlenbeck extension the mean matches ridge but the stationary covariance is nonzero due to accumulated noise and reset-timing variance.
  • Stylized experiments directly compare the deterministic filters and show when non-exponential cases diverge from ridge predictions.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The same stationary-age mechanism could be approximated in discrete-time optimization by periodic random restarts to produce tunable regularization.
  • Direction-dependent reset rates might induce anisotropic spectral filters that regularize differently along different eigendirections.
  • Testing the identity on objectives with state-dependent noise would clarify how far the mean-matching property survives beyond the additive OU case.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

0 major / 2 minor

Summary. The paper connects stochastic resetting from non-equilibrium statistical physics to ridge regularization in statistical learning. For continuous-time linear gradient flow on quadratic objectives with isotropic Poisson resetting at rate r, the stationary mean equals the ridge estimator (X^T X + r I)^{-1} X^T y. This reinterprets the known Laplace-transform link between ridge and exponential averaging via the stationary age of the renewal process. The exponential reset distribution is shown to be the unique renewal law reproducing the exact ridge filter 1/(λ + r) in every eigendirection for all λ > 0; non-exponential renewals yield alternative spectral filters. An additive Ornstein-Uhlenbeck extension models fluctuations (as a stylized SGD approximation), where equality holds only at the mean level due to nonzero stationary covariance. Results are scoped to quadratic objectives, continuous-time flow, and isotropic resetting; stylized experiments illustrate filter differences.

Significance. If the central derivation holds, the work supplies a renewal-theoretic interpretation of spectral regularization, showing how Poisson resetting recovers ridge exactly and how other renewal laws generate distinct filters. It explicitly separates the mean identity from covariance behavior and leverages a known Laplace-transform relationship without introducing free parameters or circularity. This framing may suggest new regularization schemes derived from renewal processes and clarifies when mean-level equivalence does or does not extend to fluctuations. The explicit scoping to linear gradient flow on quadratics and the uniqueness argument for the exponential case are strengths that keep the claims falsifiable and contained.

minor comments (2)
  1. [Abstract] Abstract, final sentence: the scoping statement is helpful but could be repeated verbatim in §1 or the conclusion to ensure readers do not over-generalize the filter identity beyond quadratic objectives.
  2. [OU extension paragraph] The OU-fluctuation analysis is presented as a separate stylized model; a brief remark on why state-independent diffusion is chosen (versus state-dependent noise more typical of SGD) would aid interpretability without altering the mean result.

Simulated Author's Rebuttal

0 responses · 0 unresolved

We thank the referee for the careful reading, the accurate summary of the contribution, and the recommendation to accept.

Circularity Check

0 steps flagged

No significant circularity identified

full rationale

The paper's central derivation reinterprets the known external Laplace-transform link between ridge regression and exponential averaging of gradient flow through the stationary age of a renewal process. The uniqueness claim for the exponential reset law follows directly from imposing the exact filter identity 1/(λ + r) for every curvature λ > 0; this is a mathematical consequence of the requirement rather than a reduction of the target result to its own inputs. No parameters are fitted to data and then renamed as predictions, no self-citations are load-bearing, and the claims are explicitly scoped to continuous-time linear gradient flow on quadratic objectives with isotropic resetting. The OU-fluctuation analysis is separated and acknowledges that equality fails at the covariance level by construction. The derivation is therefore self-contained against external benchmarks.

Axiom & Free-Parameter Ledger

0 free parameters · 2 axioms · 0 invented entities

The paper relies on standard mathematical relationships from probability (renewal theory, stationary age) and a known Laplace-transform identity from prior optimization literature; no new free parameters are fitted to data and no new physical entities are postulated.

axioms (2)
  • domain assumption Laplace-transform relationship between ridge regression and exponential-time averaging of gradient flow
    Invoked to equate resetting stationary mean with ridge estimator.
  • standard math Existence and properties of stationary age distribution for renewal processes
    Used to interpret exponential time as stationary age under Poisson resetting.

pith-pipeline@v0.9.1-grok · 5791 in / 1421 out tokens · 54661 ms · 2026-06-29T09:09:21.630274+00:00 · methodology

0 comments
read the original abstract

We connect stochastic resetting from non-equilibrium statistical physics with ridge regularization in statistical learning. For linear gradient flow, resetting to the origin at rate $r$ produces stationary mean $(X^\top X+rI)^{-1}X^\top y$, exactly the ridge estimator with penalty $\lambda=r$. This uses the known Laplace-transform relationship between ridge regression and exponential-time averaging of gradient flow, with the exponential time now interpreted as the stationary age associated with Poisson resetting. We then extend this identity to general renewal reset laws: the exponential reset time distribution is the unique renewal law whose stationary mean reproduces scalar ridge in every eigendirection as an exact filter identity for every positive curvature, while non-exponential renewal laws generate alternative spectral filters. At the fluctuation level, we study a separate additive Ornstein-Uhlenbeck extension with constant diffusion, interpreted as a stylized SGD approximation. In this setting, the equality holds only at the level of the mean, since the reset process has a nonzero stationary covariance from accumulated OU noise and reset-timing variance, whereas deterministic ridge is a fixed estimator with the same center. Stylized experiments compare the deterministic renewal-induced filters directly and illustrate when filters induced by non-exponential reset-time laws can differ predictively from ridge. The results for the stationary mean and the induced spectral filters are established for continuous-time gradient flow with isotropic resetting on quadratic objectives; the covariance and risk formulas additionally assume additive noise with state-independent covariance.

Figures

Figures reproduced from arXiv: 2605.30059 by Petar Jolakoski.

Figure 1
Figure 1. Figure 1: Gradient-flow trajectories with and without Poisson resetting in a two-parameter linear regression problem. The left panel shows a weakly identified coefficient, corresponding to a predictor with low variance, while the right panel shows a strongly identified coefficient. Blue curves are trajectories with resetting at rate r, the black curve is their ensemble mean, and the orange curve is the corresponding… view at source ↗
Figure 2
Figure 2. Figure 2: Two-mode comparison of ridge and renewal-induced shrinkage. Each panel fixes two representative directions, a weak feature with µτ = 1 and a strong feature with µτ = 10, where τ = E[S] is the mean reset interval. The solid curves show the retained fraction that ordinary ridge would produce on those two features as the dimensionless ridge strength λτ varies. The dashed horizontal lines show the retained fra… view at source ↗
Figure 3
Figure 3. Figure 3: Renewal law risk beyond the mean. Top row: risk decomposition at fixed mean reset interval τ as the reset-time law is varied from Poisson/exponential resetting (k = 1) through Erlang laws to the periodic limit (k → ∞). Solid curves show the exact renewal formulas after averaging over observation noise, and markers show empirical estimates from exact snapshots. Colors denote total risk (black), mean-estimat… view at source ↗
Figure 4
Figure 4. Figure 4: Validation-tuned non-exponential filters on the spiked covariance model. (a) Representative filter curves at γ = 1.5, with periodic and Erlang-3 filters each using its own validation￾tuned mean reset interval. The gray fill shows the Marchenko-Pastur continuous density, blue ticks mark positive sample spikes, gray ticks mark positive sample bulk eigenvalues, and the red dashed line marks the Marchenko-Past… view at source ↗
Figure 5
Figure 5. Figure 5: summarizes the resulting mechanism. Panel (a) plots the validation-tuned Gamma￾family filters on a representative split, while panel (b) shows the corresponding population eigenvalue distribution together with the signal-carrying modes. The true signal lives on a small subset of the spectrum, while nuisance modes occupy a broader lower-eigenvalue band. On the weak part of the spectrum, the fitted curves fo… view at source ↗
Figure 6
Figure 6. Figure 6: Scalar conditional-risk landscapes for individual observed datasets. Each panel fixes the realised data-dependent value b and varies the reset rate r. The conditional scalar risk is R(r) =  b µ+r − β0 2 + σ 2 n 2µ+r + rb2 (µ+r) 2(2µ+r) , decomposed into bias2 , SGD variance, and reset variance, together with empirical estimates from stationary simulations. The black star marks the theoretical minimiser o… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

46 extracted references · 2 canonical work pages · 1 internal anchor

  1. [1]

    Ridge regression: Biased estimation for nonorthogonal problems

    Arthur E Hoerl and Robert W Kennard. Ridge regression: Biased estimation for nonorthogonal problems. Technometrics, 12(1):55–67, 1970

  2. [2]

    Generalizations of mean square error applied to ridge regression

    Chris M Theobald. Generalizations of mean square error applied to ridge regression. Journal of the Royal Statistical Society Series B: Statistical Methodology , 36(1):103–106, 1974

  3. [3]

    Further results on the mean square error of ridge regression

    R W Farebrother. Further results on the mean square error of ridge regression. Journal of the Royal Statistical Society Series B: Statistical Methodology , 38(3):248–250, 1976

  4. [4]

    Anders Krogh and John A. Hertz. A simple weight decay can improve generalization. In Advances in Neural Information Processing Systems , volume 4, pages 950–957, 1991

  5. [5]

    Spectral methods for regularization in learning theory

    Lorenzo Rosasco, Ernesto De Vito, and Alessandro Verri. Spectral methods for regularization in learning theory. Technical Report DISI-TR-05-18, DISI, Università degli Studi di Genova, 2005

  6. [6]

    On regularization algorithms in learning theory

    Frank Bauer, Sergei Pereverzev, and Lorenzo Rosasco. On regularization algorithms in learning theory. Journal of complexity , 23(1):52–72, 2007

  7. [7]

    Lo Gerfo, Lorenzo Rosasco, Francesca Odone, Ernesto De Vito, and Alessandro Verri

    L. Lo Gerfo, Lorenzo Rosasco, Francesca Odone, Ernesto De Vito, and Alessandro Verri. Spectral algorithms for supervised learning. Neural Computation , 20(7):1873–1897, 2008

  8. [8]

    On early stopping in gradient descent learning

    Yuan Yao, Lorenzo Rosasco, and Andrea Caponnetto. On early stopping in gradient descent learning. Constructive approximation , 26(2):289–315, 2007

  9. [9]

    A continuous-time view of early stopping for least squares regression

    Alnur Ali, J Zico Kolter, and Ryan J Tibshirani. A continuous-time view of early stopping for least squares regression. In The 22nd international conference on artificial intelligence and statistics , pages 1370–1378. PMLR, 2019. 33

  10. [10]

    Optimal adaptation for early stopping in statistical inverse problems

    Gilles Blanchard, Marc Hoffmann, and Markus Reiß. Optimal adaptation for early stopping in statistical inverse problems. SIAM/ASA Journal on Uncertainty Quantification , 6(3):1043–1075, 2018

  11. [11]

    On regularization via early stopping for least squares regression

    Rishi Sonthalia, Jackie Lok, and Elizaveta Rebrova. On regularization via early stopping for least squares regression. arXiv preprint arXiv:2406.04425 , 2024

  12. [12]

    Kernel ridge vs

    Lee H Dicker, Dean P Foster, and Daniel Hsu. Kernel ridge vs. principal component regression: Minimax bounds and the qualification of regularization operators. Electronic Journal of Statistics, 11(1):1022–1047, 2017

  13. [13]

    Comparing classes of estima- tors: When does gradient descent beat ridge regression in linear models? arXiv preprint arXiv:2108.11872, 2021

    Dominic Richards, Edgar Dobriban, and Patrick Rebeschini. Comparing classes of estima- tors: When does gradient descent beat ridge regression in linear models? arXiv preprint arXiv:2108.11872, 2021

  14. [14]

    On the saturation effects of spectral algorithms in large dimensions

    Weihao Lu, Haobo Zhang, Yicheng Li, and Qian Lin. On the saturation effects of spectral algorithms in large dimensions. Advances in Neural Information Processing Systems , 37:7011– 7059, 2024

  15. [15]

    A differential equation for modeling nes- terov’s accelerated gradient method: Theory and insights

    Weijie Su, Stephen Boyd, and Emmanuel J Candès. A differential equation for modeling nes- terov’s accelerated gradient method: Theory and insights. Journal of Machine Learning Research, 17(153):1–43, 2016

  16. [16]

    Adaptive restart for accelerated gradient schemes

    Brendan O’Donoghue and Emmanuel Candès. Adaptive restart for accelerated gradient schemes. Foundations of computational mathematics , 15(3):715–732, 2015

  17. [17]

    Monotonicity and restart in fast gradient methods

    Pontus Giselsson and Stephen Boyd. Monotonicity and restart in fast gradient methods. In 53rd IEEE Conference on Decision and Control , pages 5058–5063. IEEE, 2014

  18. [18]

    Stochastic gradient descent as approx- imate bayesian inference

    Stephan Mandt, Matthew D Hoffman, and David M Blei. Stochastic gradient descent as approx- imate bayesian inference. Journal of Machine Learning Research , 18(134):1–35, 2017

  19. [19]

    Implicit regularization or implicit conditioning? exact risk trajectories of sgd in high dimensions

    Courtney Paquette, Elliot Paquette, Ben Adlam, and Jeffrey Pennington. Implicit regularization or implicit conditioning? exact risk trajectories of sgd in high dimensions. Advances in Neural Information Processing Systems , 35:35984–35999, 2022

  20. [20]

    The implicit regularization of stochastic gra- dient flow for least squares

    Alnur Ali, Edgar Dobriban, and Ryan Tibshirani. The implicit regularization of stochastic gra- dient flow for least squares. In International conference on machine learning , pages 233–244. PMLR, 2020

  21. [21]

    Foster, and Sham M

    Difan Zou, Jingfeng Wu, Vladimir Braverman, Quanquan Gu, Dean P. Foster, and Sham M. Kakade. The benefits of implicit regularization from sgd in least squares problems. In Advances in Neural Information Processing Systems , volume 34, pages 5456–5468, 2021

  22. [22]

    Implicit gradient regularization

    David GT Barrett and Benoit Dherin. Implicit gradient regularization. In International Confer- ence on Learning Representations , 2021

  23. [23]

    Implicit regularization for deep neural networks driven by an ornstein-uhlenbeck like process

    Guy Blanc, Neha Gupta, Gregory Valiant, and Paul Valiant. Implicit regularization for deep neural networks driven by an ornstein-uhlenbeck like process. In Proceedings of Thirty Third Conference on Learning Theory , volume 125 of Proceedings of Machine Learning Research, pages 483–513. PMLR, 2020

  24. [24]

    Diffusion with stochastic resetting

    Martin R Evans and Satya N Majumdar. Diffusion with stochastic resetting. Physical review letters, 106(16):160601, 2011

  25. [25]

    Stochastic resetting and applications

    Martin R Evans, Satya N Majumdar, and Grégory Schehr. Stochastic resetting and applications. Journal of Physics A: Mathematical and Theoretical , 53(19):193001, 2020. 34

  26. [26]

    First passage under restart

    Arnab Pal and Shlomi Reuveni. First passage under restart. Physical review letters, 118(3):030603, 2017

  27. [27]

    Universal performance bounds of restart

    Dmitry Starkov and Sergey Belan. Universal performance bounds of restart. Physical Review E , 107(6):L062101, 2023

  28. [28]

    Constructing efficient strategies for the process optimization by restart

    Ilia Nikitin and Sergey Belan. Constructing efficient strategies for the process optimization by restart. Physical Review E , 109(5):054117, 2024

  29. [29]

    Random search with resetting: a unified renewal approach

    Aleksei Chechkin and Igor M Sokolov. Random search with resetting: a unified renewal approach. Physical review letters , 121(5):050601, 2018

  30. [30]

    Keidar, Ofir Blumer, Barak Hirshberg, and Shlomi Reuveni

    Tommer D. Keidar, Ofir Blumer, Barak Hirshberg, and Shlomi Reuveni. Adaptive resetting for informed search strategies and the design of non-equilibrium steady-states. Nature Communica- tions, 16:7259, 2025

  31. [31]

    Sgdr: Stochastic gradient descent with warm restarts

    Ilya Loshchilov and Frank Hutter. Sgdr: Stochastic gradient descent with warm restarts. In International Conference on Learning Representations , 2017

  32. [32]

    The primacy bias in deep reinforcement learning

    Evgenii Nikishin, Max Schwarzer, Pierluca D’Oro, Pierre-Luc Bacon, and Aaron Courville. The primacy bias in deep reinforcement learning. In International conference on machine learning , pages 16828–16847. PMLR, 2022

  33. [33]

    Keidar, Shlomi Reuveni, and Barak Hirshberg

    Sagi Meir, Tommer D. Keidar, Shlomi Reuveni, and Barak Hirshberg. First-passage approach to optimizing perturbations for improved training of machine learning models. Machine Learning: Science and Technology, 6(2):025053, 2025

  34. [34]

    Stochastic resetting mitigates latent gradient bias of sgd from label noise

    Youngkyoung Bae, Yeongwoo Song, and Hawoong Jeong. Stochastic resetting mitigates latent gradient bias of sgd from label noise. Machine Learning: Science and Technology , 6(1):015062, 2025

  35. [35]

    Non-stationary learning of neural networks with automatic soft pa- rameter reset

    Alexandre Galashov, Michalis Titsias, András György, Clare Lyle, Razvan Pascanu, Yee Whye Teh, and Maneesh Sahani. Non-stationary learning of neural networks with automatic soft pa- rameter reset. Advances in Neural Information Processing Systems , 37:83197–83234, 2024

  36. [36]

    Gradient flow, laplace transforms, and infinitesimal steepest descent: Partial results and open directions

    Ryan J Tibshirani. Gradient flow, laplace transforms, and infinitesimal steepest descent: Partial results and open directions. In Mathematical Foundations of Robust and Generalizable Learning , volume 19 of Oberwolfach Reports, pages 2683–2684. EMS Press, 2022. Report No. 46/2022

  37. [37]

    Sheldon M. Ross. Introduction to Probability Models . Academic Press, 11th edition, 2014

  38. [38]

    Gallager

    Robert G. Gallager. Discrete Stochastic Processes , volume 321 of The Springer International Series in Engineering and Computer Science . Springer, New York, NY, 1996

  39. [39]

    Schilling, Renming Song, and Zoran Vondraček

    René L. Schilling, Renming Song, and Zoran Vondraček. Bernstein Functions: Theory and Applications. De Gruyter, 2nd edition, 2012

  40. [40]

    Optimal shrinkage of eigenvalues in the spiked covariance model

    David L Donoho, Matan Gavish, and Iain M Johnstone. Optimal shrinkage of eigenvalues in the spiked covariance model. Annals of statistics , 46(4):1742–1778, 2018

  41. [41]

    Bayesian interpolation

    David JC MacKay. Bayesian interpolation. Neural computation, 4(3):415–447, 1992

  42. [42]

    Pattern recognition and machine learning

    Christopher M Bishop. Pattern recognition and machine learning . Information Science and Statistics. Springer, 2006

  43. [43]

    An Introduction to Statis- tical Learning: with Applications in R , volume 103 of Springer Texts in Statistics

    Gareth James, Daniela Witten, Trevor Hastie, and Robert Tibshirani. An Introduction to Statis- tical Learning: with Applications in R , volume 103 of Springer Texts in Statistics . Springer, New York, NY, 2013. 35

  44. [44]

    The Elements of Statistical Learning: Data Mining, Inference, and Prediction

    Trevor Hastie, Robert Tibshirani, and Jerome Friedman. The Elements of Statistical Learning: Data Mining, Inference, and Prediction . Springer, 2009

  45. [45]

    Principal components regression in exploratory statistical research

    William F Massy. Principal components regression in exploratory statistical research. Journal of the American Statistical Association , 60(309):234–256, 1965

  46. [46]

    Truncated singular value decomposition solutions to discrete ill-posed prob- lems with ill-determined numerical rank

    Per Christian Hansen. Truncated singular value decomposition solutions to discrete ill-posed prob- lems with ill-determined numerical rank. SIAM Journal on Scientific and Statistical Computing , 11(3):503–518, 1990. 36