REVIEW 1 major objections 4 minor 231 references
The Fourth Quadrant: A Stylized View of Benign Misfitting
T0 review · 1 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read This paper claims that in a deterministic single-spike regression model, within the window $d/\gamma^2 \ll n \ll d/\gamma$, any span predictor with small test error must have large training error, and proves the exact tradeoff curve.
desk verdict An exact, honest stylized model that makes a real conceptual point; the central proofs check out and the caveats are about scope, not logic. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing identity is the scalar reduction of Lemma 3.1: for a span predictor $\vec u=\sum_i \alpha_i \vec x_i$, with $s=\sum_i\alpha_i$ and $q=\sum_i\alpha_i^2$, orthogonality gives $E_{\rm test}=(\gamma s-1)^2+dq$ and $E_{\rm train}=(\gamma s-1)^2+2ds(\gamma s-1)/n+d^2q/n$. The defining asymmetry is residue amplification: the nuisance quadratic enters training error with an extra factor $d/n$ compared with test error, so the same shared signal that helps fresh points forces per-training-point overshoot by a factor $\sim d/(\gamma n)=1/\rho$. The proof of the frontier completes a square in the signed training residual $h=(\gamma+d/n)s-1$; the one-pass SGD analysis uses the geometric
What would settle it
Run the deterministic model with $d=50{,}000$, $\gamma=100$, $n=100$: if the best span predictor's test error is not near $d/(\gamma^2 n)=0.05$ while its training error is not near $(d/(\gamma n))^2=25$, or the minimum-norm interpolator's test error is already small, the fourth-quadrant claim fails. Alternatively, in the random same-distribution model, if interpolation reaches a fixed target error at the same sample count as the best span predictor, the threshold separation fails.
Extended reading notes
Core claim
Within the deterministic orthogonal model of Section 3.1, the central discovery is an exact train–test Pareto frontier for predictors in the span of the training vectors. Every Pareto-optimal span predictor is a scalar multiple of the minimum-norm interpolator, with coefficient sum $s$ and energy $q=s^2/n$, and the frontier runs from the interpolator $(0,g_{\rm int})$ to the test-optimal point $(t_\star,g_{\rm min})$, where $g_{\rm int}=d(d+n)/(d+\gamma n)^2$ and $g_{\rm min}=d/(d+\gamma^2 n)$. The necessity result, Corollary 3.3, states that in the window $\rho=\gamma n/d\to0$, $R=\gamma^2 n/d\to\infty$, any span predictor with $E_{\rm test}\le \varepsilon$ must satisfy $E_{\rm train}\ge \m
Load-bearing premise
The exact-orthogonality design of equation (9)—training nuisance vectors mutually orthogonal with equal squared norm $d$—is the load-bearing premise; the paper's own Section 7.1 shows that without it the first-useful threshold becomes spectrum-dependent, and the announced same-distribution companion is not proved in this preprint.
Editorial extensions
If this is right
- Interpolation is not a safe stopping rule in this regime: fitting every label requires the spike coefficient $\rho/(1+\rho)$, which under-calibrates the shared signal and leaves test error near baseline until $n\gg d/\gamma$.
- Sample-efficient use of the span requires overshooting, equivalently a negative ridge penalty $-\frac{d}{n}(1-1/\gamma)$; in this model that is an exact algebraic statement, not a numerical artifact.
- One-pass SGD with a constant large learning rate in the window $1/(\gamma n)\ll\eta\ll\gamma/d$ has $E_{\rm test}\to0$ and $E_{\rm train}\to\infty$; the tuned rate $\eta=\log(\gamma^2 n/d)/(2\gamma n)$ matches the best span predictor up to a factor $\frac14\log(\gamma^2 n/d)$.
- The clean-test threshold $n_{\rm clean}\asymp d/\gamma^2$ precedes the low-adversarial-sensitivity threshold $n_{\rm adv}\asymp \delta^2 d^2/(\gamma^2\tau^2)$: a predictor can look accurate on random test points while remaining vulnerable to label-preserving perturbations aligned with its nuisance residue.
- Negative generalization gaps are forced by the geometry of the span in this window, so a method that treats low training error as evidence of good generalization is using the wrong diagnostic.
Reading between the lines
- If the announced same-distribution Gaussian companion behaves as claimed, the qualitatively similar separation should be visible in ordinary Gaussian designs: one can test whether the best ridge or min-norm predictor's training residual exceeds the zero-predictor baseline in the interval between $d/\gamma^2$ and $d/\gamma$.
- The fresh-versus-reused stability contrast suggests a practical diagnostic for real models: compare the loss on never-seen fresh examples with the loss on reused training examples at the same optimizer step; a first epoch where fresh loss falls while reused loss spikes is the algorithmic signature of misfitting.
- The participation accounting of Appendix B.5 gives a testable reinterpretation of linear mini-batch scaling: batch size should be scaled so that per-example learning rate and effective participation are preserved, rather than to match gradient-noise temperature.
- One could extend the necessity to classification by replacing squared error with a margin-based loss: the same span argument predicts that early good classifiers may have worse-than-baseline empirical margin on their own training points, mirroring the regression overshoot.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies a deterministic (d+1)-dimensional single-spike linear regression model in which n training vectors share a spike coordinate of amplitude sqrt(gamma), gamma>1, and have mutually orthogonal, equal-norm nuisance components; all training labels are 1, and test points are Gaussian with covariance diag(gamma,1,...,1). Restricting to predictors in the training span, the paper derives an exact train-test Pareto frontier (Lemma 3.1, Theorem 3.2), identifies a 'benign misfitting' window d/gamma^2 << n << d/gamma in which every span predictor with small test error must have training error above the zero-predictor baseline and in fact diverging (Corollary 3.3), proves that one-pass SGD with a suitably large constant learning rate reaches this region within a logarithmic factor of the frontier (Theorem 4.3, Proposition 4.4), and connects the nuisance residue to adversarial sensitivity (Theorem 5.1, Appendix C). Appendices provide stability boundaries, learning-rate robustness, an explicit step-by-step training-error ascent, a nonlinear median construction escaping the span obstruction, and comparisons with stochastic convex optimization.
Significance. If the result stands, this is a clean and exact algebraic demonstration that a negative generalization gap can be forced by span geometry rather than by noise, augmentation, or regularization. The paper gives explicit finite-n formulas, has an unusually clear statement of its own scope, and ships the full derivations for the main theorems; the deterministic caricature makes the mechanism auditable. Strengths include: the exact scalar reduction to (s,q), the closed-form frontier and sample-complexity separation, the geometric one-pass SGD recursion, and the honest labeling of what is conjectural (Section 7.1 random orientations) or deferred to companion papers. The stress-test concern about exact orthogonality is a scope limitation, not an internal inconsistency: the paper repeatedly states that necessity claims concern the training span in the fixed orthogonal design, and Section 7.1 explicitly says the random setting can shift the threshold. The central deterministic claims are internally consistent.
major comments (1)
- [§4.2, Remark 4.5] The remark asserts that optimizing over stable constant learning rates 'cannot' eliminate the logarithmic gap as long as log R = o(n), gives a heuristic three-case argument, and then states 'The full calculation, including the sharp leading constant, is omitted here.' This is a theorem-like optimality claim with no proof supplied. It is used to attribute the gap to a participation deficit and to justify the title of Subsection B.4. Please either provide the proof in an appendix or explicitly downgrade the statement to a conjecture/qualified heuristic. The main reachability result, Theorem 4.3, does not depend on this optimality claim, so the paper's central conclusion is unaffected.
minor comments (4)
- [After Eq. (77)] The sentence saying both risks are 'a factor 1/4 log R (1+o(1))' from the frontier endpoints is ambiguous. For Etest the ratio to gmin is 1 + (1/4)log R (1+o(1)), while for Etrain the ratio to t* is (1/4)log R (1+o(1)); the wording suggests the same multiplicative factor on both axes. Please rephrase precisely.
- [§7.1, Non-flat nuisance tails] The paragraph asserts, without proof, that under covariance-weighted approximate isotropy the first-useful-span threshold becomes n >> tr(Omega^2)/gamma^2 and that non-flatness narrows the window. The random-orientation version is honestly labeled a conjecture, but these weighted-isotropy claims are stated as facts. Please add a short derivation or label them as conjectural; they are outside the main deterministic theorem but should not be presented as proven.
- [§6 and Introduction] The persistence claims for the same-distribution companion [160], multispike [158], and Oja [159] settings are announced but not proved in this paper. This is acceptable if they are clearly identified as results of separate preprints; consider adding an explicit sentence in Section 6 that the current paper does not prove those extensions.
- [General notation] The same symbol r is used for spike calibration (33) and for the effective-rank family r_k in Section 7.1, while the same symbol R is used for the signal-to-nuisance ratio and the effective-rank family R_k. The paper notes this, but a visual distinction (e.g., calligraphic letters for the effective ranks) would reduce reader confusion, especially in a paper with as many symbols as this one.
Circularity Check
No significant circularity: the claimed frontier and one-pass SGD reachability are derived from the stated orthogonal geometry, not assumed or fitted.
full rationale
The paper's central claims are self-contained algebraic consequences of its explicitly stated deterministic model. Lemma 3.1 reduces the two risks to scalars s and q using only orthogonality (9) and the definitions (13)-(14); the proof is direct and does not presuppose the fourth-quadrant conclusion. Theorem 3.2 derives the exact Pareto frontier by completing the square and comparing branches with γ>1; Corollary 3.3 uses only E_test ≥ (r−1)^2 and q ≥ s^2/n, so the forced misfit bound follows for every span predictor, not just the constructed optimum. The one-pass SGD result is also not fitted: Lemma 4.1 derives the geometric coefficient profile exactly from the fresh-example orthogonality identity (69), and Theorem 4.3/Proposition 4.4 give finite-n expressions from which the chosen learning rate η_pass arises as an interior point of a broad analytical window (113), with a whole interval of successful rates certified rather than a single tuned value matching a target. Self-citations to companion papers [158,159,160] appear only as context, terminology, or announced future work; the same-distribution companion is explicitly stated to be announced rather than proved here, and no load-bearing step cites it. The paper's own limitations—e.g., the unproved 'cannot eliminate the logarithmic gap' remark and the spectrum-dependence discussion in Section 7.1—are openly flagged and do not smuggle in the main result. Nothing in the derivation chain reduces by construction to its inputs.
Assumptions & free parameters
free parameters (1)
- one-pass SGD learning rate eta_pass =
log(gamma^2 n/d) / (2 gamma n)
assumptions (6)
- domain assumption Training nuisance vectors are exactly orthogonal with equal squared norm d: v_i^T v_j = 0 for i != j and ||v_i||_2^2 = d (Section 3.1, equation (9)).
- domain assumption All training labels are 1 and each training vector has spike coordinate sqrt(gamma) (Section 3.1, equations (10)-(11)).
- domain assumption The test law is x_test ~ N(0, diag(gamma,1,...,1)) with y_test = x_test[1]/sqrt(gamma) (Section 3.1, equation (12)).
- domain assumption gamma > 1 and n <= d (standing assumptions, Section 3.1).
- standard math Standard linear algebra, Cauchy-Schwarz, and Gaussian moments are used freely in Lemma 3.1, Theorem 3.2, and Lemma 4.1.
- domain assumption For the adversarial claim, the adversary is restricted to label-preserving perturbations orthogonal to the spike direction (Section 5.1).
Cite this review
Pith. "Pith review of The Fourth Quadrant: A Stylized View of Benign Misfitting." pith.science (2026). https://pith.science/paper/YMAAKVYJ
@misc{pith2026260801032,
author = {Pith},
title = {Pith review of: The Fourth Quadrant: A Stylized View of Benign Misfitting},
year = {2026},
howpublished = {\url{https://pith.science/paper/YMAAKVYJ}},
note = {Machine review of arXiv:2608.01032}
}
abstract
Training error is what we can observe on a training set; test error is the quantity we actually care about. We study linear regression with squared-error in a deterministic $(d+1)$-dimensional single-spike model. Each stylized training vector has the same informative spike coordinate, of amplitude $\sqrt{\gamma}$ with $\gamma>1$. The remaining directions are nuisance, and the nuisance components of distinct training vectors all have equal norm and are mutually orthogonal. The training labels are all $1$. Fresh test points are drawn from $\vec{x}_{\rm test} \sim \mathcal{N}(\vec{0},\operatorname{diag}(\gamma,1,\ldots,1))$, with the noise-free test labels being the normalized spike coordinate $x_{\rm test}[1]/\sqrt{\gamma}$. We focus on linear predictors in the span of the training vectors, the class naturally reached by zero-initialized linear gradient methods. We exhibit a range of training-set sizes $n$ in which every span predictor that generalizes well must fit the training data \emph{worse} than the zero predictor. We call this regime \emph{benign misfitting}, or the fourth quadrant. The best span predictor begins to generalize when $n\gg d/\gamma^2$, while interpolation does not generalize until the later threshold $n\gg d/\gamma$. In the window $d/\gamma^2 \ll n \ll d/\gamma$, useful prediction within the linear span lies beyond interpolation: predictions on the training points overshoot the labels. We show that one-pass stochastic gradient descent (SGD), with a large constant learning rate, reaches small test error throughout this window---matching the best span predictor up to a logarithmic factor. We also verify directly that it indeed has \emph{large} empirical training error (despite the descent premise in its name). Finally, we show that the unavoidable nuisance component responsible for the training misfit also controls the predictor's adversarial sensitivity.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
Brendan McMahan, Ilya Mironov, Kunal Talwar, and Li Zhang
Martin Abadi, Andy Chu, Ian Goodfellow, H. Brendan McMahan, Ilya Mironov, Kunal Talwar, and Li Zhang. Deep learning with differential privacy. InProceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security, pages 308–318, 2016
2016
-
[2]
The merged-staircase property: A necessary and nearly sufficient condition for SGD learning of sparse functions on two-layer neural networks
Emmanuel Abbe, Enric Boix Adser` a, and Theodor Misiakiewicz. The merged-staircase property: A necessary and nearly sufficient condition for SGD learning of sparse functions on two-layer neural networks. InProceedings of Thirty Fifth Conference on Learning Theory, volume 178 ofProceedings of Machine Learning Research, pages 4782–4887. PMLR, 2022. URLhttps...
2022
-
[3]
SGD learning on neural networks: Leap complexity and saddle-to-saddle dynamics
Emmanuel Abbe, Enric Boix Adser` a, and Theodor Misiakiewicz. SGD learning on neural networks: Leap complexity and saddle-to-saddle dynamics. InProceedings of Thirty Sixth Conference on Learning Theory, volume 195 ofProceedings of Machine Learning Research, pages 2552–2623. PMLR, 2023. URLhttps://proceedings.mlr.press/v195/abbe23a.html
2023
-
[4]
Ishaq Aden-Ali, Mikael Møller Høgsgaard, Kasper Green Larsen, and Nikita Zhivotovskiy. Majority-of-three: The simplest optimal learner? InProceedings of the Thirty Seventh Conference on Learning Theory, volume 247 ofProceedings of Machine Learning Research, pages 22–45. PMLR, 2024
2024
-
[5]
Madhu S. Advani and Andrew M. Saxe. High-dimensional dynamics of generalization error in neural networks.Neural Networks, 132:428–446, 2020. doi: 10.1016/j.neunet.2020.08.022
-
[6]
Understanding the unstable convergence of gradient descent
Kwangjun Ahn, Jingzhao Zhang, and Suvrit Sra. Understanding the unstable convergence of gradient descent. InProceedings of the 39th International Conference on Machine Learning, volume 162 ofProceedings of Machine Learning Research, pages 247–257. PMLR, 2022. URL https://proceedings.mlr.press/v162/ahn22a.html
2022
-
[7]
Zico Kolter, and Ryan J
Alnur Ali, J. Zico Kolter, and Ryan J. Tibshirani. A continuous-time view of early stopping for least squares regression. InProceedings of the Twenty-Second International Conference on Artificial Intelligence and Statistics, volume 89 ofProceedings of Machine Learning Research, pages 1370–1378. PMLR, 2019. URLhttps://proceedings.mlr.press/v89/ali19a.html
2019
-
[8]
The implicit regularization of stochastic gradient flow for least squares
Alnur Ali, Edgar Dobriban, and Ryan Tibshirani. The implicit regularization of stochastic gradient flow for least squares. InProceedings of the 37th International Conference on Machine Learning, volume 119 ofProceedings of Machine Learning Research, pages 233–244. PMLR,
Show all 231 references
-
[9]
The space complexity of approximating the frequency moments.Journal of Computer and System Sciences, 58(1):137–147, 1999
Noga Alon, Yossi Matias, and Mario Szegedy. The space complexity of approximating the frequency moments.Journal of Computer and System Sciences, 58(1):137–147, 1999
1999
-
[10]
Never go full batch (in stochastic convex optimization)
Idan Amir, Yair Carmon, Tomer Koren, and Roi Livni. Never go full batch (in stochastic convex optimization). InAdvances in Neural Information Processing Systems, volume 34, 2021
2021
-
[11]
SGD generalizes better than GD (and regularization doesn’t help)
Idan Amir, Tomer Koren, and Roi Livni. SGD generalizes better than GD (and regularization doesn’t help). InProceedings of the Thirty Fourth Conference on Learning Theory, volume 134 ofProceedings of Machine Learning Research, pages 63–92. PMLR, 2021. 61
2021
-
[12]
Thinking outside the ball: Optimal learning with gradient descent for generalized linear stochastic convex optimization
Idan Amir, Roi Livni, and Nathan Srebro. Thinking outside the ball: Optimal learning with gradient descent for generalized linear stochastic convex optimization. InAdvances in Neural Information Processing Systems, volume 35, 2022
2022
-
[13]
Edge of stochastic stability: Revisiting the edge of stability for SGD, 2024
Arseniy Andreyev and Pierfrancesco Beneventano. Edge of stochastic stability: Revisiting the edge of stability for SGD, 2024. URLhttps://arxiv.org/abs/2412.20553
2024
-
[14]
Understanding gradient descent on the edge of stability in deep learning
Sanjeev Arora, Zhiyuan Li, and Abhishek Panigrahi. Understanding gradient descent on the edge of stability in deep learning. InInternational Conference on Machine Learning, pages 948–1024, 2022
2022
-
[15]
Idan Attias, Gintare Karolina Dziugaite, Mahdi Haghifam, Roi Livni, and Daniel M. Roy. Information complexity of stochastic convex optimization: Applications to generalization, memorization, and tracing. InProceedings of the 41st International Conference on Machine Learning, 2...
2024 arXiv
-
[16]
Robust linear least squares regression.The Annals of Statistics, 39(5):2766–2794, 2011
Jean-Yves Audibert and Olivier Catoni. Robust linear least squares regression.The Annals of Statistics, 39(5):2766–2794, 2011
2011
-
[17]
Diggavi, and David N
Amir Salman Avestimehr, Suhas N. Diggavi, and David N. C. Tse. Wireless network infor- mation flow: A deterministic approach.IEEE Transactions on Information Theory, 57(4): 1872–1905, 2011. doi: 10.1109/TIT.2011.2110110
1905
-
[18]
Phase transition of the largest eigenvalue for nonnull complex sample covariance matrices.The Annals of Probability, 33(5):1643–1697,
Jinho Baik, G´ erard Ben Arous, and Sandrine P´ ech´ e. Phase transition of the largest eigenvalue for nonnull complex sample covariance matrices.The Annals of Probability, 33(5):1643–1697,
-
[19]
David G. T. Barrett and Benoit Dherin. Implicit gradient regularization. InInternational Conference on Learning Representations, 2021. URL https://openreview.net/forum?id= 3q5IqUrkcF
2021
-
[20]
Bartlett, Philip M
Peter L. Bartlett, Philip M. Long, G´ abor Lugosi, and Alexander Tsigler. Benign overfitting in linear regression.Proceedings of the National Academy of Sciences, 117(48):30063–30070,
-
[21]
Fit without fear: Remarkable mathematical phenomena of deep learn- ing through the prism of interpolation.Acta Numerica, 30:203–248, 2021
Mikhail Belkin. Fit without fear: Remarkable mathematical phenomena of deep learn- ing through the prism of interpolation.Acta Numerica, 30:203–248, 2021. doi: 10.1017/ S0962492921000039
2021
-
[22]
Hsu, and Partha P
Mikhail Belkin, Daniel J. Hsu, and Partha P. Mitra. Overfitting or perfect fitting? risk bounds for classification and regression rules that interpolate. InAdvances in Neural Information Processing Systems 31, pages 2300–2311, 2018. URL https://papers.nips.cc/paper/7498
2018
-
[23]
doi: 10.1073/pnas.1907378117
-
[24]
Reconciling modern machine- learning practice and the classical bias–variance trade-off.Proceedings of the National Academy of Sciences, 116(32):15849–15854, 2019
Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal. Reconciling modern machine- learning practice and the classical bias–variance trade-off.Proceedings of the National Academy of Sciences, 116(32):15849–15854, 2019. doi: 10.1073/pnas.1903070116. 62
2019 doi
-
[25]
Tsybakov
Mikhail Belkin, Alexander Rakhlin, and Alexandre B. Tsybakov. Does data interpolation contradict statistical optimality? InProceedings of the Twenty-Second International Conference on Artificial Intelligence and Statistics, volume 89 ofProceedings of Machine Learning Research,...
2019
-
[26]
To understand deep learning we need to understand kernel learning
Mikhail Belkin, Siyuan Ma, and Soumik Mandal. To understand deep learning we need to understand kernel learning. InProceedings of the 35th International Conference on Machine Learning, volume 80 ofProceedings of Machine Learning Research, pages 541–549. PMLR,
-
[27]
Bertsekas
Dimitri P. Bertsekas. Incremental gradient, subgradient, and proximal methods for convex optimization: A survey. In Suvrit Sra, Sebastian Nowozin, and Stephen J. Wright, editors, Optimization for Machine Learning, pages 85–120. MIT Press, 2011. doi: 10.7551/mitpress/ 8996.003....
2011 arXiv
-
[28]
Learning single-index models with shallow neural networks
Alberto Bietti, Joan Bruna, Clayton Sanford, and Min Jae Song. Learning single-index models with shallow neural networks. InAdvances in Neural Information Processing Systems 35, pages 9768–9783, 2022. URL https://papers.nips.cc/paper_files/paper/2022/hash/ 3fb6c52aeb11e09053c1...
2022
-
[29]
Implicit regularization for deep neural networks driven by an Ornstein–Uhlenbeck like process
Guy Blanc, Neha Gupta, Gregory Valiant, and Paul Valiant. Implicit regularization for deep neural networks driven by an Ornstein–Uhlenbeck like process. InProceedings of the Thirty Third Conference on Learning Theory, volume 125 ofProceedings of Machine Learning Research, page...
2020
-
[30]
Two models of double descent for weak features.SIAM Journal on Mathematics of Data Science, 2(4):1167–1180, 2020
Mikhail Belkin, Daniel Hsu, and Ji Xu. Two models of double descent for weak features.SIAM Journal on Mathematics of Data Science, 2(4):1167–1180, 2020. doi: 10.1137/20M1336072
2020 doi
-
[31]
The tradeoffs of large scale learning
L´ eon Bottou and Olivier Bousquet. The tradeoffs of large scale learning. InAdvances in Neural Information Processing Systems, volume 20, 2007
2007
-
[32]
Curtis, and Jorge Nocedal
L´ eon Bottou, Frank E. Curtis, and Jorge Nocedal. Optimization methods for large-scale machine learning.SIAM Review, 60(2):223–311, 2018. doi: 10.1137/16M1080173
2018 doi
-
[33]
Stability and generalization.Journal of Machine Learning Research, 2:499–526, 2002
Olivier Bousquet and Andr´ e Elisseeff. Stability and generalization.Journal of Machine Learning Research, 2:499–526, 2002
2002
-
[34]
Euclidean information theory
Shashi Borade and Lizhong Zheng. Euclidean information theory. InInternational Zurich Seminar on Communications, pages 14–17, 2008
2008
-
[35]
Bagging predictors.Machine Learning, 24(2):123–140, 1996
Leo Breiman. Bagging predictors.Machine Learning, 24(2):123–140, 1996
1996
-
[36]
Gavin Brown, Mark Bun, Vitaly Feldman, Adam Smith, and Kunal Talwar. When is memorization of irrelevant training data necessary for high-accuracy learning? InProceedings of the 53rd Annual ACM SIGACT Symposium on Theory of Computing, STOC 2021, pages 123–132, 2021. doi: 10.114...
2021
-
[37]
Survey on algorithms for multi-index models.Statistical Science, 40(3):378–391, 2025
Joan Bruna and Daniel Hsu. Survey on algorithms for multi-index models.Statistical Science, 40(3):378–391, 2025. 63
2025
-
[38]
Proper learning, Helly number, and an optimal SVM bound
Olivier Bousquet, Steve Hanneke, Shay Moran, and Nikita Zhivotovskiy. Proper learning, Helly number, and an optimal SVM bound. InProceedings of the Thirty Third Conference on Learning Theory, volume 125 ofProceedings of Machine Learning Research, pages 582–609. PMLR, 2020
2020
-
[39]
A universal law of robustness via isoperimetry
S´ ebastien Bubeck and Mark Sellke. A universal law of robustness via isoperimetry. InAdvances in Neural Information Processing Systems, volume 34, 2021
2021
-
[40]
All ERMs can fail in stochastic convex optimization lower bounds in linear dimension, 2026
Tal Burla and Roi Livni. All ERMs can fail in stochastic convex optimization lower bounds in linear dimension, 2026
2026
-
[41]
The secret sharer: Evaluating and testing unintended memorization in neural networks
Nicholas Carlini, Chang Liu, ´Ulfar Erlingsson, Jernej Kos, and Dawn Song. The secret sharer: Evaluating and testing unintended memorization in neural networks. In28th USENIX Security Symposium, pages 267–284, 2019
2019
-
[42]
Veeravalli
Yuheng Bu, Shaofeng Zou, and Venugopal V. Veeravalli. Tightening mutual information-based bounds on generalization error.IEEE Journal on Selected Areas in Information Theory, 1(1): 121–130, 2020. doi: 10.1109/JSAIT.2020.2991139
2020
-
[43]
Membership inference attacks from first principles
Nicholas Carlini, Steve Chien, Milad Nasr, Shuang Song, Andreas Terzis, and Florian Tram` er. Membership inference attacks from first principles. InIEEE Symposium on Security and Privacy (SP), pages 1897–1914, 2022. doi: 10.1109/SP46214.2022.9833649
1914
-
[44]
The sample complexity of ERMs in stochastic convex optimization
Daniel Carmon, Roi Livni, and Amir Yehudayoff. The sample complexity of ERMs in stochastic convex optimization. InProceedings of the 27th International Conference on Artificial Intelligence and Statistics, volume 238 ofProceedings of Machine Learning Research, pages 3799–3807....
2024
-
[45]
Kamalika Chaudhuri, Claire Monteleoni, and Anand D. Sarwate. Differentially private empirical risk minimization.Journal of Machine Learning Research, 12:1069–1109, 2011
2011
-
[46]
Extracting training data from large language models
Nicholas Carlini, Florian Tram` er, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Kather- ine Lee, Adam Roberts, Tom Brown, Dawn Song, ´Ulfar Erlingsson, Alina Oprea, and Colin Raffel. Extracting training data from large language models. In30th USENIX Security Symposium, 2021
2021
-
[47]
Understanding gradient clipping in private SGD: A geometric perspective
Xiangyi Chen, Steven Z Wu, and Mingyi Hong. Understanding gradient clipping in private SGD: A geometric perspective. InAdvances in Neural Information Processing Systems, volume 33, 2020
2020
-
[48]
From one-pass SGD to data reuse: Mini-batch scaling laws in sketched linear regression, 2026
Ziyan Chen, Zhongzhu Zhou, and Ding-Xuan Zhou. From one-pass SGD to data reuse: Mini-batch scaling laws in sketched linear regression, 2026. arXiv:2605.24316
2026 arXiv
-
[49]
On the robustness of minimum norm interpolators and regularized empirical risk minimizers.The Annals of Statistics, 50(4): 2306–2333, 2022
Geoffrey Chinot, Matthias L¨ offler, and Sara van de Geer. On the robustness of minimum norm interpolators and regularized empirical risk minimizers.The Annals of Statistics, 50(4): 2306–2333, 2022. doi: 10.1214/22-AOS2190
2022 doi
-
[50]
Sarwate, and Kaushik Sinha
Kamalika Chaudhuri, Anand D. Sarwate, and Kaushik Sinha. Near-optimal algorithms for differentially-private principal components. InAdvances in Neural Information Processing Systems, volume 26, 2013
2013
-
[51]
Cohen, Simran Kaur, Yuanzhi Li, J
Jeremy M. Cohen, Simran Kaur, Yuanzhi Li, J. Zico Kolter, and Ameet Talwalkar. Gradient descent on neural networks typically occurs at the edge of stability. InInternational Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=jh-rTtvkGeM. 64
2021
-
[52]
High-dimensional limit of one-pass SGD on least squares.Electronic Communications in Probability, 29, 2024
Elizabeth Collins-Woodfin and Elliot Paquette. High-dimensional limit of one-pass SGD on least squares.Electronic Communications in Probability, 29, 2024. doi: 10.1214/23-ECP571. arXiv:2304.06847
2024 arXiv
-
[53]
Cubuk, Barret Zoph, Dandelion Mane, Vijay Vasudevan, and Quoc V
Ekin D. Cubuk, Barret Zoph, Dandelion Mane, Vijay Vasudevan, and Quoc V. Le. Au- toAugment: Learning augmentation strategies from data. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019
2019
-
[54]
Choquette-Choo, Florian Tram` er, Nicholas Carlini, and Nicolas Papernot
Christopher A. Choquette-Choo, Florian Tram` er, Nicholas Carlini, and Nicolas Papernot. Label-only membership inference attacks. InInternational Conference on Machine Learning, volume 139 ofProceedings of Machine Learning Research, pages 1964–1974, 2021
1964
-
[55]
Baraniuk
Yehuda Dar, Vidya Muthukumar, and Richard G. Baraniuk. A farewell to the bias-variance tradeoff? an overview of the theory of overparameterized machine learning, 2021. URL https://arxiv.org/abs/2109.02355
2021 arXiv
-
[56]
Oliveira
Luc Devroye, Matthieu Lerasle, G´ abor Lugosi, and Roberto I. Oliveira. Sub-Gaussian mean estimators.The Annals of Statistics, 44(6):2695–2725, 2016
2016
-
[57]
Sharp minima can generalize for deep nets
Laurent Dinh, Razvan Pascanu, Samy Bengio, and Yoshua Bengio. Sharp minima can generalize for deep nets. InProceedings of the 34th International Conference on Machine Learning, volume 70 ofProceedings of Machine Learning Research, pages 1019–1028. PMLR,
-
[58]
Self-stabilization: The implicit bias of gradient descent at the edge of stability
Alex Damian, Eshaan Nichani, and Jason D Lee. Self-stabilization: The implicit bias of gradient descent at the edge of stability. InInternational Conference on Learning Representations, 2023
2023
-
[59]
Robust linear regression: Phase-transitions and precise tradeoffs for general norms.arXiv preprint arXiv:2308.00556, 2023
Elvis Dohmatob and Meyer Scetbon. Robust linear regression: Phase-transitions and precise tradeoffs for general norms.arXiv preprint arXiv:2308.00556, 2023
2023 arXiv
-
[60]
Draper and R
Norman R. Draper and R. Craig Van Nostrand. Ridge regression and james–stein estimation: Review and comments.Technometrics, 21(4):451–466, 1979. doi: 10.1080/00401706.1979. 10489815
1979
-
[61]
Learning single-index models in gaussian space
Rishabh Dudeja and Daniel Hsu. Learning single-index models in gaussian space. InProceedings of the 31st Conference on Learning Theory, volume 75 ofProceedings of Machine Learning Research, pages 1887–1930. PMLR, 2018. URL https://proceedings.mlr.press/v75/ dudeja18a.html
1930
-
[62]
The algorithmic foundations of differential privacy
Cynthia Dwork and Aaron Roth. The algorithmic foundations of differential privacy. Foundations and Trends in Theoretical Computer Science, 9(3–4):211–407, 2014. doi: 10.1561/0400000042
2014 doi
-
[63]
Generalized no free lunch theorem for adversarial robustness
Elvis Dohmatob. Generalized no free lunch theorem for adversarial robustness. InProceedings of the 36th International Conference on Machine Learning, volume 97 ofProceedings of Machine Learning Research, pages 1646–1654, 2019
2019
-
[64]
Analyze gauss: Optimal bounds for privacy-preserving principal component analysis
Cynthia Dwork, Kunal Talwar, Abhradeep Thakurta, and Li Zhang. Analyze gauss: Optimal bounds for privacy-preserving principal component analysis. InProceedings of the 46th Annual ACM Symposium on Theory of Computing, pages 11–20, 2014. 65
2014
-
[65]
James–stein estimation and ridge regression
Bradley Efron and Trevor Hastie. James–stein estimation and ridge regression. InComputer Age Statistical Inference, pages 91–107. Cambridge University Press, 2016. doi: 10.1017/ CBO9781316576533.008
2016
-
[66]
Stein’s paradox in statistics.Scientific American, 236(5): 119–127, 1977
Bradley Efron and Carl Morris. Stein’s paradox in statistics.Scientific American, 236(5): 119–127, 1977. doi: 10.1038/scientificamerican0577-119
1977 doi
-
[67]
Etkin, David N
Raul H. Etkin, David N. C. Tse, and Hua Wang. Gaussian interference channel capacity to within one bit.IEEE Transactions on Information Theory, 54(12):5534–5562, 2008. doi: 10.1109/TIT.2008.2006447
2008
-
[68]
Calibrating noise to sensitivity in private data analysis
Cynthia Dwork, Frank McSherry, Kobbi Nissim, and Adam Smith. Calibrating noise to sensitivity in private data analysis. InTheory of Cryptography Conference, pages 265–284. Springer, 2006
2006
-
[69]
Analysis of classifiers’ robustness to adversarial perturbations.Machine Learning, 107:481–508, 2018
Alhussein Fawzi, Omar Fawzi, and Pascal Frossard. Analysis of classifiers’ robustness to adversarial perturbations.Machine Learning, 107:481–508, 2018
2018
-
[70]
Generalization of ERM in stochastic convex optimization: The dimension strikes back
Vitaly Feldman. Generalization of ERM in stochastic convex optimization: The dimension strikes back. InAdvances in Neural Information Processing Systems, volume 29, 2016
2016
-
[71]
Does learning require memorization? a short tale about a long tail
Vitaly Feldman. Does learning require memorization? a short tale about a long tail. In Proceedings of the 52nd Annual ACM SIGACT Symposium on Theory of Computing, STOC 2020, pages 954–959, 2020. doi: 10.1145/3357713.3384290
2020
-
[72]
Figiel, J
T. Figiel, J. Lindenstrauss, and V. D. Milman. The dimension of almost spherical sections of convex bodies.Acta Mathematica, 139(1–2):53–94, 1977
1977
-
[73]
(S)GD over diagonal linear networks: Implicit bias, large stepsizes and edge of stability
Mathieu Even, Scott Pesme, Suriya Gunasekar, and Nicolas Flammarion. (S)GD over diagonal linear networks: Implicit bias, large stepsizes and edge of stability. InAdvances in Neural Information Processing Systems 36, 2023. URL https://openreview.net/forum? id=uAyElhYKxg
2023
-
[74]
Shrinkage to infinity: Reducing test error by inflating the minimum norm interpolator in linear models, 2025
Jake Freeman. Shrinkage to infinity: Reducing test error by inflating the minimum norm interpolator in linear models, 2025. arXiv:2510.19206; v2 of April 30, 2026
2025 arXiv
-
[75]
Bartlett, and Nati Srebro
Spencer Frei, Gal Vardi, Peter L. Bartlett, and Nati Srebro. The double-edged sword of implicit bias: Generalization vs. robustness in ReLU networks. InAdvances in Neural Information Pro- cessing Systems, volume 36, 2023. URL https://proceedings.neurips.cc/paper_files/ paper/2...
2023
-
[76]
Gallager.Information Theory and Reliable Communication
Robert G. Gallager.Information Theory and Reliable Communication. John Wiley & Sons, New York, 1968
1968
-
[77]
Escaping from saddle points—online stochastic gradient for tensor decomposition
Rong Ge, Furong Huang, Chi Jin, and Yang Yuan. Escaping from saddle points—online stochastic gradient for tensor decomposition. InProceedings of the 28th Conference on Learning Theory, pages 797–842, 2015
2015
-
[78]
Foster, Satyen Kale, Haipeng Luo, Mehryar Mohri, and Karthik Sridharan
Dylan J. Foster, Satyen Kale, Haipeng Luo, Mehryar Mohri, and Karthik Sridharan. Logistic regression: The importance of being improper. InProceedings of the 31st Conference on Learning Theory, volume 75 ofProceedings of Machine Learning Research, pages 167–208. PMLR, 2018
2018
-
[79]
A loss curvature perspective on training instabilities of deep learning models
Justin Gilmer, Behrooz Ghorbani, Ankush Garg, Sneha Kudugunta, Behnam Neyshabur, David Cardoze, George Edward Dahl, Zachary Nado, and Orhan Firat. A loss curvature perspective on training instabilities of deep learning models. InInternational Conference on Learning Representat...
2022
-
[80]
Goodfellow, Jonathon Shlens, and Christian Szegedy
Ian J. Goodfellow, Jonathon Shlens, and Christian Szegedy. Explaining and harnessing adversarial examples. InInternational Conference on Learning Representations, 2015
2015
-
[81]
Accurate, large minibatch SGD: Training ImageNet in 1 hour, 2017
Priya Goyal, Piotr Doll´ ar, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He. Accurate, large minibatch SGD: Training ImageNet in 1 hour, 2017. URLhttps://arxiv.org/abs/1706.02677
2017 arXiv
-
[82]
Characterizing implicit bias in terms of optimization geometry
Suriya Gunasekar, Jason Lee, Daniel Soudry, and Nathan Srebro. Characterizing implicit bias in terms of optimization geometry. InProceedings of the 35th International Conference on Machine Learning, volume 80 ofProceedings of Machine Learning Research, pages 1832–1841. PMLR, 2...
2018
-
[83]
Schoenholz, Maithra Raghu, Martin Wattenberg, and Ian Goodfellow
Justin Gilmer, Luke Metz, Fartash Faghri, Samuel S. Schoenholz, Maithra Raghu, Martin Wattenberg, and Ian Goodfellow. Adversarial spheres. InInternational Conference on Learning Representations Workshop, 2018. 66
2018
-
[84]
The surprising harmfulness of benign overfitting for adversarial robustness.arXiv preprint arXiv:2401.12236, 2024
Yifan Hao and Tong Zhang. The surprising harmfulness of benign overfitting for adversarial robustness.arXiv preprint arXiv:2401.12236, 2024
2024 arXiv
-
[85]
HaoChen, Colin Wei, Jason D
Jeff Z. HaoChen, Colin Wei, Jason D. Lee, and Tengyu Ma. Shape matters: Understanding the implicit bias of the noise covariance. InProceedings of the Thirty Fourth Conference on Learning Theory, volume 134 ofProceedings of Machine Learning Research, pages 2315–2357. PMLR, 2021
2021
-
[86]
On the geometry of differential privacy
Moritz Hardt and Kunal Talwar. On the geometry of differential privacy. InProceedings of the Forty-Second ACM Symposium on Theory of Computing, pages 705–714, 2010. doi: 10.1145/1806689.1806786
2010
-
[87]
Train faster, generalize better: Stability of stochastic gradient descent
Moritz Hardt, Benjamin Recht, and Yoram Singer. Train faster, generalize better: Stability of stochastic gradient descent. InProceedings of the 33rd International Conference on Machine Learning, pages 1225–1234, 2016
2016
-
[88]
The optimal sample complexity of PAC learning.Journal of Machine Learning Research, 17(38):1–15, 2016
Steve Hanneke. The optimal sample complexity of PAC learning.Journal of Machine Learning Research, 17(38):1–15, 2016
2016
-
[89]
Tibshirani
Trevor Hastie, Andrea Montanari, Saharon Rosset, and Ryan J. Tibshirani. Surprises in high-dimensional ridgeless least squares interpolation.The Annals of Statistics, 50(2):949–986,
-
[90]
Herman.Fundamentals of Computerized Tomography: Image Reconstruction from Projections
Gabor T. Herman.Fundamentals of Computerized Tomography: Image Reconstruction from Projections. Advances in Pattern Recognition. Springer, London, 2nd edition, 2009. doi: 10.1007/978-1-84628-723-7. Algebraic reconstruction techniques and relaxation: Chapter 11, pp. 193–216
2009 doi
-
[91]
Flat minima.Neural Computation, 9(1):1–42, 1997
Sepp Hochreiter and J¨ urgen Schmidhuber. Flat minima.Neural Computation, 9(1):1–42, 1997. doi: 10.1162/neco.1997.9.1.1. 67
1997 doi
-
[92]
Euclidean information theory of networks
Shao-Lun Huang, Changho Suh, and Lizhong Zheng. Euclidean information theory of networks. IEEE Transactions on Information Theory, 61(12):6795–6814, 2015. doi: 10.1109/TIT.2015. 2484066
2015 doi
-
[93]
Hochwald
Babak Hassibi and Bertrand M. Hochwald. How much training is needed in multiple-antenna wireless links?IEEE Transactions on Information Theory, 49(4):951–963, 2003. doi: 10.1109/ TIT.2003.809594
2003
-
[94]
Optimizer dynamics at the edge of stability with differential privacy, 2025
Ayana Hussain and Ricky Fang. Optimizer dynamics at the edge of stability with differential privacy, 2025. URLhttps://arxiv.org/abs/2512.19019
2025
-
[95]
Adversarial examples are not bugs, they are features
Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras, Logan Engstrom, Brandon Tran, and Aleksander Madry. Adversarial examples are not bugs, they are features. InAdvances in Neural Information Processing Systems, volume 32, 2019
2019
-
[96]
Kakade, Rahul Kidambi, Praneeth Netrapalli, and Aaron Sidford
Prateek Jain, Sham M. Kakade, Rahul Kidambi, Praneeth Netrapalli, and Aaron Sidford. Parallelizing stochastic gradient descent for least squares regression: Mini-batching, averaging, and model misspecification.Journal of Machine Learning Research, 18(223):1–42, 2018. URL http:...
2018
-
[97]
James and Charles Stein
W. James and Charles Stein. Estimation with quadratic loss. InProceedings of the Fourth Berkeley Symposium on Mathematical Statistics and Probability, volume 1, pages 361–379. University of California Press, 1961
1961
-
[98]
Three factors influencing minima in SGD, 2017
Stanislaw Jastrzebski, Zachary Kenton, Devansh Arpit, Nicolas Ballas, Asja Fischer, Yoshua Bengio, and Amos Storkey. Three factors influencing minima in SGD, 2017. URL https: //arxiv.org/abs/1711.04623
2017 arXiv
-
[99]
Wornell, and Lizhong Zheng
Shao-Lun Huang, Anuran Makur, Gregory W. Wornell, and Lizhong Zheng. Universal features for high-dimensional learning and inference: Information theoretic and geometric perspectives. Foundations and Trends in Communications and Information Theory, 21(1–2):1–299, 2024. doi: 10....
2024 doi
-
[100]
Jerrum, Leslie G
Mark R. Jerrum, Leslie G. Valiant, and Vijay V. Vazirani. Random generation of combinatorial structures from a uniform distribution.Theoretical Computer Science, 43:169–188, 1986
1986
-
[101]
Kakade, and Michael I
Chi Jin, Rong Ge, Praneeth Netrapalli, Sham M. Kakade, and Michael I. Jordan. How to escape saddle points efficiently. InProceedings of the 34th International Conference on Machine Learning, pages 1724–1732, 2017
2017
-
[102]
Johnstone
Iain M. Johnstone. On the distribution of the largest eigenvalue in principal components analysis.The Annals of Statistics, 29(2):295–327, 2001
2001
-
[103]
Angen¨ aherte Aufl¨ osung von Systemen linearer Gleichungen.Bulletin International de l’Acad´ emie Polonaise des Sciences et des Lettres, Classe A, 35:355–357, 1937
Stefan Kaczmarz. Angen¨ aherte Aufl¨ osung von Systemen linearer Gleichungen.Bulletin International de l’Acad´ emie Polonaise des Sciences et des Lettres, Classe A, 35:355–357, 1937
1937
-
[104]
SGD: The role of implicit regularization, batch-size and multiple epochs
Satyen Kale, Ayush Sekhari, and Karthik Sridharan. SGD: The role of implicit regularization, batch-size and multiple epochs. InAdvances in Neural Information Processing Systems, volume 34, pages 27422–27433, 2021
2021
-
[105]
Precise tradeoffs in adversarial training for linear regression
Adel Javanmard, Mahdi Soltanolkotabi, and Hamed Hassani. Precise tradeoffs in adversarial training for linear regression. InProceedings of Thirty Third Conference on Learning Theory, volume 125 ofProceedings of Machine Learning Research, pages 2034–2078, 2020
-
[106]
B. S. Kashin. The widths of certain finite-dimensional sets and classes of smooth functions. Izvestiya Akademii Nauk SSSR. Seriya Matematicheskaya, 41:334–351, 1977
1977
-
[107]
Lower bounds for differential privacy from gaussian width
Assimakis Kattis and Aleksandar Nikolov. Lower bounds for differential privacy from gaussian width. In33rd International Symposium on Computational Geometry, volume 77 ofLeibniz International Proceedings in Informatics, pages 45:1–45:16, 2017. doi: 10.4230/LIPIcs.SoCG. 2017.45
2017 doi
-
[108]
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang. On large-batch training for deep learning: Generalization gap and sharp minima. InInternational Conference on Learning Representations, 2017. URL https: //openreview.net/for...
2017
-
[109]
Membership inference attacks beyond overfitting
Mona Khalil, Alberto Blanco-Justicia, Najeeb Jebreel, and Josep Domingo-Ferrer. Membership inference attacks beyond overfitting. InComputer Security. ESORICS 2025 International Workshops, volume 16231 ofLecture Notes in Computer Science, pages 32–48. Springer, 2026. doi: 10.10...
2025
-
[110]
Kimeldorf and Grace Wahba
George S. Kimeldorf and Grace Wahba. Some results on Tchebycheffian spline func- tions.Journal of Mathematical Analysis and Applications, 33(1):82–95, 1971. doi: 10.1016/0022-247X(71)90184-3
1971 doi
-
[111]
New lower bounds for private esti- mation and a generalized fingerprinting lemma
Gautam Kamath, Argyris Mouzakis, and Vikrant Singhal. New lower bounds for private esti- mation and a generalized fingerprinting lemma. InAdvances in Neural Information Processing 68 Systems, volume 35, 2022. URL https://proceedings.neurips.cc/paper_files/paper/ 2022/hash/9a6b...
2022
-
[112]
The optimal ridge penalty for real- world high-dimensional data can be zero or negative due to the implicit ridge regularization
Dmitry Kobak, Jonathan Lomond, and Benoit Sanchez. The optimal ridge penalty for real- world high-dimensional data can be zero or negative due to the implicit ridge regularization. Journal of Machine Learning Research, 21(169):1–16, 2020
2020
-
[113]
Sutherland, and Nathan Srebro
Frederic Koehler, Lijia Zhou, Danica J. Sutherland, and Nathan Srebro. Uniform convergence of interpolators: Gaussian width, norm bounds and benign overfitting. InAdvances in Neural Information Processing Systems 34, 2021
2021
-
[114]
Benign underfitting of stochastic gradient descent
Tomer Koren, Roi Livni, Yishay Mansour, and Uri Sherman. Benign underfitting of stochastic gradient descent. InAdvances in Neural Information Processing Systems, volume 35, pages 19605–19617, 2022. arXiv:2202.13361
2022 arXiv
-
[115]
V. A. Kotel’nikov.The Theory of Optimum Noise Immunity. McGraw–Hill, New York, 1959. English translation of the Russian edition; based on the author’s 1947 doctoral dissertation
1959
-
[116]
Bagging is an optimal PAC learner
Kasper Green Larsen. Bagging is an optimal PAC learner. InProceedings of the Thirty Sixth Conference on Learning Theory, volume 195 ofProceedings of Machine Learning Research, pages 450–468. PMLR, 2023
2023
-
[117]
John Wiley & Sons, New York, 1965
Leslie Kish.Survey Sampling. John Wiley & Sons, New York, 1965
1965
-
[118]
The large learning rate phase of deep learning: The catapult mechanism, 2020
Aitor Lewkowycz, Yasaman Bahri, Ethan Dyer, Jascha Sohl-Dickstein, and Guy Gur-Ari. The large learning rate phase of deep learning: The catapult mechanism, 2020. URL https: //arxiv.org/abs/2003.02218. 69
2020 arXiv
-
[119]
Feature averaging: An implicit bias of gradient descent leading to non-robustness in neural networks
Binghui Li, Zhixuan Pan, Kaifeng Lyu, and Jian Li. Feature averaging: An implicit bias of gradient descent leading to non-robustness in neural networks. InInternational Conference on Learning Representations, 2025. URLhttps://openreview.net/forum?id=zPHra4V5Mc
2025
-
[120]
Towards explaining the regularization effect of initial large learning rate in training neural networks
Yuanzhi Li, Colin Wei, and Tengyu Ma. Towards explaining the regularization effect of initial large learning rate in training neural networks. InAdvances in Neural Information Processing Systems 32, 2019
2019
-
[121]
Kakade, Peter L
Licong Lin, Jingfeng Wu, Sham M. Kakade, Peter L. Bartlett, and Jason D. Lee. Scaling laws in linear regression: Compute, parameters, and data. InAdvances in Neural Information Processing Systems, volume 37, 2024. arXiv:2406.08466
2024 arXiv
-
[122]
Bartlett
Licong Lin, Jingfeng Wu, and Peter L. Bartlett. Improved scaling laws in linear regression via data reuse. InAdvances in Neural Information Processing Systems, 2025. arXiv:2506.08415
2025
-
[123]
A new characterization of the edge of stability based on a sharpness measure aware of batch gradient distribution
Sungyoon Lee and Cheongjae Jang. A new characterization of the edge of stability based on a sharpness measure aware of batch gradient distribution. InInternational Conference on Learning Representations, 2023. URLhttps://openreview.net/forum?id=bH-kCY6LdKg
2023
-
[124]
The sample complexity of gradient descent in stochastic convex optimization
Roi Livni. The sample complexity of gradient descent in stochastic convex optimization. In Advances in Neural Information Processing Systems, volume 37, 2024. arXiv:2404.04931
2024 arXiv
-
[125]
Krijthe, and David M
Marco Loog, Tom Viering, Alexander Mey, Jesse H. Krijthe, and David M. J. Tax. A brief prehistory of double descent.Proceedings of the National Academy of Sciences, 117(20): 10625–10626, 2020. doi: 10.1073/pnas.2001875117
2020 doi
-
[126]
Mean estimation and regression under heavy-tailed distributions: A survey.Foundations of Computational Mathematics, 19(5):1145–1190, 2019
G´ abor Lugosi and Shahar Mendelson. Mean estimation and regression under heavy-tailed distributions: A survey.Foundations of Computational Mathematics, 19(5):1145–1190, 2019
2019
-
[127]
Sub-Gaussian estimators of the mean of a random vector.The Annals of Statistics, 47(2):783–794, 2019
G´ abor Lugosi and Shahar Mendelson. Sub-Gaussian estimators of the mean of a random vector.The Annals of Statistics, 47(2):783–794, 2019
2019
-
[128]
Regularization, sparse recovery, and median-of-means tournaments.Bernoulli, 25(3):2075–2106, 2019
G´ abor Lugosi and Shahar Mendelson. Regularization, sparse recovery, and median-of-means tournaments.Bernoulli, 25(3):2075–2106, 2019
-
[129]
Making progress based on false discoveries
Roi Livni. Making progress based on false discoveries. In15th Innovations in Theoretical Computer Science Conference (ITCS 2024), volume 287 ofLeibniz International Proceedings in Informatics (LIPIcs), pages 76:1–76:18, 2024. doi: 10.4230/LIPIcs.ITCS.2024.76
2024 doi
-
[130]
On the SDEs and scaling rules for adaptive gradient algorithms
Sadhika Malladi, Kaifeng Lyu, Abhishek Panigrahi, and Sanjeev Arora. On the SDEs and scaling rules for adaptive gradient algorithms. InAdvances in Neural Information Processing Systems, volume 35, 2022. URL https://papers.neurips.cc/paper_files/paper/2022/ hash/32ac710102f0620...
2022
-
[131]
Simon, Amirhesam Abedsoltan, Parthe Pandit, Misha Belkin, and Preetum Nakkiran
Neil Mallinar, James B. Simon, Amirhesam Abedsoltan, Parthe Pandit, Misha Belkin, and Preetum Nakkiran. Benign, tempered, or catastrophic: Toward a re- fined taxonomy of overfitting. InAdvances in Neural Information Processing Sys- tems 35, 2022. URL https://proceedings.neurip...
2022
-
[132]
Hoffman, and David M
Stephan Mandt, Matthew D. Hoffman, and David M. Blei. Stochastic gradient descent as approximate Bayesian inference.Journal of Machine Learning Research, 18(134):1–35, 2017. URLhttps://jmlr.org/papers/v18/17-214.html. 70
2017
-
[133]
An empirical model of large-batch training, 2018
Sam McCandlish, Jared Kaplan, Dario Amodei, and OpenAI Dota Team. An empirical model of large-batch training, 2018. URLhttps://arxiv.org/abs/1812.06162
2018 arXiv
-
[134]
McRae, Santhosh Karnik, Mark Davenport, and Vidya K
Andrew D. McRae, Santhosh Karnik, Mark Davenport, and Vidya K. Muthukumar. Harmless interpolation in regression and classification with structured features. InProceedings of The 25th International Conference on Artificial Intelligence and Statistics, volume 151 of Proceedings ...
2022
-
[135]
Towards deep learning models resistant to adversarial attacks
Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu. Towards deep learning models resistant to adversarial attacks. InInternational Conference on Learning Representations, 2018
2018
-
[136]
Hancheng Min and Ren´ e Vidal. Can implicit bias imply adversarial robustness? InProceedings of the 41st International Conference on Machine Learning, volume 235 ofProceedings of Machine Learning Research, pages 35687–35718. PMLR, 2024. URL https://proceedings. mlr.press/v235/...
2024
-
[137]
Geometric median and robust estimation in Banach spaces.Bernoulli, 21 (4):2308–2335, 2015
Stanislav Minsker. Geometric median and robust estimation in Banach spaces.Bernoulli, 21 (4):2308–2335, 2015
2015
-
[138]
Random reshuffling: Simple analysis with vast improvements
Konstantin Mishchenko, Ahmed Khaled, and Peter Richt´ arik. Random reshuffling: Simple analysis with vast improvements. InAdvances in Neural Information Processing Systems, volume 33, pages 17309–17320, 2020
2020
-
[139]
VC classes are adversarially robustly learnable, but only improperly
Omar Montasser, Steve Hanneke, and Nathan Srebro. VC classes are adversarially robustly learnable, but only improperly. InProceedings of the Thirty-Second Conference on Learning Theory, volume 99 ofProceedings of Machine Learning Research, pages 2512–2530. PMLR, 2019
2019
-
[140]
Harmless interpolation of noisy data in regression
Vidya Muthukumar, Kailas Vodrahalli, and Anant Sahai. Harmless interpolation of noisy data in regression. InIEEE International Symposium on Information Theory, pages 2299–2303,
-
[141]
Gallager
Muriel M´ edard and Robert G. Gallager. Bandwidth scaling for fading multipath channels. IEEE Transactions on Information Theory, 48(4):840–852, 2002. doi: 10.1109/18.992769
2002 doi
-
[142]
Classification vs regression in overparameterized regimes: Does the loss function matter?Journal of Machine Learning Research, 22(222):1–69, 2021
Vidya Muthukumar, Adhyyan Narang, Vignesh Subramanian, Mikhail Belkin, Daniel Hsu, and Anant Sahai. Classification vs regression in overparameterized regimes: Does the loss function matter?Journal of Machine Learning Research, 22(222):1–69, 2021. URL https: //jmlr.org/papers/v...
2021
-
[143]
Implicit bias of the step size in linear diagonal neural networks
Mor Shpigel Nacson, Kavya Ravichandran, Nathan Srebro, and Daniel Soudry. Implicit bias of the step size in linear diagonal neural networks. InProceedings of the 39th International Conference on Machine Learning, volume 162 ofProceedings of Machine Learning Research, pages 162...
2022
-
[144]
Zico Kolter
Vaishnavh Nagarajan and J. Zico Kolter. Uniform convergence may be unable to explain generalization in deep learning. InAdvances in Neural Information Processing Systems, volume 32, 2019. 71
2019
-
[145]
Deep double descent: Where bigger models and more data hurt
Preetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang, Boaz Barak, and Ilya Sutskever. Deep double descent: Where bigger models and more data hurt. InInternational Conference on Learning Representations, 2020. URLhttps://openreview.net/forum?id=B1g5sA4twr
2020
-
[146]
The deep bootstrap framework: Good online learners are good offline generalizers
Preetum Nakkiran, Behnam Neyshabur, and Hanie Sedghi. The deep bootstrap framework: Good online learners are good offline generalizers. InInternational Conference on Learning Representations, 2021
2021
-
[147]
Classification and adversarial examples in an overparameterized linear model: A signal processing perspective, 2021
Adhyyan Narang, Vidya Muthukumar, and Anant Sahai. Classification and adversarial examples in an overparameterized linear model: A signal processing perspective, 2021. URL https://arxiv.org/abs/2109.13215
2021 arXiv
-
[148]
Harmless interpolation of noisy data in regression.IEEE Journal on Selected Areas in Information Theory, 1(1):67–83, 2020
Vidya Muthukumar, Kailas Vodrahalli, Vignesh Subramanian, and Anant Sahai. Harmless interpolation of noisy data in regression.IEEE Journal on Selected Areas in Information Theory, 1(1):67–83, 2020. doi: 10.1109/JSAIT.2020.2984716
2020
-
[149]
Robust stochastic approximation approach to stochastic programming.SIAM Journal on Optimization, 19(4): 1574–1609, 2009
Arkadi Nemirovski, Anatoli Juditsky, Guanghui Lan, and Alexander Shapiro. Robust stochastic approximation approach to stochastic programming.SIAM Journal on Optimization, 19(4): 1574–1609, 2009. doi: 10.1137/070704277
2009 doi
-
[150]
Nemirovsky and David B
Arkadi S. Nemirovsky and David B. Yudin.Problem Complexity and Method Efficiency in Optimization. Wiley-Interscience Series in Discrete Mathematics. John Wiley & Sons, New York, 1983
1983
-
[151]
Simplified neuron model as a principal component analyzer.Journal of Mathematical Biology, 15(3):267–273, 1982
Erkki Oja. Simplified neuron model as a principal component analyzer.Journal of Mathematical Biology, 15(3):267–273, 1982. doi: 10.1007/BF00275687
1982 doi
-
[152]
Deep learning on a data diet: Finding important examples early in training
Mansheej Paul, Surya Ganguli, and Gintare Karolina Dziugaite. Deep learning on a data diet: Finding important examples early in training. InAdvances in Neural Information Processing Systems 34, 2021
2021
-
[153]
Implicit bias of SGD for diagonal linear networks: a provable benefit of stochasticity
Scott Pesme, Loucas Pillaud-Vivien, and Nicolas Flammarion. Implicit bias of SGD for diagonal linear networks: a provable benefit of stochasticity. InAdvances in Neural Information Processing Systems 34, 2021
2021
-
[154]
Statistical optimality of stochastic gradient descent on hard learning problems through multiple passes
Loucas Pillaud-Vivien, Alessandro Rudi, and Francis Bach. Statistical optimality of stochastic gradient descent on hard learning problems through multiple passes. InAdvances in Neural Information Processing Systems, volume 31, 2018
2018
-
[155]
Stochastic gradient descent, weighted sampling, and the randomized Kaczmarz algorithm
Deanna Needell, Nathan Srebro, and Rachel Ward. Stochastic gradient descent, weighted sampling, and the randomized Kaczmarz algorithm. InAdvances in Neural Information Processing Systems, volume 27, 2014
2014
-
[156]
Polyak and Anatoli B
Boris T. Polyak and Anatoli B. Juditsky. Acceleration of stochastic approximation by averaging. SIAM Journal on Control and Optimization, 30(4):838–855, 1992
1992
-
[157]
Under- standing and mitigating the tradeoff between robustness and accuracy
Aditi Raghunathan, Sang Michael Xie, Fanny Yang, John Duchi, and Percy Liang. Under- standing and mitigating the tradeoff between robustness and accuracy. InProceedings of the 37th International Conference on Machine Learning, 2020
2020
-
[158]
The fourth quadrant: Benign misfitting across multiple spikes, 2026
Gireeja Ranade and Anant Sahai. The fourth quadrant: Benign misfitting across multiple spikes, 2026. In preparation
2026
-
[159]
The fourth quadrant: One-pass direction recovery at the BBP scale, 2026
Gireeja Ranade and Anant Sahai. The fourth quadrant: One-pass direction recovery at the BBP scale, 2026. In preparation. 72
2026
-
[160]
The fourth quadrant: Same-distribution tests of benign misfitting, 2026
Gireeja Ranade and Anant Sahai. The fourth quadrant: Same-distribution tests of benign misfitting, 2026. In preparation
2026
-
[161]
Understanding the generalization benefits of late learning rate decay
Yinuo Ren, Chao Ma, and Lexing Ying. Understanding the generalization benefits of late learning rate decay. InProceedings of The 27th International Conference on Artificial In- telligence and Statistics, volume 238 ofProceedings of Machine Learning Research, pages 4465–4473. P...
2024
-
[162]
Polyak.Introduction to Optimization
Boris T. Polyak.Introduction to Optimization. Optimization Software, Inc., New York, 1987
1987
-
[163]
A stochastic approximation method.The Annals of Mathematical Statistics, 22(3):400–407, 1951
Herbert Robbins and Sutton Monro. A stochastic approximation method.The Annals of Mathematical Statistics, 22(3):400–407, 1951. doi: 10.1214/aoms/1177729586
1951
-
[164]
Efficient estimations from a slowly convergent Robbins–Monro process
David Ruppert. Efficient estimations from a slowly convergent Robbins–Monro process. Technical Report 781, Cornell University, School of Operations Research and Industrial Engineering, 1988
1988
-
[165]
Controlling bias in adaptive data analysis using information theory
Daniel Russo and James Zou. Controlling bias in adaptive data analysis using information theory. InProceedings of the 19th International Conference on Artificial Intelligence and Statistics, volume 51 ofProceedings of Machine Learning Research, pages 1232–1240. PMLR, 2016
2016
-
[166]
Hashimoto, and Percy Liang
Shiori Sagawa, Pang Wei Koh, Tatsunori B. Hashimoto, and Percy Liang. Distributionally robust neural networks for group shifts: On the importance of regularization for worst-case generalization. InInternational Conference on Learning Representations, 2020
2020
-
[167]
Raj Rajagopalan, and H
Lalitha Sankar, S. Raj Rajagopalan, and H. Vincent Poor. Utility-privacy tradeoffs in databases: An information-theoretic approach.IEEE Transactions on Information Forensics and Security, 8(6):838–852, 2013
2013
-
[168]
Saxe, James L
Andrew M. Saxe, James L. McClelland, and Surya Ganguli. Exact solutions to the nonlinear dynamics of learning in deep linear neural networks. InInternational Conference on Learning Representations, 2014
2014
-
[169]
Ribeiro and Thomas B
Antˆ onio H. Ribeiro and Thomas B. Sch¨ on. Overparameterized linear regression under adversarial attacks.IEEE Transactions on Signal Processing, 71:601–614, 2023. doi: 10.1109/ TSP.2023.3246228
2023
-
[170]
The dimension strikes back with gradients: Generalization of gradient methods in stochastic convex optimization
Matan Schliserman, Uri Sherman, and Tomer Koren. The dimension strikes back with gradients: Generalization of gradient methods in stochastic convex optimization. InProceedings of the 36th International Conference on Algorithmic Learning Theory, volume 272 ofProceedings of Mach...
2025 arXiv
-
[171]
Adversarially robust generalization requires more data
Ludwig Schmidt, Shibani Santurkar, Dimitris Tsipras, Kunal Talwar, and Aleksander Madry. Adversarially robust generalization requires more data. InAdvances in Neural Information Processing Systems, volume 31, 2018
2018
-
[172]
Bernhard Sch¨ olkopf, Ralf Herbrich, and Alex J. Smola. A generalized representer theorem. In Computational Learning Theory, volume 2111 ofLecture Notes in Computer Science, pages 416–426, Berlin, Heidelberg, 2001. Springer. doi: 10.1007/3-540-44581-1 27. 73
2001 doi
-
[173]
The pitfalls of simplicity bias in neural networks
Harshay Shah, Kaustav Tamuly, Aditi Raghunathan, Prateek Jain, and Praneeth Netrapalli. The pitfalls of simplicity bias in neural networks. InAdvances in Neural Information Process- ing Systems, volume 33, pages 9573–9585, 2020. URL https://proceedings.neurips.cc/ paper_files/...
2020
-
[174]
Learnability, stability and uniform convergence.Journal of Machine Learning Research, 11:2635–2670, 2010
Shai Shalev-Shwartz, Ohad Shamir, Nathan Srebro, and Karthik Sridharan. Learnability, stability and uniform convergence.Journal of Machine Learning Research, 11:2635–2670, 2010
2010
-
[175]
Shallue, Jaehoon Lee, Joseph Antognini, Jascha Sohl-Dickstein, Roy Frostig, and George E
Christopher J. Shallue, Jaehoon Lee, Joseph Antognini, Jascha Sohl-Dickstein, Roy Frostig, and George E. Dahl. Measuring the effects of data parallelism on neural network training,
-
[176]
Robust linear regression: Gradient-descent, early- stopping, and beyond
Meyer Scetbon and Elvis Dohmatob. Robust linear regression: Gradient-descent, early- stopping, and beyond. InProceedings of the 26th International Conference on Artificial Intelligence and Statistics, volume 206 ofProceedings of Machine Learning Research. PMLR, 2023
2023
-
[177]
Membership inference attacks against machine learning models
Reza Shokri, Marco Stronati, Congzheng Song, and Vitaly Shmatikov. Membership inference attacks against machine learning models. In2017 IEEE Symposium on Security and Privacy (SP), pages 3–18. IEEE, 2017
2017
-
[178]
Certifying some distributional robustness with principled adversarial training
Aman Sinha, Hongseok Namkoong, and John Duchi. Certifying some distributional robustness with principled adversarial training. InInternational Conference on Learning Representations, 2018
2018
-
[179]
Leslie N. Smith. Cyclical learning rates for training neural networks. In2017 IEEE Winter Conference on Applications of Computer Vision (WACV), pages 464–472, 2017. doi: 10.1109/ W ACV.2017.58
2017
-
[180]
Smith and Nicholay Topin
Leslie N. Smith and Nicholay Topin. Super-convergence: Very fast training of neural networks using large learning rates. InArtificial Intelligence and Machine Learning for Multi-Domain Operations Applications, volume 11006, page 1100612. SPIE, 2019. doi: 10.1117/12.2520589
2019 doi
-
[181]
Smith and Quoc V
Samuel L. Smith and Quoc V. Le. A Bayesian perspective on generalization and stochastic gradient descent. InInternational Conference on Learning Representations, 2018. URL https://openreview.net/forum?id=BJij4yg0Z
2018
-
[182]
Smith, Pieter-Jan Kindermans, Chris Ying, and Quoc V
Samuel L. Smith, Pieter-Jan Kindermans, Chris Ying, and Quoc V. Le. Don’t decay the learning rate, increase the batch size. InInternational Conference on Learning Representations,
-
[183]
URLhttps://arxiv.org/abs/1811.03600
-
[184]
Claude E. Shannon. Communication in the presence of noise.Proceedings of the IRE, 37(1): 10–21, 1949. doi: 10.1109/JRPROC.1949.232969
1949
-
[185]
Ben Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli, and Ari S. Morcos. Beyond neural scaling laws: Beating power law scaling via data pruning. InAdvances in Neural Information Processing Systems 35, 2022. 74
2022
-
[186]
The implicit bias of gradient descent on separable data.Journal of Machine Learning Research, 19 (70):1–57, 2018
Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and Nathan Srebro. The implicit bias of gradient descent on separable data.Journal of Machine Learning Research, 19 (70):1–57, 2018. URLhttps://www.jmlr.org/papers/v19/18-188.html
2018
-
[187]
A jamming transition from under- to over-parametrization affects generalization in deep learning.Journal of Physics A: Mathematical and Theoretical, 52(47):474001, 2019
Stefano Spigler, Mario Geiger, St´ ephane d’Ascoli, Levent Sagun, Giulio Biroli, and Matthieu Wyart. A jamming transition from under- to over-parametrization affects generalization in deep learning.Journal of Physics A: Mathematical and Theoretical, 52(47):474001, 2019. doi: 1...
2019 doi
-
[188]
Inadmissibility of the usual estimator for the mean of a multivariate normal distribution
Charles Stein. Inadmissibility of the usual estimator for the mean of a multivariate normal distribution. InProceedings of the Third Berkeley Symposium on Mathematical Statistics and Probability, volume 1, pages 197–206. University of California Press, 1956
1956
-
[189]
Reasoning about generalization via conditional mutual information
Thomas Steinke and Lydia Zakynthinou. Reasoning about generalization via conditional mutual information. InProceedings of Thirty Third Conference on Learning Theory, volume 125 ofProceedings of Machine Learning Research, pages 3437–3452. PMLR, 2020
2020
-
[190]
A randomized Kaczmarz algorithm with exponential convergence.Journal of Fourier Analysis and Applications, 15(2):262–278, 2009
Thomas Strohmer and Roman Vershynin. A randomized Kaczmarz algorithm with exponential convergence.Journal of Fourier Analysis and Applications, 15(2):262–278, 2009. doi: 10.1007/ s00041-008-9030-4
2009
-
[191]
URLhttps://openreview.net/forum?id=B1Yy1BxCZ
-
[192]
Smith, Benoit Dherin, David G
Samuel L. Smith, Benoit Dherin, David G. T. Barrett, and Soham De. On the origin of implicit regularization in stochastic gradient descent. InInternational Conference on Learning Representations, 2021. URLhttps://openreview.net/forum?id=rq_Qr0c1Hyo
2021
-
[193]
Evading the curse of dimensionality in unconstrained private GLMs
Shuang Song, Thomas Steinke, Om Thakkar, and Abhradeep Thakurta. Evading the curse of dimensionality in unconstrained private GLMs. InInternational Conference on Artificial Intelligence and Statistics, pages 2223–2234, 2021
2021
-
[194]
Emre Telatar and David N. C. Tse. Capacity and mutual information of wideband multipath fading channels.IEEE Transactions on Information Theory, 46(4):1384–1400, 2000. doi: 10.1109/18.850674
-
[195]
Mariya Toneva, Alessandro Sordoni, Remi Tachet des Combes, Adam Trischler, Yoshua Bengio, and Geoffrey J. Gordon. An empirical study of example forgetting during deep neural network learning. InInternational Conference on Learning Representations, 2019
2019
-
[196]
Fixing the train-test resolution discrepancy
Hugo Touvron, Andrea Vedaldi, Matthijs Douze, and Herv´ e J´ egou. Fixing the train-test resolution discrepancy. InAdvances in Neural Information Processing Systems 32, 2019
2019
-
[197]
Bartlett
Alexander Tsigler and Peter L. Bartlett. Benign overfitting in ridge regression.Journal of Machine Learning Research, 24(123):1–76, 2023
2023
-
[198]
The price of implicit bias in adversarially robust generalization, 2024
Nikolaos Tsilivis, Natalie Frank, Nathan Srebro, and Julia Kempe. The price of implicit bias in adversarially robust generalization, 2024. URLhttps://arxiv.org/abs/2406.04981
2024 arXiv
-
[199]
Robustness may be at odds with accuracy
Dimitris Tsipras, Shibani Santurkar, Logan Engstrom, Alexander Turner, and Aleksander Madry. Robustness may be at odds with accuracy. InInternational Conference on Learning Representations, 2019. 75
2019
-
[200]
Generalization for multiclass classifica- tion with overparameterized linear models
Vignesh Subramanian, Rahul Arya, and Anant Sahai. Generalization for multiclass classifica- tion with overparameterized linear models. InAdvances in Neural Information Processing Systems 35, 2022. URLhttps://openreview.net/forum?id=ikWvMRVQBWW
2022
-
[201]
Subramanian and Bruce Hajek
Vijay G. Subramanian and Bruce Hajek. Broad-band fading channels: Signal burstiness and capacity.IEEE Transactions on Information Theory, 48(4):809–827, 2002. doi: 10.1109/18. 992768
2002 doi
-
[202]
Intriguing properties of neural networks
Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus. Intriguing properties of neural networks. InInternational Conference on Learning Representations, 2014
2014
-
[203]
Cambridge University Press, 2018
Roman Vershynin.High-Dimensional Probability: An Introduction with Applications in Data Science. Cambridge University Press, 2018
2018
-
[204]
Benign overfitting in multiclass classification: All roads lead to interpolation
Ke Wang, Vidya Muthukumar, and Christos Thrampoulidis. Benign overfitting in multiclass classification: All roads lead to interpolation. InAdvances in Neural Information Processing Systems 34, pages 24164–24179, 2021. URL https://proceedings.neurips.cc/paper/ 2021/hash/caaa29e...
2021
-
[205]
Good regularity creates large learning rate implicit biases: Edge of stability, balancing, and catapult.Journal of Machine Learning Research, 26(273):1–68, 2025
Yuqing Wang, Zhenghao Xu, Tuo Zhao, and Molei Tao. Good regularity creates large learning rate implicit biases: Edge of stability, balancing, and catapult.Journal of Machine Learning Research, 26(273):1–68, 2025. URLhttps://jmlr.org/papers/v26/23-1691.html
2025
-
[206]
Near-interpolators: Rapid norm growth and the trade-off between interpolation and generalization
Yutong Wang, Rishi Sonthalia, and Wei Hu. Near-interpolators: Rapid norm growth and the trade-off between interpolation and generalization. InInternational Conference on Artificial Intelligence and Statistics, volume 238 ofProceedings of Machine Learning Research, 2024. arXiv:...
2024 arXiv
-
[207]
Stearns.Adaptive Signal Processing
Bernard Widrow and Samuel D. Stearns.Adaptive Signal Processing. Prentice-Hall, Englewood Cliffs, NJ, 1985
1985
-
[208]
Wozencraft and Irwin Mark Jacobs.Principles of Communication Engineering
John M. Wozencraft and Irwin Mark Jacobs.Principles of Communication Engineering. John Wiley & Sons, New York, 1965
1965
-
[209]
Van Trees.Detection, Estimation, and Modulation Theory, Part I
Harry L. Van Trees.Detection, Estimation, and Modulation Theory, Part I. John Wiley & Sons, New York, 1968
1968
-
[210]
Rapid overfitting of multi-pass SGD in stochastic convex optimization
Shira Vansover-Hager, Tomer Koren, and Roi Livni. Rapid overfitting of multi-pass SGD in stochastic convex optimization. InProceedings of the 42nd International Conference on Machine Learning, volume 267 ofProceedings of Machine Learning Research, pages 60905– 60923. PMLR, 202...
2025 arXiv
-
[211]
Spectral efficiency in the wideband regime.IEEE Transactions on Information Theory, 48(6):1319–1343, 2002
Sergio Verd´ u. Spectral efficiency in the wideband regime.IEEE Transactions on Information Theory, 48(6):1319–1343, 2002. doi: 10.1109/TIT.2002.1003824
2002 arXiv
-
[212]
Lei Wu and Weijie J. Su. The implicit regularization of dynamical stability in stochastic gradient descent. InProceedings of the 40th International Conference on Machine Learning, volume 202 ofProceedings of Machine Learning Research, pages 37656–37684. PMLR, 2023. URLhttps://...
2023
-
[213]
How SGD selects the global minima in over-parameterized learning: A dynamical stability perspective
Lei Wu, Chao Ma, and Weinan E. How SGD selects the global minima in over-parameterized learning: A dynamical stability perspective. InAdvances in Neural Information Processing Systems 31, pages 8279–8288, 2018
2018
-
[214]
Adversarially robust estimate and risk analysis in linear regression
Yue Xing, Ruizhi Zhang, and Guang Cheng. Adversarially robust estimate and risk analysis in linear regression. InProceedings of the 24th International Conference on Artificial Intelligence and Statistics, pages 514–522, 2021
2021
-
[215]
Information-theoretic analysis of generalization capability of learning algorithms
Aolin Xu and Maxim Raginsky. Information-theoretic analysis of generalization capability of learning algorithms. InAdvances in Neural Information Processing Systems, volume 30, pages 2524–2533, 2017
2017
-
[216]
Privacy risk in machine learning: Analyzing the connection to overfitting
Samuel Yeom, Irene Giacomelli, Matt Fredrikson, and Somesh Jha. Privacy risk in machine learning: Analyzing the connection to overfitting. InIEEE 31st Computer Security Foundations Symposium (CSF), pages 268–282. IEEE, 2018
2018
-
[217]
CutMix: Regularization strategy to train strong classifiers with localizable features
Sangdoo Yun, Dongyoon Han, Seong Joon Oh, Sanghyuk Chun, Junsuk Choe, and Youngjoon Yoo. CutMix: Regularization strategy to train strong classifiers with localizable features. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2019
2019
-
[218]
Precise asymptotic generalization for multiclass classification with overparameterized linear models
David Xing Wu and Anant Sahai. Precise asymptotic generalization for multiclass classification with overparameterized linear models. InAdvances in Neural Information Processing Systems 36, 2023. URLhttps://openreview.net/forum?id=cRGINXQWem
2023
-
[219]
On the optimal weighted ℓ2 regularization in overparameterized linear regression
Denny Wu and Ji Xu. On the optimal weighted ℓ2 regularization in overparameterized linear regression. InAdvances in Neural Information Processing Systems, volume 33, pages 10112–10123, 2020
2020
-
[220]
Jingfeng Wu, Difan Zou, Vladimir Braverman, Quanquan Gu, and Sham M. Kakade. Last iterate risk bounds of SGD with decaying stepsize for overparameterized linear regression. In International Conference on Machine Learning (ICML), 2022. arXiv:2110.06198
2022 arXiv
-
[221]
The anisotropic noise in stochastic gradient descent: Its behavior of escaping from sharp minima and regularization effects
Zhanxing Zhu, Jingfeng Wu, Bing Yu, Lei Wu, and Jinwen Ma. The anisotropic noise in stochastic gradient descent: Its behavior of escaping from sharp minima and regularization effects. InProceedings of the 36th International Conference on Machine Learning, volume 97 ofProceedin...
2019
-
[222]
Difan Zou, Jingfeng Wu, Vladimir Braverman, Quanquan Gu, and Sham M. Kakade. Benign overfitting of constant-stepsize SGD for linear regression.Journal of Machine Learning Research, 24(326):1–58, 2023. URL http://jmlr.org/papers/v24/21-1297.html. Short version in Proc. 34th Con...
2023
-
[227]
Dauphin, and David Lopez-Paz
Hongyi Zhang, Moustapha Cisse, Yann N. Dauphin, and David Lopez-Paz. mixup: Beyond empirical risk minimization. InInternational Conference on Learning Representations, 2018
2018
-
[228]
Lizhong Zheng and David N. C. Tse. Communication on the grassmann manifold: A geometric approach to the noncoherent multiple-antenna channel.IEEE Transactions on Information Theory, 48(2):359–383, 2002. doi: 10.1109/18.978730
2002 doi
-
[229]
Sutherland, and Nathan Srebro
Lijia Zhou, Frederic Koehler, Danica J. Sutherland, and Nathan Srebro. Optimistic rates: A unifying theory for interpolation learning and regularization in linear regression.ACM/IMS Journal of Data Science, 1(2):1–51, 2024. doi: 10.1145/3594234
2024 doi
-
[2005]
doi: 10.1214/009117905000000233
-
[2017]
URLhttps://proceedings.mlr.press/v70/dinh17b.html
-
[2018]
URLhttps://proceedings.mlr.press/v80/belkin18a.html
-
[2019]
doi: 10.1109/ISIT.2019.8849614
2019
-
[2020]
URLhttps://proceedings.mlr.press/v119/ali20a.html
-
[2022]
doi: 10.1214/21-AOS2133
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.