Pith. sign in

REVIEW 3 major objections 5 minor 60 references

The Zero Pattern of a Design Matrix Drives Multiple Descent in Over-parameterized Regression

T0 review · 3 major / 5 minor · reviewed 2026-07-31 · deepseek-v4-flash

Pith's one-line read Multiple-descent peaks in over-parameterized regression are set by the zero pattern of the design's variance profile — a matching rule on a bipartite graph — so a single heterogeneous or dependent design can show several risk peaks with not

desk verdict A genuinely new combinatorial mechanism for multiple descent, with a real gap the authors flag themselves: the theorem covers exact square configurations, while the advertised limiting peaks are left as candidates. read the letter →

arxiv 2607.24041 v1 pith:QEEGFMEF submitted 2026-07-27 math.ST cs.LGstat.MLstat.TH

classification math.STcs.LGstat.MLstat.TH MSC 62J0760B2005C7062J05
keywords multipledescentover-parameterizedregressionvarianceprofileDulmage–MendelsohndecompositionstrongHallpropertyridgedataaugmentationrandommatrixlocallaw
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper aims to show that multiple descent — several risk peaks in the test-error curve — can arise from the structure of the data alone, and to give an exact rule for where the peaks sit. For Gaussian designs whose covariances share an eigenbasis and may be rank-deficient, both heterogeneous and dependent designs collapse to a single variance-profile model; the vanishing-ridge risk then has explicit deterministic equivalents. The bias is shown to live exactly on the Dulmage–Mendelsohn degenerate columns of the profile's bipartite graph, and the variance diverges exactly when the residual profile loses the column-side strong Hall property. Extra peaks therefore come from the zeros of the covariance, not from architecture, feature blocks, or kernel scales. The authors conjecture that positive-definite covariances, even non-commuting ones, retain only the classical single peak at p/n = 1.

What carries the argument

The variance profile S, an n×p matrix of per-entry variances, with its bipartite variance graph G_S; maximum matchings of G_S (whose size is the structural rank); the Dulmage–Mendelsohn decomposition, which yields the degenerate column set J_S; and the strong Hall properties (row-side for the bias, column-side for the residual). The fixed-point system 1 = r_l(λ) + Σ_i S_il r_l(λ)/(λ + Σ_j S_ij r_j(λ)) carries the risk; r(0) and ∂r(0) at the hard edge feed the bias and variance. A re-proved local law tracking the spectral parameter down to zero makes the vanishing-ridge limit tractable.

What would settle it

Take the Section 6.2 setup but push anisotropy to an extreme (e.g., eigenvalues 10 and 10^-3 with non-commuting rotations): the paper's conjecture says the risk still peaks only at p/n = 1. Any second peak away from γ = 1 would refute the dichotomy that zeros, not scale ratios, drive multiple descent. Conversely, the two-group near-miss of Remark 10 (ratio approaching a candidate equality without ever attaining it) tests whether divergence requires the exact square configuration.

Watch

Extended reading notes

Core claim

On the paper's own terms: for a variance-profile Gaussian design with vanishing ridge penalty, the limiting prediction risk is governed by a fixed-point vector r(λ): bias mass sits on coordinates with r_l(0) > 0, and the variance is a weighted sum of derivatives ∂r_l(0). The paper identifies r_l(0) > 0 with membership in the degenerate column set J_S of the variance graph (the set some maximum matching leaves unmatched), and proves the variance diverges precisely at switching configurations where the residual submatrix fails the column-side strong Hall property. Consequently the number and location of multiple-descent peaks are fixed by the support pattern of the design — a combinatorial inv

Load-bearing premise

The reduction to a variance profile — and hence the entire matching rule — requires the design's covariance matrices to share a common eigenbasis; without a shared basis there is no S for the graph to act on, and the positive-definite non-commuting case is left as a conjecture.

Editorial extensions

If this is right

  • A single rank-deficient heterogeneous design, or a single dependent design from data augmentation, already produces multiple descent; the peaks sit where the support rule predicts (Sections 6.1 and 6.3).
  • Real degenerate designs — a pretrained transformer's token-embedding block augmented with Gaussian columns — exhibit double descent matching the theory (Section 6.4).
  • Decomposable profiles split blockwise, so the overall risk is the sum of block risks: each block contributes its own candidate peaks, which merge when thresholds coincide.
  • The deterministic equivalents extend beyond Gaussian entries to any independent entries with matching variance profile (Remark 1); heavy-tailed simulations reproduce the Gaussian peak structure.
  • Positive-definite, shared-eigenbasis covariances are non-degenerate and show only the classical peak at p/n = 1; the non-commuting case is conjectured to behave the same.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The matching rule gives a pre-training diagnostic: estimate the support of the covariance profile, run a maximum matching, and predict where risk will spike before fitting.
  • The exact-zero theory plus the paper's near-zero simulations suggest a continuity question: do peak locations shift continuously as zero eigenvalues are replaced by σ→0 scales? A quantitative version would extend the theory to approximate degeneracy.
  • The mechanism ties double descent to structural-rank theory; the same Dulmage–Mendelsohn language may transfer to minimum-norm interpolation in classification or to kernel models with degenerate kernel matrices.
  • If the Section 5 conjecture holds, multiple descent becomes a signal of exact rank deficiency in feature matrices — information about the data, not about the model family.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper studies vanishing-ridge linear regression for Gaussian designs in which observations are heterogeneous or dependent, with covariance matrices that are simultaneously diagonalizable and possibly rank-deficient. It reduces both Model 1 (heterogeneous) and Model 2 (finite-rank dependent) to a variance-profile model, and derives deterministic equivalents for the bias and variance of the prediction risk (Theorem 3, Corollary 4). The paper's central structural claim is that the zero pattern of the variance profile S controls multiple descent: the bias support is exactly the set of columns left unmatched by some maximum matching of the bipartite variance graph (Theorem 9), and the variance diverges exactly when the residual profile loses column-side strong Hall slack, so switching configurations locate the peaks (Theorem 14). A positive-definite, non-commuting extension is left as a conjecture supported by simulations, and an explicit-z local-law extension of Alt–Erdős–Krüger is developed in Appendix C.

Significance. If the main claims hold, the paper makes a substantial contribution: it goes beyond the i.i.d./non-degenerate setting of ridgeless regression, exposes a genuinely combinatorial mechanism for multiple descent, and supplies a new local-law ingredient with explicit spectral-parameter dependence that is of independent interest. The paper is also unusually transparent: it states what is conjectural (Section 5), what requires a separate stability estimate (Remark 10), and what remains open. The deterministic equivalents are derived rather than fitted, and the graph dictionary is elegant. However, the gap between the proved fixed-(n,p) statements and the advertised asymptotic peak-location rule is load-bearing, and the simulation protocol violates the stated signal condition of the main theorem. These issues are fixable within the manuscript's scope, but they currently prevent acceptance.

major comments (3)
  1. [§4.2, Theorem 14, Corollary 60, Remark 10] The central asymptotic claim that the zero pattern fixes multiple-descent peaks is not proved. Theorem 14 is a fixed-(n,p) statement: ∂r(0) is finite iff the switching condition fails, and it certifies divergence only at exact switching configurations. Corollary 60 and Remark 10 explicitly concede that if, say, |J_1 \ J_2| = n/2 + 1 for all n, then ∂r(0) < ∞ at every n, and that proving divergence on approach would need a separate stability estimate for (4). Consequently, the finite-n sweeps in Section 6 and Appendix F, which vary p or support sizes at finite n without enforcing the exact equalities, are not consequences of Theorem 14; they show near-singular finite-n behavior. The paper needs either a stability/continuity result that lets a limiting peak be inferred from approach to a switching configuration, or the claims must be reframed as exact-square finite-n statements with 'candi
  2. [§3.5 Theorem 3, §6.1] The main theorem is proved under the signal condition ∥β∥_ℓ1 ≤ n^{1/4−δ} (Theorem 3; Corollary 4 for the rotated signal). The simulations, however, use the flat signal β = 1_p/√p, whose ℓ1 norm is √p ≍ n^{1/2} — far above the n^{1/4−δ} threshold. Thus the numerical evidence in Figures 2–5 and Appendix F does not validate the theorem in the regime in which it is stated. The paper should either run the experiments with admissible sparse or ℓ1-small signals, or extend the proof to cover the flat-signal regime; as written, the empirical support for the central claim is outside the theorem's hypotheses.
  3. [§4, Assumption 1 and Remark 5] The main theorems are stated for profiles satisfying the irreducibility/connectivity Assumption 1, but the flagship multiple-descent examples (Example 5, Section 6.1, Appendix E) are decomposable profiles. The paper says these are handled 'blockwise' (Remark 5, Appendix E), but no theorem or proof is given that the blockwise decomposition of r(0), ∂r(0), and the bias/variance equivalents is valid in the vanishing-ridge limit for the risk itself. This is likely fixable, but it needs a formal statement: either prove the decomposition as a corollary of Theorem 3, or state the blockwise version explicitly.
minor comments (5)
  1. [Theorem 3] The notation 'sup_p ∥β∥ < ∞' should presumably be 'sup_n ∥β_n∥ < ∞' over a sequence of signals; the index p is fixed in that expression. Please clarify.
  2. [Assumption 1 / Appendix D.1] Assumption 1 in the main text and Assumption D.1 in the appendix are identical but stated separately. Unify them to avoid confusion.
  3. [Figure 1 / Section 6.3] The caption says n=100 and k=5 copies, with b=2,3, but does not explain how γ=p/n is swept when n is fixed. Clarify whether p is varied or n is subsampled, so the reader can map the horizontal axis to the theoretical aspect ratio.
  4. [Section 3.1] Equation for β̂_λ is written with (1/n)Σ Xi Xi^T + λ I_p; the conventional ridge form divides by n. It is correct given the later definition of W_n, but the notation is easy to misread.
  5. [Remark 10] The list of candidate ratios in Corollary 60 is called 'candidates' only in the remark; the main text and abstract would benefit from the same caution, since Theorem 14 itself does not justify calling every listed ratio a proved peak.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the risk and peak derivations are self-contained; admitted stability/conjecture gaps are correctness limitations, not circular reductions.

full rationale

I checked every load-bearing step. Models 1 and 2 are reduced to the variance-profile Model 3 by orthogonal invariance (Lemmas 17–18), and Theorem 3's deterministic equivalents are derived from the variance-profile local law with explicit spectral-parameter dependence (Appendix C), solving the fixed-point system (4) rather than fitting it. Theorem 9 identifies the support of r(0) with the Dulmage–Mendelsohn column set J_S by a convex-potential/matching proof, not by definition; Theorem 14 derives divergence of ∂r(0) from failure of the column-side strong Hall property for the residual profile, again via Lemma 31. No fitted parameter is renamed as a prediction, and no input quantity is defined in terms of an output claim. Self-citations ([32], [40]) occur in related work, in the data-augmentation motivation, and in the phrase 'standing assumption of [4, 32]' for Assumption 1, but are not load-bearing: the local law is taken from external [4] and extended in the paper, and the matching theorems are classical. The paper's own limitations—Remark 10/Corollary 60 stating that exact switching equalities are required and that a stability estimate for (4) is needed for approach-to-peak divergence, and Section 5's explicitly conjectural positive-definite claim—are proof/scope gaps rather than circular reductions. Simulations are demonstrations of the proved exact-configuration rule, not inputs to the theorems.

Assumptions & free parameters 0 free parameters · 6 assumptions · 0 invented entities

No new particles, mediators, or forces are introduced. The variance profile and the Dulmage–Mendelsohn decomposition are existing mathematical objects. All remaining assumptions are stated as explicit model restrictions, and the positive-definite case is honestly labeled a conjecture.

assumptions (6)
  • domain assumption Model 1 and Model 2 covariances are simultaneously diagonalizable, and the design is Gaussian for the reduction to an independent-entry variance profile.
    Invoked in Lemmas 17–18 and Section 3.6; without a shared eigenbasis the profile S does not exist and the matching rule has no object to act on.
  • domain assumption Assumption 1: the variance profile is flat and connected through power bounds on SS^T and S^T S.
    Stated in Section 3.4; the main theorems require this irreducibility, with reducible cases handled only blockwise.
  • domain assumption The test point X_new is drawn from a uniform mixture of the training marginals.
    Equation (1) and Appendix A; this defines the prediction risk and generalizes the i.i.d. test mechanism, but a different test distribution would change the risk formulas.
  • standard math Existence and uniqueness of the quadratic-vector-equation solution, and the local law of Alt–Erdős–Krüger [4] extended with explicit z-dependence in Appendix C.
    The deterministic equivalents in Theorem 3 rest on this random-matrix machinery.
  • standard math Classical matching facts: Hall's theorem, Berge's theorem, Dulmage–Mendelsohn decomposition, and strong Hall properties.
    Used in Lemmas 8, 13, 24, 30 and the proof of Theorems 9 and 14.
  • domain assumption A vanishing ridge level λ_n ↓ 0 with λ_n in a polynomially decaying window approximates the ridgeless interpolator; the exact λ=0 limit is not established.
    Remark 2 explicitly leaves the exact ridgeless limit open; all theorems are stated for the vanishing-ridge sequence.

how reviews work

0 comments
Cite this review

Pith. "Pith review of The Zero Pattern of a Design Matrix Drives Multiple Descent in Over-parameterized Regression." pith.science (2026). https://pith.science/paper/QEEGFMEF

@misc{pith2026260724041,
  author       = {Pith},
  title        = {Pith review of: The Zero Pattern of a Design Matrix Drives Multiple Descent in Over-parameterized Regression},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/QEEGFMEF}},
  note         = {Machine review of arXiv:2607.24041}
}
read the original abstract

Over-parameterized linear regression has been widely studied over the last decade. However, most existing works assume that the covariates are independent and that their covariance matrices are non-degenerate. In this paper, we relax both assumptions and derive deterministic equivalents for the prediction risk in a vanishing-ridge regime. We show that degeneracy of the covariance matrices and dependence can lead to multiple descent, and characterize where the corresponding peaks can occur. Our proofs use a novel graph representation of the variance profile. We show that maximum matchings and the Dulmage--Mendelsohn decomposition of the associated bipartite graph identify the configurations at which the variance becomes singular.

Figures

Figures reproduced from arXiv: 2607.24041 by the authors.

Figure 1
Figure 1. Multiple descent from data augmentation (Model [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Multiple descent from latent covariance support (Model [PITH_FULL_IMAGE:figures/full_fig_p019_2.png] view at source ↗
Figure 3
Figure 3. Non-Gaussian universality of multiple descent. The same sweep of Figure [PITH_FULL_IMAGE:figures/full_fig_p020_3.png] view at source ↗
Figures from the paper (12 more)
Figure 4
Figure 4. Figure 4: Positive-definite heterogeneity has only the classical peak (Model [PITH_FULL_IMAGE:figures/full_fig_p021_4.png]
Figure 5
Figure 5. Figure 5: Multiple descent from a text-embedding design. Held-out ridge test error against [PITH_FULL_IMAGE:figures/full_fig_p022_5.png]
Figure 6
Figure 6. Figure 6: Two groups of n/2 observations, n = 200. Group 1 is N (0, diag{1,. . ., 1, 0, .. . , 0}) with the top p/2 coordinates active; group 2 is N (0, diag{1, .. . , 1, 0, . .. , 0}) with the top ρp coordinates active. Test risk against p/n; each curve is one value of ρ (ratio…
Figure 7
Figure 7. Figure 7: As Figure [PITH_FULL_IMAGE:figures/full_fig_p076_7.png]
Figure 8
Figure 8. Figure 8: Three groups of n/3 observations, n = 300. Group 1 activates the top p/3 coordinates, group 2 the bottom p/2, and group 3 a middle band of width p/6 starting at coordinate ρp. Test risk against p/n; each curve is one value of ρ (pos= ρ). 2 4 6 8 10 0 2 4 6 8 eval_size=…
Figure 9
Figure 9. Figure 9: Two groups of n/2, n = 200; test risk (log10 scale) against p/n. Group 1 has covariance diag{1, .. . , 1, σ2 , .. . , σ2 } (top p/2 equal to 1); group 2 has diag{σ 2 , .. . , σ2 , 1,. . . , 1} (bottom p/4 equal to 1). Each curve is one value of σ (eval size= σ) [PITH_…
Figure 10
Figure 10. Figure 10: As Figure [PITH_FULL_IMAGE:figures/full_fig_p078_10.png]
Figure 11
Figure 11. Figure 11: Overlapping-support version of Figure [PITH_FULL_IMAGE:figures/full_fig_p078_11.png]
Figure 12
Figure 12. Figure 12: As Figure [PITH_FULL_IMAGE:figures/full_fig_p079_12.png]
Figure 13
Figure 13. Figure 13: Full-rank, non-commuting control. Two groups of [PITH_FULL_IMAGE:figures/full_fig_p079_13.png]
Figure 14
Figure 14. Figure 14: As Figure [PITH_FULL_IMAGE:figures/full_fig_p080_14.png]
Figure 15
Figure 15. Figure 15: Full-rank sanity check with varying scale. Two groups of [PITH_FULL_IMAGE:figures/full_fig_p080_15.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

60 extracted references · 4 linked inside Pith

  1. [1]

    and PENNINGTON, J

    ADLAM, B. and PENNINGTON, J. (2020). The neural tangent kernel in high dimensions: Triple descent and a multi-scale theory of generalization. InInternational Conference on Machine Learning74–84. PMLR

  2. [2]

    and KR ¨UGER, T

    AJANKI, O., ERD ˝OS, L. and KR ¨UGER, T. (2019). Quadratic vector equations on complex upper half-plane. Mem. Amer. Math. Soc.261

  3. [3]

    H., ERD ˝OS, L

    AJANKI, O. H., ERD ˝OS, L. and KR ¨UGER, T. (2017). Universality for general Wigner-type matrices. Probab. Theory Related Fields169667–727

  4. [4]

    and KR ¨UGER, T

    ALT, J., ERD ˝OS, L. and KR ¨UGER, T. (2017). Local law for random Gram matrices.Electron. J. Probab.22 1–41

  5. [5]

    L., LONG, P

    BARTLETT, P. L., LONG, P. M., LUGOSI, G. and TSIGLER, A. (2020). Benign overfitting in linear regres- sion.Proc. Natl. Acad. Sci. USA11730063–30070

  6. [6]

    L., MONTANARI, A

    BARTLETT, P. L., MONTANARI, A. and RAKHLIN, A. (2021). Deep learning: a statistical viewpoint.Acta Numer.3087–201

  7. [7]

    and MANDAL, S

    BELKIN, M., HSU, D., MA, S. and MANDAL, S. (2019). Reconciling modern machine-learning practice and the classical bias–variance trade-off.Proc. Natl. Acad. Sci. USA11615849–15854

  8. [8]

    BERGE, C. (1957). Two theorems in graph theory.Proc. Natl. Acad. Sci. USA43842–844

Show all 60 references
  1. [9]

    and FEDER, M

    BIBAS, K. and FEDER, M. (2021). Distribution free uncertainty for the minimum norm solution of over- parameterized linear regression. InWorkshop on Distribution-Free Uncertainty Quantification ICML

  2. [10]

    and MALE, C

    BIGOT, J., DABO, I.-M. and MALE, C. (2026). High-dimensional analysis of ridge regression for non-identically distributed data with a variance profile.SIAM J. Math. Data Sci.To appear. arXiv:2403.20200

  3. [11]

    BRUALDI, R. A. and SHADER, B. L. (1994). Strong Hall matrices.SIAM J. Matrix Anal. Appl.15359–365

  4. [12]

    and KARBASI, A

    CHEN, L., MIN, Y., BELKIN, M. and KARBASI, A. (2021). Multiple descent: Design your own general- ization curve.Adv. Neural Inf. Process. Syst.348898–8912

  5. [13]

    and MONTANARI, A

    CHENG, C. and MONTANARI, A. (2024). Dimension free ridge regression.Ann. Statist.522879–2912. 24

  6. [14]

    F., EDENBRANDT, A

    COLEMAN, T. F., EDENBRANDT, A. and GILBERT, J. R. (1986). Predicting fill for sparse orthogonal factorization.J. ACM33517–532

  7. [15]

    and LIAO, Z

    COUILLET, R. and LIAO, Z. (2022).Random Matrix Methods for Machine Learning. Cambridge University Press

  8. [16]

    and BIGOT, J

    DABO, I.-M. and BIGOT, J. (2025). High-dimensional ridge regression with random features for non- identically distributed data with a variance profile.arXiv preprint arXiv:2504.03035. [17]D’ASCOLI, S., SAGUN, L. and BIROLI, G. (2020). Triple descent and the two kinds of overfi...

  9. [18]

    and WAGER, S

    DOBRIBAN, E. and WAGER, S. (2018). High-dimensional asymptotics of prediction: Ridge regression and classification.Ann. Statist.46247–279

  10. [19]

    DULMAGE, A. L. and MENDELSOHN, N. S. (1958). Coverings of Bipartite Graphs.Canad. J. Math.10 517–534

  11. [20]

    and SCHWARTZ, J

    DUNFORD, N. and SCHWARTZ, J. T. (1988).Linear operators, part 1: general theory. John Wiley & Sons

  12. [21]

    ERDOS, L., KNOWLES, A., YAU, H.-T., YIN, J. et al. (2013). The local semicircle law for a general class of random matrices.Electron. J. Probab181–58

  13. [22]

    ETHAYARAJH, K. (2019). How contextual are contextualized word representations? Comparing the geome- try of BERT, ELMo, and GPT-2 embeddings. InEmpirical Methods in Natural Language Processing (EMNLP)

  14. [23]

    and LIU, T.-Y

    GAO, J., HE, D., TAN, X., QIN, T., WANG, L. and LIU, T.-Y. (2019). Representation degeneration problem in training natural language generation models. InInternational Conference on Learning Representa- tions (ICLR)

  15. [24]

    GORDON, Y. (1985). Some inequalities for Gaussian processes and applications.Israel J. Math.50265– 289

  16. [25]

    and NAJIM, J

    HACHEM, W., HARDY, A. and NAJIM, J. (2016). Large Complex Correlated Wishart Matrices: The Pearcey Kernel and Expansion at the Hard Edge.Electron. J. Probab.211–36

  17. [26]

    and NAJIM, J

    HACHEM, W., LOUBATON, P. and NAJIM, J. (2007). Deterministic equivalents for certain functionals of large random matrices.Ann. Appl. Probab.17875–930

  18. [27]

    HALL, P. (1935). On representatives of subsets.J. London Math. Soc.1026–30

  19. [28]

    and SHEN, Y

    HAN, Q. and SHEN, Y. (2023). Universality of regularized regression estimators in high dimensions.Ann. Statist.511799–1823

  20. [29]

    and TIBSHIRANI, R

    HASTIE, T., MONTANARI, A., ROSSET, S. and TIBSHIRANI, R. J. (2022). Surprises in high-dimensional ridgeless least squares interpolation.Ann. Statist.50949–986

  21. [30]

    and ROSENTHAL, R

    HE, Y., KNOWLES, A. and ROSENTHAL, R. (2018). Isotropic self-consistent equations for mean-field random matrices.Probab. Theory Related Fields171203–249

  22. [31]

    and LU, Y

    HU, H. and LU, Y. M. (2023). Universality laws for high-dimensional learning with random features.IEEE Trans. Inform. Theory691932–1964

  23. [32]

    H., ORBANZ, P

    HUANG, K. H., ORBANZ, P. and AUSTERN, M. (2026). Gaussian and non-Gaussian universality of data augmentation.Ann. Statist.Forthcoming. arXiv:2202.09134

  24. [33]

    and YIN, J

    KNOWLES, A. and YIN, J. (2017). Anisotropic local laws for random matrices.Probab. Theory Related Fields169257–352

  25. [34]

    and SANCHEZ, B

    KOBAK, D., LOMOND, J. and SANCHEZ, B. (2020). The optimal ridge penalty for real-world high- dimensional data can be zero or negative due to the implicit ridge regularization.J. Mach. Learn. Res.211–16

  26. [35]

    and SUR, P

    LAHIRY, S. and SUR, P. (2024). Universality in block dependent linear models with applications to nonlin- ear regression.IEEE Trans. Inform. Theory708975–9000

  27. [36]

    and LI, L

    LI, B., ZHOU, H., HE, J., WANG, M., YANG, Y. and LI, L. (2020). On the sentence embeddings from pre-trained language models. InEmpirical Methods in Natural Language Processing (EMNLP)9119– 9130

  28. [37]

    and ZHAI, X

    LIANG, T., RAKHLIN, A. and ZHAI, X. (2020). On the multiple descent of minimum-norm interpolants and restricted lower isometry of kernels. InConference on Learning Theory2683–2711. PMLR

  29. [38]

    and COUILLET, R

    LOUART, C. and COUILLET, R. (2021). Spectral properties of sample covariance matrices arising from ran- dom matrices with independent non identically distributed columns.arXiv preprint arXiv:2109.02644

  30. [39]

    and PLUMMER, M

    LOV ´ASZ, L. and PLUMMER, M. D. (2009).Matching theory367. American Mathematical Soc

  31. [40]

    E., HUANG, K

    MALLORY, M. E., HUANG, K. H. and AUSTERN, M. (2025). Universality of High-Dimensional Logistic Regression and a Novel CGMT under Dependence with Applications to Data Augmentation. InThe Thirty Eighth Annual Conference on Learning Theory1799–1918. PMLR

  32. [41]

    and MONTANARI, A

    MEI, S. and MONTANARI, A. (2022). The generalization error of random features regression: Precise asymptotics and the double descent curve.Comm. Pure Appl. Math.75667–766. 25

  33. [42]

    and GANGULI, S

    MEL, G. and GANGULI, S. (2021). A theory of high dimensional regression with arbitrary correlations between input features and target functions: sample complexity, multiple descent curves and a hier- archy of phase transitions. InProceedings of the 38th International Conferenc...

  34. [43]

    and CAO, Y

    MENG, X., YAO, J. and CAO, Y. (2024). Multiple descent in the multiple random feature model.J. Mach. Learn. Res.251–49

  35. [44]

    and HASSANI, H

    MONIRI, B. and HASSANI, H. (2025). Asymptotics of Linear Regression with Linearly Dependent Data. In7th Annual Learning for Dynamics\& Control Conference72–85. PMLR

  36. [45]

    and SAEED, B

    MONTANARI, A. and SAEED, B. N. (2022). Universality of empirical risk minimization. InConference on Learning Theory4310–4312. PMLR

  37. [46]

    and VISWANATH, P

    MU, J., BHAT, S. and VISWANATH, P. (2018). All-but-the-top: Simple and effective postprocessing for word representations. InInternational Conference on Learning Representations (ICLR)

  38. [47]

    and SAHAI, A

    MUTHUKUMAR, V., VODRAHALLI, K., SUBRAMANIAN, V. and SAHAI, A. (2020). Harmless interpola- tion of noisy data in regression.IEEE J. Sel. Areas Inf. Theory167–83

  39. [48]

    and SUTSKEVER, I

    NAKKIRAN, P., KAPLUN, G., BANSAL, Y., YANG, T., BARAK, B. and SUTSKEVER, I. (2021). Deep double descent: Where bigger models and more data hurt.J. Stat. Mech. Theory Exp.2021124003

  40. [49]

    NAKKIRAN, P., VENKAT, P., KAKADE, S. M. and MA, T. (2021). Optimal regularization can mitigate double descent. InInternational Conference on Learning Representations

  41. [50]

    and EDUNOV, S

    NG, N., YEE, K., BAEVSKI, A., OTT, M., AULI, M. and EDUNOV, S. (2019). Facebook FAIR’s WMT19 news translation task submission. InProceedings of the Fourth Conference on Machine Translation (WMT)

  42. [51]

    and FAN, C.-J

    POTHEN, A. and FAN, C.-J. (1990). Computing the block triangular form of a sparse matrix.ACM Trans. Math. Software16303–324

  43. [52]

    PULLEYBLANK, W. R. (1996). Matchings and extensions. InHandbook of combinatorics (vol. 1)179–232

  44. [53]

    and ROSASCO, L

    RICHARDS, D., MOURTADA, J. and ROSASCO, L. (2021). Asymptotics of ridge(less) regression under general source condition. InInternational Conference on Artificial Intelligence and Statistics3889–

  45. [54]

    and SUR, P

    SONG, Y., BHATTACHARYA, S. and SUR, P. (2024). Generalization error of min-norm interpolators in transfer learning.arXiv preprint arXiv:2406.13944

  46. [55]

    and HASSIBI, B

    THRAMPOULIDIS, C., ABBASI, E. and HASSIBI, B. (2018). Precise error analysis of regularizedM- estimators in high dimensions.IEEE Trans. Inform. Theory645592–5628

  47. [56]

    TRACY, C. A. and WIDOM, H. (1994). Level Spacing Distributions and the Bessel Kernel.Comm. Math. Phys.161289–309

  48. [57]

    and BARTLETT, P

    TSIGLER, A. and BARTLETT, P. L. (2023). Benign overfitting in ridge regression.J. Mach. Learn. Res.24 1–76

  49. [58]

    and XU, J

    WU, D. and XU, J. (2020). On the optimal weightedℓ 2 regularization in overparameterized linear regres- sion.Adv. Neural Inf. Process. Syst.3310112–10123

  50. [59]

    YIN, Y. (2020). On the singular value distribution of large-dimensional data matrices whose columns have different correlations.Statistics54353–374. Appendices The appendices are organized as follows: • Appendix A restates the setup and proves the equivalence of the models; • ...

  51. [60]

    SupposeI t is non-empty. Recall that by Lemma 24, the induced sub-matrixS(JS)satisfies the row-side strong Hall property, which implies |It|+ 1≤ j∈J S Sij >0for somei∈I t = j∈J S Sij >0for somei∈I J S (S)withN i(J S)⊆A t ≤ |At|, in which case|A t| − |It| ≥1again. In summary, w...

  52. [61]

    Equivalently the group-1block ofW n equals(en/n)fWen, with fWen in theen-sample normalization of Theorem 3; this rescaling is harmless and does not affect the switching condition. • Ifa/en→γ 1 ∈(0,∞), the block is a complete (all-positive) variance profile with entries of orde...

Pith tools

Reviewed July 31, 2026 · model on record in the stance chip above.