REVIEW 3 major objections 5 minor 60 references
The Zero Pattern of a Design Matrix Drives Multiple Descent in Over-parameterized Regression
T0 review · 3 major / 5 minor · reviewed 2026-07-31 · deepseek-v4-flash
Pith's one-line read Multiple-descent peaks in over-parameterized regression are set by the zero pattern of the design's variance profile — a matching rule on a bipartite graph — so a single heterogeneous or dependent design can show several risk peaks with not
desk verdict A genuinely new combinatorial mechanism for multiple descent, with a real gap the authors flag themselves: the theorem covers exact square configurations, while the advertised limiting peaks are left as candidates. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The variance profile S, an n×p matrix of per-entry variances, with its bipartite variance graph G_S; maximum matchings of G_S (whose size is the structural rank); the Dulmage–Mendelsohn decomposition, which yields the degenerate column set J_S; and the strong Hall properties (row-side for the bias, column-side for the residual). The fixed-point system 1 = r_l(λ) + Σ_i S_il r_l(λ)/(λ + Σ_j S_ij r_j(λ)) carries the risk; r(0) and ∂r(0) at the hard edge feed the bias and variance. A re-proved local law tracking the spectral parameter down to zero makes the vanishing-ridge limit tractable.
What would settle it
Take the Section 6.2 setup but push anisotropy to an extreme (e.g., eigenvalues 10 and 10^-3 with non-commuting rotations): the paper's conjecture says the risk still peaks only at p/n = 1. Any second peak away from γ = 1 would refute the dichotomy that zeros, not scale ratios, drive multiple descent. Conversely, the two-group near-miss of Remark 10 (ratio approaching a candidate equality without ever attaining it) tests whether divergence requires the exact square configuration.
Extended reading notes
Core claim
On the paper's own terms: for a variance-profile Gaussian design with vanishing ridge penalty, the limiting prediction risk is governed by a fixed-point vector r(λ): bias mass sits on coordinates with r_l(0) > 0, and the variance is a weighted sum of derivatives ∂r_l(0). The paper identifies r_l(0) > 0 with membership in the degenerate column set J_S of the variance graph (the set some maximum matching leaves unmatched), and proves the variance diverges precisely at switching configurations where the residual submatrix fails the column-side strong Hall property. Consequently the number and location of multiple-descent peaks are fixed by the support pattern of the design — a combinatorial inv
Load-bearing premise
The reduction to a variance profile — and hence the entire matching rule — requires the design's covariance matrices to share a common eigenbasis; without a shared basis there is no S for the graph to act on, and the positive-definite non-commuting case is left as a conjecture.
Editorial extensions
If this is right
- A single rank-deficient heterogeneous design, or a single dependent design from data augmentation, already produces multiple descent; the peaks sit where the support rule predicts (Sections 6.1 and 6.3).
- Real degenerate designs — a pretrained transformer's token-embedding block augmented with Gaussian columns — exhibit double descent matching the theory (Section 6.4).
- Decomposable profiles split blockwise, so the overall risk is the sum of block risks: each block contributes its own candidate peaks, which merge when thresholds coincide.
- The deterministic equivalents extend beyond Gaussian entries to any independent entries with matching variance profile (Remark 1); heavy-tailed simulations reproduce the Gaussian peak structure.
- Positive-definite, shared-eigenbasis covariances are non-degenerate and show only the classical peak at p/n = 1; the non-commuting case is conjectured to behave the same.
Reading between the lines
- The matching rule gives a pre-training diagnostic: estimate the support of the covariance profile, run a maximum matching, and predict where risk will spike before fitting.
- The exact-zero theory plus the paper's near-zero simulations suggest a continuity question: do peak locations shift continuously as zero eigenvalues are replaced by σ→0 scales? A quantitative version would extend the theory to approximate degeneracy.
- The mechanism ties double descent to structural-rank theory; the same Dulmage–Mendelsohn language may transfer to minimum-norm interpolation in classification or to kernel models with degenerate kernel matrices.
- If the Section 5 conjecture holds, multiple descent becomes a signal of exact rank deficiency in feature matrices — information about the data, not about the model family.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies vanishing-ridge linear regression for Gaussian designs in which observations are heterogeneous or dependent, with covariance matrices that are simultaneously diagonalizable and possibly rank-deficient. It reduces both Model 1 (heterogeneous) and Model 2 (finite-rank dependent) to a variance-profile model, and derives deterministic equivalents for the bias and variance of the prediction risk (Theorem 3, Corollary 4). The paper's central structural claim is that the zero pattern of the variance profile S controls multiple descent: the bias support is exactly the set of columns left unmatched by some maximum matching of the bipartite variance graph (Theorem 9), and the variance diverges exactly when the residual profile loses column-side strong Hall slack, so switching configurations locate the peaks (Theorem 14). A positive-definite, non-commuting extension is left as a conjecture supported by simulations, and an explicit-z local-law extension of Alt–Erdős–Krüger is developed in Appendix C.
Significance. If the main claims hold, the paper makes a substantial contribution: it goes beyond the i.i.d./non-degenerate setting of ridgeless regression, exposes a genuinely combinatorial mechanism for multiple descent, and supplies a new local-law ingredient with explicit spectral-parameter dependence that is of independent interest. The paper is also unusually transparent: it states what is conjectural (Section 5), what requires a separate stability estimate (Remark 10), and what remains open. The deterministic equivalents are derived rather than fitted, and the graph dictionary is elegant. However, the gap between the proved fixed-(n,p) statements and the advertised asymptotic peak-location rule is load-bearing, and the simulation protocol violates the stated signal condition of the main theorem. These issues are fixable within the manuscript's scope, but they currently prevent acceptance.
major comments (3)
- [§4.2, Theorem 14, Corollary 60, Remark 10] The central asymptotic claim that the zero pattern fixes multiple-descent peaks is not proved. Theorem 14 is a fixed-(n,p) statement: ∂r(0) is finite iff the switching condition fails, and it certifies divergence only at exact switching configurations. Corollary 60 and Remark 10 explicitly concede that if, say, |J_1 \ J_2| = n/2 + 1 for all n, then ∂r(0) < ∞ at every n, and that proving divergence on approach would need a separate stability estimate for (4). Consequently, the finite-n sweeps in Section 6 and Appendix F, which vary p or support sizes at finite n without enforcing the exact equalities, are not consequences of Theorem 14; they show near-singular finite-n behavior. The paper needs either a stability/continuity result that lets a limiting peak be inferred from approach to a switching configuration, or the claims must be reframed as exact-square finite-n statements with 'candi
- [§3.5 Theorem 3, §6.1] The main theorem is proved under the signal condition ∥β∥_ℓ1 ≤ n^{1/4−δ} (Theorem 3; Corollary 4 for the rotated signal). The simulations, however, use the flat signal β = 1_p/√p, whose ℓ1 norm is √p ≍ n^{1/2} — far above the n^{1/4−δ} threshold. Thus the numerical evidence in Figures 2–5 and Appendix F does not validate the theorem in the regime in which it is stated. The paper should either run the experiments with admissible sparse or ℓ1-small signals, or extend the proof to cover the flat-signal regime; as written, the empirical support for the central claim is outside the theorem's hypotheses.
- [§4, Assumption 1 and Remark 5] The main theorems are stated for profiles satisfying the irreducibility/connectivity Assumption 1, but the flagship multiple-descent examples (Example 5, Section 6.1, Appendix E) are decomposable profiles. The paper says these are handled 'blockwise' (Remark 5, Appendix E), but no theorem or proof is given that the blockwise decomposition of r(0), ∂r(0), and the bias/variance equivalents is valid in the vanishing-ridge limit for the risk itself. This is likely fixable, but it needs a formal statement: either prove the decomposition as a corollary of Theorem 3, or state the blockwise version explicitly.
minor comments (5)
- [Theorem 3] The notation 'sup_p ∥β∥ < ∞' should presumably be 'sup_n ∥β_n∥ < ∞' over a sequence of signals; the index p is fixed in that expression. Please clarify.
- [Assumption 1 / Appendix D.1] Assumption 1 in the main text and Assumption D.1 in the appendix are identical but stated separately. Unify them to avoid confusion.
- [Figure 1 / Section 6.3] The caption says n=100 and k=5 copies, with b=2,3, but does not explain how γ=p/n is swept when n is fixed. Clarify whether p is varied or n is subsampled, so the reader can map the horizontal axis to the theoretical aspect ratio.
- [Section 3.1] Equation for β̂_λ is written with (1/n)Σ Xi Xi^T + λ I_p; the conventional ridge form divides by n. It is correct given the later definition of W_n, but the notation is easy to misread.
- [Remark 10] The list of candidate ratios in Corollary 60 is called 'candidates' only in the remark; the main text and abstract would benefit from the same caution, since Theorem 14 itself does not justify calling every listed ratio a proved peak.
Circularity Check
No circularity: the risk and peak derivations are self-contained; admitted stability/conjecture gaps are correctness limitations, not circular reductions.
full rationale
I checked every load-bearing step. Models 1 and 2 are reduced to the variance-profile Model 3 by orthogonal invariance (Lemmas 17–18), and Theorem 3's deterministic equivalents are derived from the variance-profile local law with explicit spectral-parameter dependence (Appendix C), solving the fixed-point system (4) rather than fitting it. Theorem 9 identifies the support of r(0) with the Dulmage–Mendelsohn column set J_S by a convex-potential/matching proof, not by definition; Theorem 14 derives divergence of ∂r(0) from failure of the column-side strong Hall property for the residual profile, again via Lemma 31. No fitted parameter is renamed as a prediction, and no input quantity is defined in terms of an output claim. Self-citations ([32], [40]) occur in related work, in the data-augmentation motivation, and in the phrase 'standing assumption of [4, 32]' for Assumption 1, but are not load-bearing: the local law is taken from external [4] and extended in the paper, and the matching theorems are classical. The paper's own limitations—Remark 10/Corollary 60 stating that exact switching equalities are required and that a stability estimate for (4) is needed for approach-to-peak divergence, and Section 5's explicitly conjectural positive-definite claim—are proof/scope gaps rather than circular reductions. Simulations are demonstrations of the proved exact-configuration rule, not inputs to the theorems.
Assumptions & free parameters
assumptions (6)
- domain assumption Model 1 and Model 2 covariances are simultaneously diagonalizable, and the design is Gaussian for the reduction to an independent-entry variance profile.
- domain assumption Assumption 1: the variance profile is flat and connected through power bounds on SS^T and S^T S.
- domain assumption The test point X_new is drawn from a uniform mixture of the training marginals.
- standard math Existence and uniqueness of the quadratic-vector-equation solution, and the local law of Alt–Erdős–Krüger [4] extended with explicit z-dependence in Appendix C.
- standard math Classical matching facts: Hall's theorem, Berge's theorem, Dulmage–Mendelsohn decomposition, and strong Hall properties.
- domain assumption A vanishing ridge level λ_n ↓ 0 with λ_n in a polynomially decaying window approximates the ridgeless interpolator; the exact λ=0 limit is not established.
Cite this review
Pith. "Pith review of The Zero Pattern of a Design Matrix Drives Multiple Descent in Over-parameterized Regression." pith.science (2026). https://pith.science/paper/QEEGFMEF
@misc{pith2026260724041,
author = {Pith},
title = {Pith review of: The Zero Pattern of a Design Matrix Drives Multiple Descent in Over-parameterized Regression},
year = {2026},
howpublished = {\url{https://pith.science/paper/QEEGFMEF}},
note = {Machine review of arXiv:2607.24041}
}
read the original abstract
Over-parameterized linear regression has been widely studied over the last decade. However, most existing works assume that the covariates are independent and that their covariance matrices are non-degenerate. In this paper, we relax both assumptions and derive deterministic equivalents for the prediction risk in a vanishing-ridge regime. We show that degeneracy of the covariance matrices and dependence can lead to multiple descent, and characterize where the corresponding peaks can occur. Our proofs use a novel graph representation of the variance profile. We show that maximum matchings and the Dulmage--Mendelsohn decomposition of the associated bipartite graph identify the configurations at which the variance becomes singular.
Figures
Figures from the paper (12 more)
Reference graph
Works this paper leans on
-
[1]
and PENNINGTON, J
ADLAM, B. and PENNINGTON, J. (2020). The neural tangent kernel in high dimensions: Triple descent and a multi-scale theory of generalization. InInternational Conference on Machine Learning74–84. PMLR
2020
-
[2]
and KR ¨UGER, T
AJANKI, O., ERD ˝OS, L. and KR ¨UGER, T. (2019). Quadratic vector equations on complex upper half-plane. Mem. Amer. Math. Soc.261
2019
-
[3]
H., ERD ˝OS, L
AJANKI, O. H., ERD ˝OS, L. and KR ¨UGER, T. (2017). Universality for general Wigner-type matrices. Probab. Theory Related Fields169667–727
2017
-
[4]
and KR ¨UGER, T
ALT, J., ERD ˝OS, L. and KR ¨UGER, T. (2017). Local law for random Gram matrices.Electron. J. Probab.22 1–41
2017
-
[5]
L., LONG, P
BARTLETT, P. L., LONG, P. M., LUGOSI, G. and TSIGLER, A. (2020). Benign overfitting in linear regres- sion.Proc. Natl. Acad. Sci. USA11730063–30070
2020
-
[6]
L., MONTANARI, A
BARTLETT, P. L., MONTANARI, A. and RAKHLIN, A. (2021). Deep learning: a statistical viewpoint.Acta Numer.3087–201
2021
-
[7]
and MANDAL, S
BELKIN, M., HSU, D., MA, S. and MANDAL, S. (2019). Reconciling modern machine-learning practice and the classical bias–variance trade-off.Proc. Natl. Acad. Sci. USA11615849–15854
2019
-
[8]
BERGE, C. (1957). Two theorems in graph theory.Proc. Natl. Acad. Sci. USA43842–844
1957
Show all 60 references
-
[9]
and FEDER, M
BIBAS, K. and FEDER, M. (2021). Distribution free uncertainty for the minimum norm solution of over- parameterized linear regression. InWorkshop on Distribution-Free Uncertainty Quantification ICML
2021
-
[10]
and MALE, C
BIGOT, J., DABO, I.-M. and MALE, C. (2026). High-dimensional analysis of ridge regression for non-identically distributed data with a variance profile.SIAM J. Math. Data Sci.To appear. arXiv:2403.20200
2026 arXiv
-
[11]
BRUALDI, R. A. and SHADER, B. L. (1994). Strong Hall matrices.SIAM J. Matrix Anal. Appl.15359–365
1994
-
[12]
and KARBASI, A
CHEN, L., MIN, Y., BELKIN, M. and KARBASI, A. (2021). Multiple descent: Design your own general- ization curve.Adv. Neural Inf. Process. Syst.348898–8912
2021
-
[13]
and MONTANARI, A
CHENG, C. and MONTANARI, A. (2024). Dimension free ridge regression.Ann. Statist.522879–2912. 24
2024
-
[14]
F., EDENBRANDT, A
COLEMAN, T. F., EDENBRANDT, A. and GILBERT, J. R. (1986). Predicting fill for sparse orthogonal factorization.J. ACM33517–532
1986
-
[15]
and LIAO, Z
COUILLET, R. and LIAO, Z. (2022).Random Matrix Methods for Machine Learning. Cambridge University Press
2022
-
[16]
and BIGOT, J
DABO, I.-M. and BIGOT, J. (2025). High-dimensional ridge regression with random features for non- identically distributed data with a variance profile.arXiv preprint arXiv:2504.03035. [17]D’ASCOLI, S., SAGUN, L. and BIROLI, G. (2020). Triple descent and the two kinds of overfi...
2025 arXiv
-
[18]
and WAGER, S
DOBRIBAN, E. and WAGER, S. (2018). High-dimensional asymptotics of prediction: Ridge regression and classification.Ann. Statist.46247–279
2018
-
[19]
DULMAGE, A. L. and MENDELSOHN, N. S. (1958). Coverings of Bipartite Graphs.Canad. J. Math.10 517–534
1958
-
[20]
and SCHWARTZ, J
DUNFORD, N. and SCHWARTZ, J. T. (1988).Linear operators, part 1: general theory. John Wiley & Sons
1988
-
[21]
ERDOS, L., KNOWLES, A., YAU, H.-T., YIN, J. et al. (2013). The local semicircle law for a general class of random matrices.Electron. J. Probab181–58
2013
-
[22]
ETHAYARAJH, K. (2019). How contextual are contextualized word representations? Comparing the geome- try of BERT, ELMo, and GPT-2 embeddings. InEmpirical Methods in Natural Language Processing (EMNLP)
2019
-
[23]
and LIU, T.-Y
GAO, J., HE, D., TAN, X., QIN, T., WANG, L. and LIU, T.-Y. (2019). Representation degeneration problem in training natural language generation models. InInternational Conference on Learning Representa- tions (ICLR)
2019
-
[24]
GORDON, Y. (1985). Some inequalities for Gaussian processes and applications.Israel J. Math.50265– 289
1985
-
[25]
and NAJIM, J
HACHEM, W., HARDY, A. and NAJIM, J. (2016). Large Complex Correlated Wishart Matrices: The Pearcey Kernel and Expansion at the Hard Edge.Electron. J. Probab.211–36
2016
-
[26]
and NAJIM, J
HACHEM, W., LOUBATON, P. and NAJIM, J. (2007). Deterministic equivalents for certain functionals of large random matrices.Ann. Appl. Probab.17875–930
2007
-
[27]
HALL, P. (1935). On representatives of subsets.J. London Math. Soc.1026–30
1935
-
[28]
and SHEN, Y
HAN, Q. and SHEN, Y. (2023). Universality of regularized regression estimators in high dimensions.Ann. Statist.511799–1823
2023
-
[29]
and TIBSHIRANI, R
HASTIE, T., MONTANARI, A., ROSSET, S. and TIBSHIRANI, R. J. (2022). Surprises in high-dimensional ridgeless least squares interpolation.Ann. Statist.50949–986
2022
-
[30]
and ROSENTHAL, R
HE, Y., KNOWLES, A. and ROSENTHAL, R. (2018). Isotropic self-consistent equations for mean-field random matrices.Probab. Theory Related Fields171203–249
2018
-
[31]
and LU, Y
HU, H. and LU, Y. M. (2023). Universality laws for high-dimensional learning with random features.IEEE Trans. Inform. Theory691932–1964
2023
-
[32]
H., ORBANZ, P
HUANG, K. H., ORBANZ, P. and AUSTERN, M. (2026). Gaussian and non-Gaussian universality of data augmentation.Ann. Statist.Forthcoming. arXiv:2202.09134
2026
-
[33]
and YIN, J
KNOWLES, A. and YIN, J. (2017). Anisotropic local laws for random matrices.Probab. Theory Related Fields169257–352
2017
-
[34]
and SANCHEZ, B
KOBAK, D., LOMOND, J. and SANCHEZ, B. (2020). The optimal ridge penalty for real-world high- dimensional data can be zero or negative due to the implicit ridge regularization.J. Mach. Learn. Res.211–16
2020
-
[35]
and SUR, P
LAHIRY, S. and SUR, P. (2024). Universality in block dependent linear models with applications to nonlin- ear regression.IEEE Trans. Inform. Theory708975–9000
2024
-
[36]
and LI, L
LI, B., ZHOU, H., HE, J., WANG, M., YANG, Y. and LI, L. (2020). On the sentence embeddings from pre-trained language models. InEmpirical Methods in Natural Language Processing (EMNLP)9119– 9130
2020
-
[37]
and ZHAI, X
LIANG, T., RAKHLIN, A. and ZHAI, X. (2020). On the multiple descent of minimum-norm interpolants and restricted lower isometry of kernels. InConference on Learning Theory2683–2711. PMLR
2020
-
[38]
and COUILLET, R
LOUART, C. and COUILLET, R. (2021). Spectral properties of sample covariance matrices arising from ran- dom matrices with independent non identically distributed columns.arXiv preprint arXiv:2109.02644
2021 arXiv
-
[39]
and PLUMMER, M
LOV ´ASZ, L. and PLUMMER, M. D. (2009).Matching theory367. American Mathematical Soc
2009
-
[40]
E., HUANG, K
MALLORY, M. E., HUANG, K. H. and AUSTERN, M. (2025). Universality of High-Dimensional Logistic Regression and a Novel CGMT under Dependence with Applications to Data Augmentation. InThe Thirty Eighth Annual Conference on Learning Theory1799–1918. PMLR
2025
-
[41]
and MONTANARI, A
MEI, S. and MONTANARI, A. (2022). The generalization error of random features regression: Precise asymptotics and the double descent curve.Comm. Pure Appl. Math.75667–766. 25
2022
-
[42]
and GANGULI, S
MEL, G. and GANGULI, S. (2021). A theory of high dimensional regression with arbitrary correlations between input features and target functions: sample complexity, multiple descent curves and a hier- archy of phase transitions. InProceedings of the 38th International Conferenc...
2021
-
[43]
and CAO, Y
MENG, X., YAO, J. and CAO, Y. (2024). Multiple descent in the multiple random feature model.J. Mach. Learn. Res.251–49
2024
-
[44]
and HASSANI, H
MONIRI, B. and HASSANI, H. (2025). Asymptotics of Linear Regression with Linearly Dependent Data. In7th Annual Learning for Dynamics\& Control Conference72–85. PMLR
2025
-
[45]
and SAEED, B
MONTANARI, A. and SAEED, B. N. (2022). Universality of empirical risk minimization. InConference on Learning Theory4310–4312. PMLR
2022
-
[46]
and VISWANATH, P
MU, J., BHAT, S. and VISWANATH, P. (2018). All-but-the-top: Simple and effective postprocessing for word representations. InInternational Conference on Learning Representations (ICLR)
2018
-
[47]
and SAHAI, A
MUTHUKUMAR, V., VODRAHALLI, K., SUBRAMANIAN, V. and SAHAI, A. (2020). Harmless interpola- tion of noisy data in regression.IEEE J. Sel. Areas Inf. Theory167–83
2020
-
[48]
and SUTSKEVER, I
NAKKIRAN, P., KAPLUN, G., BANSAL, Y., YANG, T., BARAK, B. and SUTSKEVER, I. (2021). Deep double descent: Where bigger models and more data hurt.J. Stat. Mech. Theory Exp.2021124003
2021
-
[49]
NAKKIRAN, P., VENKAT, P., KAKADE, S. M. and MA, T. (2021). Optimal regularization can mitigate double descent. InInternational Conference on Learning Representations
2021
-
[50]
and EDUNOV, S
NG, N., YEE, K., BAEVSKI, A., OTT, M., AULI, M. and EDUNOV, S. (2019). Facebook FAIR’s WMT19 news translation task submission. InProceedings of the Fourth Conference on Machine Translation (WMT)
2019
-
[51]
and FAN, C.-J
POTHEN, A. and FAN, C.-J. (1990). Computing the block triangular form of a sparse matrix.ACM Trans. Math. Software16303–324
1990
-
[52]
PULLEYBLANK, W. R. (1996). Matchings and extensions. InHandbook of combinatorics (vol. 1)179–232
1996
-
[53]
and ROSASCO, L
RICHARDS, D., MOURTADA, J. and ROSASCO, L. (2021). Asymptotics of ridge(less) regression under general source condition. InInternational Conference on Artificial Intelligence and Statistics3889–
2021
-
[54]
and SUR, P
SONG, Y., BHATTACHARYA, S. and SUR, P. (2024). Generalization error of min-norm interpolators in transfer learning.arXiv preprint arXiv:2406.13944
2024 arXiv
-
[55]
and HASSIBI, B
THRAMPOULIDIS, C., ABBASI, E. and HASSIBI, B. (2018). Precise error analysis of regularizedM- estimators in high dimensions.IEEE Trans. Inform. Theory645592–5628
2018
-
[56]
TRACY, C. A. and WIDOM, H. (1994). Level Spacing Distributions and the Bessel Kernel.Comm. Math. Phys.161289–309
1994
-
[57]
and BARTLETT, P
TSIGLER, A. and BARTLETT, P. L. (2023). Benign overfitting in ridge regression.J. Mach. Learn. Res.24 1–76
2023
-
[58]
and XU, J
WU, D. and XU, J. (2020). On the optimal weightedℓ 2 regularization in overparameterized linear regres- sion.Adv. Neural Inf. Process. Syst.3310112–10123
2020
-
[59]
YIN, Y. (2020). On the singular value distribution of large-dimensional data matrices whose columns have different correlations.Statistics54353–374. Appendices The appendices are organized as follows: • Appendix A restates the setup and proves the equivalence of the models; • ...
2020
-
[60]
SupposeI t is non-empty. Recall that by Lemma 24, the induced sub-matrixS(JS)satisfies the row-side strong Hall property, which implies |It|+ 1≤ j∈J S Sij >0for somei∈I t = j∈J S Sij >0for somei∈I J S (S)withN i(J S)⊆A t ≤ |At|, in which case|A t| − |It| ≥1again. In summary, w...
-
[61]
Equivalently the group-1block ofW n equals(en/n)fWen, with fWen in theen-sample normalization of Theorem 3; this rescaling is harmless and does not affect the switching condition. • Ifa/en→γ 1 ∈(0,∞), the block is a complete (all-positive) variance profile with entries of orde...
Reviewed July 31, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.