Pith. sign in

REVIEW 3 major objections 4 minor 1 cited by

Ridge-Regularized Largest Root Test For High-Dimensional General Linear Hypotheses

T0 review · 3 major / 4 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read Adding a ridge term keeps Roy's largest-root test valid even when the residual covariance matrix is singular.

desk verdict The ridge-regularized largest root result is a genuine advance, but the advertised finite-moment Tracy–Widom theorem is not fully proved; the paper deserves review, with the non-Gaussian gap addressed. read the letter →

arxiv 2504.15510 v3 pith:MY7ZPLGC submitted 2025-04-22 stat.ME

classification stat.ME MSC 62H1560B2062H1062J05
keywords ridgeregularizationRoy'slargestroottestTracy-Widomdistributiongenerallinearhypotheseshigh-dimensionalmultivariateregressionF-matrixrandommatrixtheory
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper aims to make Roy's largest-root test usable when the number of response variables p is comparable to, or larger than, the residual degrees of freedom n2, a setting where the classical F-matrix $W1W2^{-1}$ is ill-conditioned or undefined. Its proposal is to replace $W2^{-1}$ by (W2 + λIp)^-1, and its main theorem states that the largest eigenvalue of this regularized ratio, after centering and scaling by explicit constants Θ1 and Θ2, converges in distribution to the Tracy-Widom law of type 1. If true, this gives calibrated p-values for a broad class of linear hypotheses, such as MANOVA, joint significance of predictors, and tests for trends or seasonal effects, using only eigenvalue information from W2 to estimate the centering and scaling. The paper also analyzes power under low-rank alternatives, proposes a data-driven choice of λ, and demonstrates the procedure on brain-volume and behavioral data from the Human Connectome Project.

What carries the argument

The machine at the center is the ridge-regularized F-matrix $F_\lambda = W_1(W_2+\lambda I_p)^{-1}$ and its symmetric counterpart $\tilde F_\lambda=U_1^TZ^TG_\lambda^{-1}ZU_1$, which shares the same nonzero eigenvalues. The argument is carried by the generalized Marchenko-Pastur equation $$\phi=\int \frac{\tau\,dF_{\Sigma_\infty}(\tau)}{\tau[1+\gamma_2\phi]^{-1}-z}+\$\lambda$,$$ whose solution $s(x)$ near the left edge $\rho$ of $G_\lambda$'s limiting spectrum controls everything: the edge parameter β solves $\beta^2 s'(\beta)=1/\gamma_1$, and the centering and scaling are $\Theta_1=(1+\gamma_1\beta s(\beta))/\beta$ and $\Theta_2=(\gamma_1^3 s''(\beta)/2+\gamma_1^2/\beta^3)^{1/3}$. Around this core, the paper develops a local law for the resolvent of $G_\lambda$, uses edge universality to transfer Gaussian Tracy-Widom fluctuations to finite-moment errors, and estimates Θ1 and Θ2 by discretizing the population spectral distribution through a linear program and then solving an ODE, using only the eigenvalues of W2.

What would settle it

Simulate the null distribution with t-distributed errors with 4 degrees of freedom, take n2=500, p=2500, n1 in {100,250,500}, and λ in {0.5,1,1.5}, then compare the empirical 95th percentile of $p^{2/3}(\ell_{\max}(F_\lambda)-\hat\Theta_1)/\hat\Theta_2$ with the 95th percentile of TW1 across many replications; a systematic exceedance beyond Monte Carlo error as p grows would contradict the finite-moment universality claim.

Watch

Extended reading notes

Core claim

The central discovery is Theorem 2.2: under the regime C1-C6 and for any fixed λ>0, $$$p^{{2/3}}$\Theta_2(\$\lambda$)^{-1}\big(\ell_{\max}(F_\$\lambda$)-\Theta_1(\$\lambda$)\big)\Rightarrow \mathrm{TW}_1,$$ where $F_\lambda=W_1(W_2+\lambda I_p)^{-1}$, $\mathrm{TW}_1$ is the Tracy-Widom distribution of type 1, and $\Theta_1,\Theta_2$ are explicit functions of an edge parameter β defined through the generalized Marchenko-Pastur equation for the regularized matrix $G_\lambda=ZU_2U_2^TZ^T+\lambda\Sigma_p^{-1}$. The theorem covers the case $p>n_2$ where W2 is singular and assumes only finite moments for the error entries. It also yields the joint Tracy-Widom GOE fluctuations of the first k largest eigenvalues. The proof proceeds through a strong local law for Gλ together with a Green-function comparison that transplants the Gaussian edge fluctuations to general error distributions.

Load-bearing premise

The theorem's promise of universality under only finite moments of the error entries rests on a strong local-law estimate for $G_\lambda$; for non-Gaussian errors the paper gives an outline referencing prior methods rather than a complete proof, so the advertised scope depends on that outline succeeding.

Editorial extensions

If this is right

  • General linear hypotheses can be tested when p exceeds n2 without imputing or inverting a singular covariance estimate, because the ridge term stabilizes W2.
  • Calibrated p-values would require no distributional assumption beyond finite moments: the centering and scaling are estimated from the eigenvalues of W2, and the reference law is the fixed Tracy-Widom type-1 distribution.
  • The test is designed for concentrated alternatives: under low-rank signals the power tends to 1, and λ can be chosen by maximizing an estimated signal-to-noise ratio rather than by cross-validation.
  • A data-driven λ choice changes the null distribution only negligibly in the paper's simulations, so the procedure can be run with λ selected on the same data, with data splitting suggested as extra protection.
  • The joint distribution of the top k eigenvalues is also Tracy-Widom/GOE, so inference using the first several roots, for example testing multiple signal directions, is a direct corollary.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the non-Gaussian local law is completed, the same proof scheme should transfer to other regularized statistics, such as ridge-regularized canonical correlations or partial least squares, because only the linearizing block structure and edge stability are used.
  • The estimation pipeline, discretize the population spectrum, fit weights by linear programming, then solve an ODE for s(x), is a reusable template for estimating edge parameters of other random matrix models, not just the ridge F-test.
  • One testable extension is a theoretical account of data-driven λ selection: the paper's simulation evidence that double-dipping is negligible could become a theorem under suitable convergence rates for λ̂ toward its limiting maximizer.
  • In the discrete-edge case where γ2ωmax>1, the smallest eigenvalue of Gλ sits exactly at λ/σmax; understanding how close the Tracy-Widom approximation remains there is a natural robustness check for covariance structures with spikes, such as factor models.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper proposes a ridge-regularized Roy's largest-root test for high-dimensional general linear hypotheses. For a fixed λ>0, the test statistic is the largest eigenvalue of F_λ = W_1(W_2+λI_p)^{-1}, where W_1 and W_2 are the hypothesis and residual sum-of-squares matrices of a multivariate linear model. The main theoretical result, Theorem 2.2, asserts that after centering and scaling by quantities Θ_1(λ) and Θ_2(λ) defined through a generalized Marchenko–Pastur equation, the statistic converges to the Tracy–Widom type-1 law under conditions C1–C6, assuming only finite-moment conditions on the errors. The paper also develops a linear-programming-based estimation procedure for Θ_1 and Θ_2 from the eigenvalues of W_2, proves consistency of this procedure under certain grid conditions, analyzes power under low-rank alternatives, reports simulation studies, and applies the method to Human Connectome Project data.

Significance. If fully established, the result would be a useful extension of Roy's largest-root testing to the regime where the number of responses and the number of tested hypotheses are comparable to the sample size, including cases with p>n_2, which classical F-tests cannot handle. The proposed regularization and the accompanying estimation algorithm for the centering and scaling constants are genuinely new and are backed by substantial numerical work. The simulations show good null coverage in many settings and clear power advantages over the compared ridge-LRT and projection-LRT baselines. However, the advertised finite-moment universality is not actually proved: the local law for G_λ is proved only under Gaussianity in Section S.5, and the non-Gaussian extension is deferred to an outline in Section S.6. In addition, the consistency of the parameter estimators is guaranteed only under grid conditions that are shown to hold for large γ_1, while some simulations (notably Toeplitz, γ_2=5) show relative estimation errors for Θ_2 that are far from negligible. The data-driven selection of λ is acknowledged to be potentially circular but is left without a theoretical justification.

major comments (3)
  1. [S.6, Step 3; Theorem 2.2; Condition C2] Theorem 2.2 is stated under the finite-moment condition C2, but the proof of the non-Gaussian Tracy–Widom limit is not supplied. Section S.5 explicitly proves the strong local law for G_λ only under SC1 (Gaussian Z) and states that the non-Gaussian extension 'can be completed following the strategy in Sections 7–10 of Knowles and Yin (2017), which is however beyond our scope'. Section S.6, Step 3, says the non-Gaussian local law follows by 'closely following' Sections 9.1, 9.2 and 10 of Han et al. (2018) with 'only minor changes', and the key estimate, Lemma S.6.1, is stated but not proved. This gap is load-bearing because the abstract's claim 'assuming only finite-moment conditions' rests entirely on this transfer. The transfer is not automatic: for non-Gaussian Z the blocks ZU_1 and ZU_2 in the linearizing matrix H(z) are dependent, whereas the Gaussian proof exploits their independence, and the Han et al. (2018) argument is for the λ=0 case where the central block is zero. The authors should either provide a complete proof of the non-Gaussian local law and Green-function comparison for the full H(z), or state Theorem 2.2 under Gaussian errors and describe the finite-moment version as a conjecture supported by simulations.
  2. [Theorem 3.2, Lemma 3.4, Table 5.2] The consistency of the proposed estimators is only established under a grid condition that Lemma 3.4 guarantees for sufficiently large γ_1. For general γ_1, including the small-γ_1 regime where β is close to ρ, no consistency result is proved. The numerical evidence in Table 5.2 shows that for the Toeplitz model with γ_2=5 and n_1=500, the scaled relative errors p^{2/3}|Θ_2hat−Θ_2|/Θ_2 have means between 4.00 and 5.67 with standard deviations between 2.32 and 3.43, which is not a small error and is difficult to reconcile with the blanket statement in Section 5.2 that 'the overall estimation precision remains well-controlled within a reasonable range for all settings under consideration'. The paper should clearly delimit the range of γ_1 and γ_2 for which the estimation consistency theorem applies and should report or explain these large-error cases explicitly.
  3. [Section 4.4, data-driven λ] The data-driven selection of λ in Section 4.4 is not covered by Theorem 2.2, which concerns a fixed λ. The text acknowledges the 'double-dipping' problem and states that simulation suggests the effect is limited, but no theorem is given for ℓ_max(F_{λhat}) when λhat is estimated from the same data. Moreover, under the null hypothesis the signal-to-noise ratio ξ(λ) is zero, so the proposed criterion λhat = argmax ξhat(λ)/Θ2hat(λ) may not converge to any well-defined limit; this identifiability issue is not discussed. The empirical size results in Table 5.3 for λhat are useful evidence, but the paper should either prove a theorem for data-dependent λ under explicit conditions on λhat or clearly label this procedure as a heuristic whose validation is numerical only.
minor comments (4)
  1. [Eq. (1.1)] In the display defining W_2, the final expression reads ':= 1/n_1 Y P_2 Y^T', but the preceding expression uses 1/n_2; this is presumably a typo and should be 1/n_2.
  2. [S.6, Lemma S.6.1] Lemma S.6.1 is stated as a central ingredient of the non-Gaussian local law, but its proof is not included; the text only says it 'closely follows' Lemma 3 of Han et al. (2018). Given that this lemma is the basis for the claimed universality, at minimum its standing as a proved versus imported result should be made precise.
  3. [Section 5.1] The paper states that results for t_4 and Poisson errors are consistent with the normal results but omits them from the manuscript. Since Theorem 2.2's non-Gaussian claim is precisely the part whose proof is incomplete, including at least one non-Gaussian simulation figure in the main text would substantially strengthen the empirical case.
  4. [Algorithm 1 and Section S.1] The truncation constant d and the recommended grid sizes K=I=500 are presented as fixed practical choices, but no sensitivity analysis for d, K, I is reported; a brief discussion of robustness to these choices would be helpful.

Circularity Check

0 steps flagged · score 0.0 of 10

No circular derivation: the Tracy-Widom limit is not identified with a fitted input, and the self-citations to Li (2024) and Han et al. (2018) are independent technical inputs; the explicitly deferred non-Gaussian local law is a proof-completeness gap, not circularity.

full rationale

I walked the derivation chain of Theorem 2.2 and the estimation procedure and found no step where a claimed prediction reduces by construction to its inputs. The centering and scaling parameters Θ1 and Θ2 are defined as deterministic functionals of the limiting population spectral distribution and the generalized Marchenko-Pastur equation, and the estimators are proven consistent to those same objects (Theorem 3.2) rather than being fitted to the largest eigenvalue whose distribution is being predicted. The proof of Theorem 2.2 does rely on the author's own Li (2024) for several technical lemmas (Lemma 2.5, Lemma 3.1, Lemma 3.2, Theorem 3.1), but those cited results concern the Marchenko-Pastur equation and limiting spectral distribution of MP-type matrices, not the TW1 fluctuation of the ridge-regularized F-matrix; they are parameter-free external results and therefore do not constitute circular support. The data-driven selection of λ is explicitly flagged as a potential 'double-dipping' problem and is only justified by simulation, not by a circular theorem. The main weakness is in S.5 and S.6, where the strong local law for Gλ is proved only under Gaussianity and the non-Gaussian extension is stated, not proved: S.5 says the extension 'can be completed following the strategy in Sections 7–10 of Knowles and Yin (2017), which is however beyond our scope,' and S.6 Step 3 says Theorem S.6.1 'can be shown by closely following' Han et al. (2018) with 'only minor changes' and then gives an outline with a stated-but-unproved Lemma S.6.1. This means Theorem 2.2 as stated under C2 is not fully proved, but an omitted proof is a correctness/completeness issue, not a circularity. I therefore assign score 0.

Assumptions & free parameters 6 free parameters · 6 assumptions · 0 invented entities

The central theory takes as input the population spectral distribution F^Sigma8 and its edge behavior (C4-C6), plus standard local-law and trace-lemma results. The estimation layer adds user-tuned grids, lambda bounds, and prior polynomial weights. The paper introduces no new physical entities or new conserved quantities.

free parameters (6)
  • lambda = fixed in theory; in practice chosen by grid or data-driven
    Regularization parameter in F_lambda = W1(W2 + lambda I)^{-1}. Theory fixes lambda; practice selects via argmax of estimated SNR over [tau_bar/50, 5 tau_bar].
  • grid K of sigma_k = approximately 500
    User-selected number of point masses for approximating F^Sigma8 in Algorithm 1; affects Theta1 and Theta2 estimates.
  • grid I of z_i = approximately 500, imaginary part 10^{-2} times inverse of largest eigenvalue of W2
    User-selected evaluation grid for the linear programming fitting in Algorithm 1.
  • truncation constant d = 2
    Threshold in Algorithm 1 to remove tiny weights and prevent floating-point instability.
  • prior polynomial coefficients pi_0, pi_1, pi_2 = pre-specified, uses r = 2
    Used for data-driven lambda selection under polynomial alternatives in Eq. (4.5); user specifies D, e.g., D = I_p.
  • lambda search bounds = tau_bar/50 and 5 tau_bar
    User-specified range for the data-driven lambda optimization in Section 4.4.
assumptions (6)
  • domain assumption Conditions C1-C5: high-dimensional asymptotic regime, finite moments, bounded spectrum of Sigma_p, stability of population ESD and of the leading edge eigenvalue.
    Define the asymptotic setting and regularity of Sigma_p required for the local laws and edge universality; violations such as factor-model spikes are excluded from C5 and only tested by simulation.
  • domain assumption Condition C6 and Definition 2.2: F^Sigma8 is regular near sigma_max.
    Ensures square-root (or atom-like) behavior of the density of G_lambda near its left edge rho, which drives the TW1 fluctuations in Lemma 2.5.
  • standard math Knowles-Yin (2017) and Han et al. (2018) local-law and edge-universality results for non-Gaussian sample covariance and F-matrices.
    S.6 Steps 3 and 4 derive the non-Gaussian local law and Green-function comparison by closely following these references; the paper gives only an outline.
  • standard math Li (2024) analysis of Marchenko-Pastur type equations (Theorems 3.2, 4.1, 5.1).
    Used to prove Lemmas 2.5, 3.1, 3.2, Corollary 3.1, and Theorem 3.1. The reference shares an author with this paper but is a separate published work.
  • standard math El Karoui and Kosters (2011) deterministic equivalent for quadratic forms of W2.
    Used in Lemma 4.1 and 4.2 to approximate q^T(W2 + lambda I)^{-1} q and traces by deterministic quantities involving phi(-lambda) and Sigma_p.
  • standard math Bai and Silverstein (1998) quadratic-form lemma for random vectors conditioned on W2.
    Used in the proof of Lemma 4.2 to replace nu^T D^{1/2}(W2 + lambda I)^{-1} D^{1/2} nu by its trace.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Ridge-Regularized Largest Root Test For High-Dimensional General Linear Hypotheses." pith.science (2026). https://pith.science/paper/MY7ZPLGC

@misc{pith2026250415510,
  author       = {Pith},
  title        = {Pith review of: Ridge-Regularized Largest Root Test For High-Dimensional General Linear Hypotheses},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/MY7ZPLGC}},
  note         = {Machine review of arXiv:2504.15510}
}
read the original abstract

A fundamental problem in multivariate analysis is testing general linear hypotheses for regression coefficients in a multivariate linear model. This framework encompasses a wide range of well-studied tasks, including MANOVA, joint significance testing of predictors, and detection of trends or seasonal effects. Among classical approaches, Roy's largest root test is particularly effective for detecting concentrated signals, relying on the largest eigenvalue of an F matrix constructed from residual covariance matrices. However, in high-dimensional settings, these matrices often become ill-conditioned or singular, rendering the test infeasible. To address this, we propose a ridge-regularized Roy's test that stabilizes the covariance estimation via a ridge term. We establish the asymptotic Tracy-Widom distribution of the largest eigenvalue of the regularized F-matrix under a high-dimensional regime, where both the dimension and hypotheses are comparable to the sample size, assuming only finite-moment conditions. A computationally efficient procedure is developed to estimate the associated centering and scaling parameters. We further analyze the power of the test under a class of low-rank alternatives and examine the influence of the regularization parameter. The method demonstrates strong performance in simulations and is applied to data from the Human Connectome Project to assess associations between volumetric brain measurements and behavioral variables.

Figures

Figures reproduced from arXiv: 2504.15510 by the authors.

Figure 4
Figure 4. , for two cases of [PITH_FULL_IMAGE:figures/full_fig_p021_4.png] view at source ↗
Figure 4.1
Figure 4.1. Ξpλq against λ when A “ 0.5 and γ1 “ 0.5. Columns (left-to-right): D “ Ip, Σp, Σ2 p ; Rows (top-to-bottom):γ2 “ 0.5, 1, 2 [PITH_FULL_IMAGE:figures/full_fig_p022_4_1.png] view at source ↗
Figure 4.2
Figure 4.2. Ξpλq against λ when A “ 0.9 and γ1 “ 0.5. Columns (left-to-right): D “ Ip, Σp, Σ2 p ; Rows (top-to-bottom):γ2 “ 0.5, 1, 2. 22 [PITH_FULL_IMAGE:figures/full_fig_p022_4_2.png] view at source ↗
Figures from the paper (4 more)
Figure 5.1
Figure 5.1. Figure 5.1: The estimated spxq, s 1 pxq and s 2 pxq (from left to right) when Σp is Poly-Decay, ˆγ2 “ 5, λ “ 0.5. Red: the true functions; Black: 5% and 95% pointwise percentile bands of the estimated functions. Σp γˆ2 λ n1 “ 500 n1 “ 250 n1 “ 100 Poly-Decay 0.5 0.5 0.06 (0.04) …
Figure 5
Figure 5. Figure 5: presents the empirical density of the regularized largest root [PITH_FULL_IMAGE:figures/full_fig_p026_5.png]
Figure 5.2
Figure 5.2. Figure 5.2: Empirical density of ℓmaxpFλq with λ “ λˆΣp . Solid lines use true parameters for normalization; dashed lines use estimated ones. Colors indicate n1: blue = 100, red = 250, purple = 500. Columns (left to right) correspond to ˆγ2 “ 0.5, 2, and 5; rows (top to bottom) …
Figure 5.3
Figure 5.3. Figure 5.3: Size-adjusted empirical power when Σ is Poly-Decay. Columns (left to right): ˆγ [PITH_FULL_IMAGE:figures/full_fig_p028_5_3.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Adaptable Regularized CCA Tests for Independence of High-Dimensional Random Vectors

    stat.ME 2026-07 accept novelty 6.0 of 10

    Ridge-plus-PC regularization of CCA yields stable high-dimensional independence tests whose null limits are normal (small k) or Tracy–Widom (large k), with consistent power under low-rank alternatives and a Bayesian-m...

Reference graph

Works this paper leans on

22 extracted references · 21 canonical work pages · cited by 1 Pith paper

  1. [1]

    Isotropic local laws for sample covariance and generalized Wigner matrices

    Alex Bloemendal, L´ aszl´ o Erd˝ os, Antti Knowles, Horng-Tzer Yau, and Jun Yin (2014). Isotropic local laws for sample covariance and generalized Wigner matrices. Electronic Journal of Probability, 19(33):1–53. Anderson, T. W. (1958). An introduction to multivariate statistical analysis , volume

  2. [2]

    Two-Sample Tests for High Dimensional Means with Thresholding and Data Transformation

    Wiley New York. Bai, Z., Chen, J., and Yao, J. (2010). On estimation of the population spectral distribution from a high- dimensional sample covariance matrix. Australian & New Zealand Journal of Statistics , 52(4):423–437. Bai, Z. and Saranadasa, H. (1996). Effect of high dimension: by an example of a two sample problem. Statistica Sinica, 6:311–329. Bai...

  3. [6]

    (3) There exists sufficiently small constants C ą 0 and C1ą 0 such that for all zPtE` iη : |E´ Θ1|ď C, 0ăηăC´1u inf τPSGλ |τ`qpzq|ě C1, whereSGλ is the support of Gλ

    Then, there exists a constant Cą 0, depending on a, such that C´1ď|qpzq|ď C, and ℑqpzqě C´1η, for all zP C` satisfying aď|z|ď a´1. (3) There exists sufficiently small constants C ą 0 and C1ą 0 such that for all zPtE` iη : |E´ Θ1|ď C, 0ăηăC´1u inf τPSGλ |τ`qpzq|ě C1, whereSGλ is the support of Gλ. It indicates that the equation (S.1) is non-singular near Θ...

  4. [7]

    Let zPD and suppose that |˜z´z|ď δpzq, where ˜z“´ 1{upzq` γ2ϕp´upzqq

    Suppose that u :DÑ C` is the Stieltjes transform of a compactly supported probability measure. Let zPD and suppose that |˜z´z|ď δpzq, where ˜z“´ 1{upzq` γ2ϕp´upzqq. Note that z“´ 1{qpzq` γ2ϕp´qpzqq and upzq“ qp˜zq. If ℑză 1, suppose also that |u´q|ď Cδ?κ`η` ? δ (S.3) holds at z` ip´5. Here, κ“|E´ Θ1|. Then, Eq. (S.3) holds at z. It is worth mentioning tha...

  5. [8]

    Let zPQ and suppose that |˜z´z|ď δpzq, where ˜z“vpzq` „ 1`γ2 ż τdF Σ8pτq ´τvpzq` λ ȷ´1

    Suppose that vpzq“ z´r 1`γ2ϱpzqs´1 where ϱpzq is the Stieltjes transform of a compactly supported probability measure. Let zPQ and suppose that |˜z´z|ď δpzq, where ˜z“vpzq` „ 1`γ2 ż τdF Σ8pτq ´τvpzq` λ ȷ´1 . Note that vpzq“ hp˜zq. If ℑză 1, suppose also that |v´h|ď Cδa |E´ρ|` η` ? δ (S.4) holds at z` ip´5. Then, Eq. (S.4) holds at z. Moreover, when z is a...

  6. [9]

    (ii) The averaged local law in Qppq holds as ˆϕppzq´ ϕppzq“ Oă ˆ 1 n2η ˙ , where ˆϕppzq“ p´1trRpzq

    (i) The entrywise local law in Qppq holds as Rpzq´ ` λΣ´1 p ´mppzqIp ˘´1 “OăpΨpzqq, uniformly in Qppq, where mppzq“ z´r 1`pp{n2qϕppzqs´1. (ii) The averaged local law in Qppq holds as ˆϕppzq´ ϕppzq“ Oă ˆ 1 n2η ˙ , where ˆϕppzq“ p´1trRpzq. (iii) The averaged local law in Qppq ´ holds as ˆϕppzq´ ϕppzq“ Oă ˆ 1 n2pκ`ηq ˙ , whereκ“|ρp´E|. 23 (iv) The averaged l...

  7. [10]

    We have }J}ď C η3, }BzJ}ď C η6

    Then, the following estimates hold for any zPQppq. We have }J}ď C η3, }BzJ}ď C η6. Furthermore, let wP Rp and vP Rn2. Then, we have the bounds pÿ u“1 |wTJp11qeu|2“ ℑwTJp11qw η , p`n2ÿ i“p`1 |vTJp22qei|2ď C}˜Z˜ZT} η ℑvTJp22qv`CvTv, pÿ u“1 |vTJp21qeu|2“ ℑvTJp22qv η , p`n2ÿ i“p`1 |wTJp12qei|2ďC}˜Z˜ZT} pÿ u“1 |wTJp11qeu|2. The estimates remain true for JpSq i...

  8. [11]

    The lemma follows from Theorem 2.10 of Alex Bloemendal et al. (2014). S.5.1.2 Weak entrywise law We first show a weak entrywise local law of J in Qppq. Proposition S.5.1. Suppose that the assumptions of Theorem S.5.1 hold. Define Λ“ max 1ďs,tďp`n2 |pJ´ Ωqst| and Λo“ max s‰t |pJ´ Ωqst|. 27 Then, Λ ăpn2ηq´1{4 uniformly in zPQppqYQppq away. Define the averag...

Show all 22 references
  1. [12]

    Let us first estimate Λ o

    Using Result (v) of Lemma S.5.2 and a simple induction argument, it is not hard to conclude that 1pΞqJpSq ss — 1, for any SĂt 1,...,p `n2u and sRS, satisfying |S|ď C. Let us first estimate Λ o. When u‰vPt 1,...,p u, using Result (ii) of Lemma S.5.2 and a large deviation estima...

  2. [13]

    Applying Lemma S.4.6 or Corollary S.4.3, when ηě 1, |ˆϕp´ϕp|ď Cp|δ1|`| δ2|q ă ΨB ăn´1{2 2

    We only need to show the diagonal elements are such that Juu´ 1 λ{σu´mp ăn´1{2 2 , u “ 1, 2,...,p, ´Jii´ 1 1`p{n2ϕp ăn´1{2 2 , i “p` 1,p ` 2,...,p `n2. Applying Lemma S.4.6 or Corollary S.4.3, when ηě 1, |ˆϕp´ϕp|ď Cp|δ1|`| δ2|q ă ΨB ăn´1{2 2 . Then, by Eq. (S.8) ´ 1 Jii “ 1`ϕp...

  3. [14]

    Then on QppqYQppq away we have 1 p pÿ u“1 1 pλ{σu´z´ 1 n2 řp`n2 i“p`1 Jiiq2p1´ Euq 1 Juu “Oă ` Υ2˘ and 1 n2 p`n2ÿ i“p`1 p1´ Eiq 1 Jii “Oă ` Υ2˘

    Suppose moreover that Λ ăn´c 2 and Λo ă Υ on QppqYQppq away. Then on QppqYQppq away we have 1 p pÿ u“1 1 pλ{σu´z´ 1 n2 řp`n2 i“p`1 Jiiq2p1´ Euq 1 Juu “Oă ` Υ2˘ and 1 n2 p`n2ÿ i“p`1 p1´ Eiq 1 Jii “Oă ` Υ2˘ . Th results are an extension of Lemma 5.6 of Knowles and Yin (2017), Le...

  4. [15]

    Therefore, to show ˇˇˇˇˇ 1 p pÿ j“1 fpℓjpGλqq´ ż fpτqdGλpτq ˇˇˇˇˇ ă 1 n2 , it suffices to show ¿ ˆR |fpzq| ˇˇˇˆϕppzq´ ϕppzq ˇˇˇdz ă 1 n2

    It follows ˇˇˇˇˇˇˇ ¿ Rz ˆR fpzqˆϕppzqdz ˇˇˇˇˇˇˇ ă 1 n2 . Therefore, to show ˇˇˇˇˇ 1 p pÿ j“1 fpℓjpGλqq´ ż fpτqdGλpτq ˇˇˇˇˇ ă 1 n2 , it suffices to show ¿ ˆR |fpzq| ˇˇˇˆϕppzq´ ϕppzq ˇˇˇdz ă 1 n2 . It is a direct consequence from the averaged local law on Qppq away in Theorem S....

  5. [16]

    local laws

    (i) The entrywise local law holds as Lpzq´ qpzqIn1“OăpΦpzqq, uniformly zPD. (ii) The averaged local law holds as ˇˇLpzq´ qpzq ˇˇ ă 1 n1η, 42 uniformly zPD. Here, Lpzq“ 1 n1 n1ÿ i“1 Liipzq and Lijpzq is the pi,jq-th element of Lpzq. Once Theorem S.5.1 is ready, we can show the ...

  6. [17]

    Let ˆΘp1“ ˆfp´ ˆβpq and ˆΦpzq“ d ℑˆqppzq n1η ` 1 n1η

    Clearly, for any fixed discrete points gj’s such that gpą 0, ˆβp exists and is unique. Let ˆΘp1“ ˆfp´ ˆβpq and ˆΦpzq“ d ℑˆqppzq n1η ` 1 n1η. Define the following domain ˆD“ ˆDpa,a1,n 1q :“tz“E` iηP C` : |E´ ˆΘp1|ď a, n´1`a1 1 ďηď 1{a1u. The following results hold by applying T...

  7. [18]

    Therefore, ´ˆqppzq is away from the support of Gλ

    The constant C only depends on F Σ8, γ1 and γ2. Therefore, ´ˆqppzq is away from the support of Gλ. Applying S.5.3, ˇˇˇˇ ż dF Gλpτq τ` ˆqppzq´ ż dGλpτq τ` ˆqppzq ˇˇˇˇ ă 1 n2 . It follows that z`Oăpn´1 1 q“´ 1 ˆqppzq`pp{n1q ż dGλpτq τ` ˆqppzq. Applying Lemma S.4.2, we obtain |ˆq...

  8. [19]

    We work on the net S satisfying the condition that E`iηl P S, l “ 0,...,L , from now on. We define Sm :“ ␣ zPS : ℑzěn´δm 1 ( corresponding to the following events: Am“ ␣ }B1pLpzq´ qpzqIn1qB˚ 2 TppZq}8 ă 1, for any zPSm ( (S.6) and Cm“ ␣ }B1pLpzq´ qpzqIn1qB˚ 2 TppZq}8 ă Φ, for ...

  9. [20]

    1 pνTD1{2pW2`λIpq´1D1{2ν`oPp1q. Using Lemma 2.7 of Bai and Silverstein (1998), conditional on W2, 1 pνTD1{2pW2`λIpq´1D1{2ν“ 1 ptr

    Under DA, 1 ps Fλ“ 1 ps Fp0q λ `dpqpqT ppW2`λIpq´1` 1 n1ps ´ BXP 1ZT Σ1{2 p ` Σ1{2 p ZP1XTBT ¯ pW2`λIpq´1. Since ℓmaxpFp0q λ q“ OPp1q as indicated by Theorem 2.2 and 1 ps{2n1 }BXP 1}2}ZT}}Σ1{2 p }}pW2`λIpq´1}“ OPp1q, we have 1 psℓmaxpFλq“ dpqT ppW2`λIpq´1qp`oPp1q. 54 Using the...

  10. [21]

    Therefore, s1pxq“ C pρ´xq2 `C ż τąρ dGλpτq τ´x Ñ8, as xÒρ

    Using Theorem 3.2 of Li (2024), Gλ is discrete at ρ. Therefore, s1pxq“ C pρ´xq2 `C ż τąρ dGλpτq τ´x Ñ8, as xÒρ. 55 When ωmaxγ2 ă 1, it is straightforward to verify that there exists a c ą 0 such that x1phq ă0 when h P rλ{σmax´c,λ{σmaxq. On the other, as hÑ´8 , x1phqÑ

  11. [22]

    Using Theorem 5.1 of Li (2024), there exists constants C1, C2 and ϵ such that C1 ?x´ρďfGλpxqď C2 ?x´ρ, x Ppρ,ρ `ϵq where fGλ is the density function of Gλ

    Using Theorem 4.1 of Li (2024), ρ“xph0q. Using Theorem 5.1 of Li (2024), there exists constants C1, C2 and ϵ such that C1 ?x´ρďfGλpxqď C2 ?x´ρ, x Ppρ,ρ `ϵq where fGλ is the density function of Gλ. It follows that s1pxqě żρ`ϵ ρ fGλpτqdτ pτ´xq2 Ñ8, as xÒρ. When F Σ8 is continuou...

  12. [197]

    Pillai, N

    John Wiley & Sons. Pillai, N. S. and Yin, J. (2014). Universality of covariance matrices. Annals of Applied Probability, 24(3):935–

  13. [1001]

    Ridge-regularized Largest Root Test for High-Dimensional General Linear Hypotheses

    Silverstein, J. W. and Bai, Z. (1995). On the empirical distribution of eigenvalues of a class of large dimensional random matrices. Journal of Multivariate analysis , 54(2):175–192. Silverstein, J. W. and Choi, S.-I. (1995). Analysis of the limiting spectral distribution of l...

  14. [2006]

    El Karoui, N. (2007). Tracy–Widom limit for the largest eigenvalue of a large class of complex sample covariance matrices. The Annals of Probability , 35(2):663–714. El Karoui, N. (2008). Spectrum estimation for large dimensional covariance matrices using random matrix theory....

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.