{"id":"ee4cb580-ee72-4a53-a77b-af7e45343333","arxiv_id":"2504.15510","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Ridge-regularized Roy's test has a Tracy-Widom type-1 limit and consistent parameter estimators, enabling largest-root testing when dimension p exceeds n2.","lead":"A new statistical test extends Roy's largest root method to high-dimensional settings where standard covariance matrices are singular, using ridge regularization. The paper derives the theoretical distribution of the test statistic and shows it works on brain-imaging data.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Non-Gaussian universality in Theorem 2.2 rests on an explicitly deferred proof: S.5 proves the local law only for Gaussian errors, and S.6 Step 3 gives only an outline for the non-Gaussian extension by 'closely following' Han et al. (2018).","rationale":"The reader's weakest_assumption is exactly the load-bearing concern: finite-moment universality behind Theorem 2.2 is incomplete. The supplement explicitly says the non-Gaussian local law is beyond scope (S.5) and gives only an outline in S.6 Step 3. I looked for an independent flaw in the fixed-lambda Gaussian argument and did not find one strong enough to overturn the claimed result; estimation consistency under large gamma1 and the data-driven lambda selection are secondary because Theorem 2.2 concerns fixed parameters. The simulations are supportive but are reported mainly for Gaussian errors, and no machine-checked proof is provided. Because the missing non-Gaussian proof affects the headline claim as stated but is plausibly repairable, the reader's conditional verdict remains appropriate: the paper should either supply the full non-Gaussian proof or restate Theorem 2.2 under Gaussianity. No verdict change from the reader is needed.","tokens_in":53704,"tokens_out":16318,"duration_ms":153376,"concrete_test":"Write out the missing non-Gaussian proof of Theorem S.6.1 and Lemma S.6.1 for the full block matrix H(z), without appealing to 'closely following'. The decisive algebraic step is the fourth-moment (Green function) comparison: for entries satisfying C2 but non-Gaussian, explicitly compute the residual after replacing Z by a Gaussian Z0; if the residual contains the fourth cumulant of z11, the cited 'minor changes' argument fails for discrete or heavy-tailed errors. A compact version: prove Lemma S.6.1 for Rademacher errors and verify that the required bound (n^{24*delta}*Phi)^k holds with the same constants as in the Gaussian case; if it does not, Theorem 2.2 as stated requires either a completed proof or a restriction to Gaussian errors.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Theorem 2.2 is stated under C2 (finite moments), but the proof that the largest eigenvalue of F_lambda has TW1 fluctuations in that generality is not supplied. In S.5, the strong local law for G_lambda is proved only under SC1 (Gaussian Z); the text says the non-Gaussian extension 'can be completed following the strategy in Sections 7–10 of Knowles and Yin (2017), which is however beyond our scope.' In S.6, Step 3, the non-Gaussian version of Theorem S.6.1 is said to follow by 'closely following' Sections 9.1, 9.2 and 10 of Han et al. (2018) with 'only minor changes', and only an outline plus a stated-but-unproved Lemma S.6.1 are given. This gap is load-bearing because the abstract's claim 'assuming only finite-moment conditions' rests entirely on that transfer. The transfer is not automatic: for non-Gaussian Z, the blocks ZU1 and ZU2 of the linearizing 3x3 block matrix H(z) are dependent, while the Gaussian proof exploits their independence. Moreover, the Han et al. (2018) argument being invoked is for the lambda=0 case where the central block is zero and the ZU2 term is absent. Until the non-Gaussian local law and Green-function comparison are actually carried out for the full H(z), Theorem 2.2 is unproved in the stated generality.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a ridge-regularized Roy's largest-root test for high-dimensional general linear hypotheses. For a fixed λ>0, the test statistic is the largest eigenvalue of F_λ = W_1(W_2+λI_p)^{-1}, where W_1 and W_2 are the hypothesis and residual sum-of-squares matrices of a multivariate linear model.  The main theoretical result, Theorem 2.2, asserts that after centering and scaling by quantities Θ_1(λ) and Θ_2(λ) defined through a generalized Marchenko–Pastur equation, the statistic converges to the Tracy–Widom type-1 law under conditions C1–C6, assuming only finite-moment conditions on the errors.  The paper also develops a linear-programming-based estimation procedure for Θ_1 and Θ_2 from the eigenvalues of W_2, proves consistency of this procedure under certain grid conditions, analyzes power under low-rank alternatives, reports simulation studies, and applies the method to Human Connectome Project data.","tokens_in":54111,"tokens_out":4016,"duration_ms":39255,"significance":"If fully established, the result would be a useful extension of Roy's largest-root testing to the regime where the number of responses and the number of tested hypotheses are comparable to the sample size, including cases with p>n_2, which classical F-tests cannot handle.  The proposed regularization and the accompanying estimation algorithm for the centering and scaling constants are genuinely new and are backed by substantial numerical work.  The simulations show good null coverage in many settings and clear power advantages over the compared ridge-LRT and projection-LRT baselines.  However, the advertised finite-moment universality is not actually proved: the local law for G_λ is proved only under Gaussianity in Section S.5, and the non-Gaussian extension is deferred to an outline in Section S.6.  In addition, the consistency of the parameter estimators is guaranteed only under grid conditions that are shown to hold for large γ_1, while some simulations (notably Toeplitz, γ_2=5) show relative estimation errors for Θ_2 that are far from negligible.  The data-driven selection of λ is acknowledged to be potentially circular but is left without a theoretical justification.","major_comments":[{"comment":"Theorem 2.2 is stated under the finite-moment condition C2, but the proof of the non-Gaussian Tracy–Widom limit is not supplied.  Section S.5 explicitly proves the strong local law for G_λ only under SC1 (Gaussian Z) and states that the non-Gaussian extension 'can be completed following the strategy in Sections 7–10 of Knowles and Yin (2017), which is however beyond our scope'.  Section S.6, Step 3, says the non-Gaussian local law follows by 'closely following' Sections 9.1, 9.2 and 10 of Han et al. (2018) with 'only minor changes', and the key estimate, Lemma S.6.1, is stated but not proved.  This gap is load-bearing because the abstract's claim 'assuming only finite-moment conditions' rests entirely on this transfer.  The transfer is not automatic: for non-Gaussian Z the blocks ZU_1 and ZU_2 in the linearizing matrix H(z) are dependent, whereas the Gaussian proof exploits their independence, and the Han et al. (2018) argument is for the λ=0 case where the central block is zero.  The authors should either provide a complete proof of the non-Gaussian local law and Green-function comparison for the full H(z), or state Theorem 2.2 under Gaussian errors and describe the finite-moment version as a conjecture supported by simulations.","section":"S.6, Step 3; Theorem 2.2; Condition C2"},{"comment":"The consistency of the proposed estimators is only established under a grid condition that Lemma 3.4 guarantees for sufficiently large γ_1.  For general γ_1, including the small-γ_1 regime where β is close to ρ, no consistency result is proved.  The numerical evidence in Table 5.2 shows that for the Toeplitz model with γ_2=5 and n_1=500, the scaled relative errors p^{2/3}|Θ_2hat−Θ_2|/Θ_2 have means between 4.00 and 5.67 with standard deviations between 2.32 and 3.43, which is not a small error and is difficult to reconcile with the blanket statement in Section 5.2 that 'the overall estimation precision remains well-controlled within a reasonable range for all settings under consideration'.  The paper should clearly delimit the range of γ_1 and γ_2 for which the estimation consistency theorem applies and should report or explain these large-error cases explicitly.","section":"Theorem 3.2, Lemma 3.4, Table 5.2"},{"comment":"The data-driven selection of λ in Section 4.4 is not covered by Theorem 2.2, which concerns a fixed λ.  The text acknowledges the 'double-dipping' problem and states that simulation suggests the effect is limited, but no theorem is given for ℓ_max(F_{λhat}) when λhat is estimated from the same data.  Moreover, under the null hypothesis the signal-to-noise ratio ξ(λ) is zero, so the proposed criterion λhat = argmax ξhat(λ)/Θ2hat(λ) may not converge to any well-defined limit; this identifiability issue is not discussed.  The empirical size results in Table 5.3 for λhat are useful evidence, but the paper should either prove a theorem for data-dependent λ under explicit conditions on λhat or clearly label this procedure as a heuristic whose validation is numerical only.","section":"Section 4.4, data-driven λ"}],"minor_comments":[{"comment":"In the display defining W_2, the final expression reads ':= 1/n_1 Y P_2 Y^T', but the preceding expression uses 1/n_2; this is presumably a typo and should be 1/n_2.","section":"Eq. (1.1)"},{"comment":"Lemma S.6.1 is stated as a central ingredient of the non-Gaussian local law, but its proof is not included; the text only says it 'closely follows' Lemma 3 of Han et al. (2018).  Given that this lemma is the basis for the claimed universality, at minimum its standing as a proved versus imported result should be made precise.","section":"S.6, Lemma S.6.1"},{"comment":"The paper states that results for t_4 and Poisson errors are consistent with the normal results but omits them from the manuscript.  Since Theorem 2.2's non-Gaussian claim is precisely the part whose proof is incomplete, including at least one non-Gaussian simulation figure in the main text would substantially strengthen the empirical case.","section":"Section 5.1"},{"comment":"The truncation constant d and the recommended grid sizes K=I=500 are presented as fixed practical choices, but no sensitivity analysis for d, K, I is reported; a brief discussion of robustness to these choices would be helpful.","section":"Algorithm 1 and Section S.1"}],"recommendation":"major_revision","confidential_remarks":"The paper relies on several lemmas and theorems from Li (2024), a paper by the same author, without reproducing proofs: Lemmas 2.5, 3.1, 3.2 and Theorem 3.1 are deferred to that source.  For a journal publication, the author should either include self-contained proofs of these results or clearly import them with explicit statements, so that the reader can verify the conditions under which they hold.  There is also a concern about scope: the core advertised contribution, finite-moment universality, is not proved in the current manuscript, so the revision must either add the missing proof or revise the claims to the Gaussian case."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: Theorem 2.2 is a real extension of the lambda=0 largest-root results to ridge-regularized F-matrices, and the paper is worth engaging with. But the abstract's claim of \"assuming only finite-moment conditions\" is not backed by a complete proof. The supplement proves the strong local law for G_lambda only under Gaussian errors (S.5), explicitly saying the non-Gaussian extension is beyond scope; S.6 Step 3 gives only an outline for the non-Gaussian local law by \"closely following\" Han et al. (2018), with a stated-but-unproved Lemma S.6.1. The transfer is not automatic: for non-Gaussian Z, the blocks ZU1 and ZU2 in the linearizing matrix are dependent, while the Gaussian proof exploits their independence, and Han et al. handles the lambda=0 case where the central block is absent. So the main theorem, as stated, is unproved in its full advertised generality.\n\nWhat is genuinely good: the regularization idea is borrowed, but the largest-root statistic for W1(W2 + lambda I)^{-1} is new, and the centering/scaling estimation procedure (Algorithms 1-3) is a real contribution. The simulations show the null distribution tracks TW1 well across settings, including the factor model where C5 is violated. The power analysis is standard but fine, and the HCP application is reasonable.\n\nSoft spots, in proportion. First, the incomplete non-Gaussian proof is load-bearing; either the proof must be supplied or the theorem should be narrowed to the Gaussian case. Second, consistency of the estimators is proven only under a large-gamma1 grid condition (Lemma 3.4), and the numerics show relative errors for Theta2 of 4-5 with big variance in the Toeplitz gamma2=5 case. That matters for users who need calibrated p-values. Third, the data-driven selection of lambda has no formal justification; the \"double-dipping\" paragraph acknowledges the issue and the simulations suggest limited damage, but a rigorous statement is missing. Fourth, several lemmas are imported from the author's own Li (2024); that is not automatically a flaw, but those results are not verified in this manuscript.\n\nBottom line: if the non-Gaussian local law can be completed, this is a solid contribution. If not, the honest fix is to state Theorem 2.2 under Gaussianity and flag the finite-moment version as a conjecture. The paper deserves a serious referee, but the referee should insist on seeing the missing proof or a narrowed claim.","headline":"The ridge-regularized largest root result is a genuine advance, but the advertised finite-moment Tracy–Widom theorem is not fully proved; the paper deserves review, with the non-Gaussian gap addressed.","tokens_in":54638,"tokens_out":1968,"would_cite":true,"duration_ms":18731,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62H15","60B20","62H10","62J05"],"pacs":[],"model":"deepseek-v4-flash","headline":"Adding a ridge term keeps Roy's largest-root test valid even when the residual covariance matrix is singular.","keywords":["ridge regularization","Roy's largest root test","Tracy-Widom distribution","general linear hypotheses","high-dimensional multivariate regression","F-matrix","random matrix theory"],"falsifier":"Simulate the null distribution with t-distributed errors with 4 degrees of freedom, take n2=500, p=2500, n1 in {100,250,500}, and λ in {0.5,1,1.5}, then compare the empirical 95th percentile of $p^{2/3}(\\ell_{\\max}(F_\\lambda)-\\hat\\Theta_1)/\\hat\\Theta_2$ with the 95th percentile of TW1 across many replications; a systematic exceedance beyond Monte Carlo error as p grows would contradict the finite-moment universality claim.","tokens_in":53520,"feed_emoji":"📊","tokens_out":11172,"duration_ms":93740,"temperature":0.7,"pith_summary":"The paper aims to make Roy's largest-root test usable when the number of response variables p is comparable to, or larger than, the residual degrees of freedom n2, a setting where the classical F-matrix $W1W2^{-1}$ is ill-conditioned or undefined. Its proposal is to replace $W2^{-1}$ by (W2 + λIp)^-1, and its main theorem states that the largest eigenvalue of this regularized ratio, after centering and scaling by explicit constants Θ1 and Θ2, converges in distribution to the Tracy-Widom law of type 1. If true, this gives calibrated p-values for a broad class of linear hypotheses, such as MANOVA, joint significance of predictors, and tests for trends or seasonal effects, using only eigenvalue information from W2 to estimate the centering and scaling. The paper also analyzes power under low-rank alternatives, proposes a data-driven choice of λ, and demonstrates the procedure on brain-volume and behavioral data from the Human Connectome Project.","feed_headline":"Ridge restores Roy's largest-root test in high dimensions","feed_subtitle":"Even with p>n2, the largest eigenvalue tracks Tracy-Widom, so p-values stay calibrated","key_machinery":"The machine at the center is the ridge-regularized F-matrix $F_\\lambda = W_1(W_2+\\lambda I_p)^{-1}$ and its symmetric counterpart $\\tilde F_\\lambda=U_1^TZ^TG_\\lambda^{-1}ZU_1$, which shares the same nonzero eigenvalues. The argument is carried by the generalized Marchenko-Pastur equation $$\\phi=\\int \\frac{\\tau\\,dF_{\\Sigma_\\infty}(\\tau)}{\\tau[1+\\gamma_2\\phi]^{-1}-z}+\\$\\lambda$,$$ whose solution $s(x)$ near the left edge $\\rho$ of $G_\\lambda$'s limiting spectrum controls everything: the edge parameter β solves $\\beta^2 s'(\\beta)=1/\\gamma_1$, and the centering and scaling are $\\Theta_1=(1+\\gamma_1\\beta s(\\beta))/\\beta$ and $\\Theta_2=(\\gamma_1^3 s''(\\beta)/2+\\gamma_1^2/\\beta^3)^{1/3}$. Around this core, the paper develops a local law for the resolvent of $G_\\lambda$, uses edge universality to transfer Gaussian Tracy-Widom fluctuations to finite-moment errors, and estimates Θ1 and Θ2 by discretizing the population spectral distribution through a linear program and then solving an ODE, using only the eigenvalues of W2.","core_discovery":"The central discovery is Theorem 2.2: under the regime C1-C6 and for any fixed λ>0, $$$p^{{2/3}}$\\Theta_2(\\$\\lambda$)^{-1}\\big(\\ell_{\\max}(F_\\$\\lambda$)-\\Theta_1(\\$\\lambda$)\\big)\\Rightarrow \\mathrm{TW}_1,$$ where $F_\\lambda=W_1(W_2+\\lambda I_p)^{-1}$, $\\mathrm{TW}_1$ is the Tracy-Widom distribution of type 1, and $\\Theta_1,\\Theta_2$ are explicit functions of an edge parameter β defined through the generalized Marchenko-Pastur equation for the regularized matrix $G_\\lambda=ZU_2U_2^TZ^T+\\lambda\\Sigma_p^{-1}$. The theorem covers the case $p>n_2$ where W2 is singular and assumes only finite moments for the error entries. It also yields the joint Tracy-Widom GOE fluctuations of the first k largest eigenvalues. The proof proceeds through a strong local law for Gλ together with a Green-function comparison that transplants the Gaussian edge fluctuations to general error distributions.","pith_inferences":["If the non-Gaussian local law is completed, the same proof scheme should transfer to other regularized statistics, such as ridge-regularized canonical correlations or partial least squares, because only the linearizing block structure and edge stability are used.","The estimation pipeline, discretize the population spectrum, fit weights by linear programming, then solve an ODE for s(x), is a reusable template for estimating edge parameters of other random matrix models, not just the ridge F-test.","One testable extension is a theoretical account of data-driven λ selection: the paper's simulation evidence that double-dipping is negligible could become a theorem under suitable convergence rates for λ̂ toward its limiting maximizer.","In the discrete-edge case where γ2ωmax>1, the smallest eigenvalue of Gλ sits exactly at λ/σmax; understanding how close the Tracy-Widom approximation remains there is a natural robustness check for covariance structures with spikes, such as factor models."],"forward_implications":["General linear hypotheses can be tested when p exceeds n2 without imputing or inverting a singular covariance estimate, because the ridge term stabilizes W2.","Calibrated p-values would require no distributional assumption beyond finite moments: the centering and scaling are estimated from the eigenvalues of W2, and the reference law is the fixed Tracy-Widom type-1 distribution.","The test is designed for concentrated alternatives: under low-rank signals the power tends to 1, and λ can be chosen by maximizing an estimated signal-to-noise ratio rather than by cross-validation.","A data-driven λ choice changes the null distribution only negligibly in the paper's simulations, so the procedure can be run with λ selected on the same data, with data splitting suggested as extra protection.","The joint distribution of the top k eigenvalues is also Tracy-Widom/GOE, so inference using the first several roots, for example testing multiple signal directions, is a direct corollary."],"supporting_citations":[{"why":"Supplies the consistent estimators of population spectral moments used for the polynomial alternatives and for estimating ξ(λ).","marker":"Bai et al. (2010)"},{"why":"Introduces the point-mass discretization of the population spectral distribution behind Algorithm 1.","marker":"El Karoui (2008)"},{"why":"Establishes the Tracy-Widom law for the unregularized F-matrix and supplies the non-Gaussian local-law strategy that Theorem 2.2 extends.","marker":"Han et al. (2018)"},{"why":"Provides the earlier Tracy-Widom result for the largest eigenvalue of F-type matrices.","marker":"Han et al. (2016)"},{"why":"Gives the type-1 Tracy-Widom limit for the largest root of the double Wishart matrix, the classical baseline the paper generalizes.","marker":"Johnstone (2008)"},{"why":"Supplies the anisotropic local-law and stability framework used to analyze Gλ and to transfer Gaussian fluctuations to general errors.","marker":"Knowles and Yin (2017)"},{"why":"Provides the Tracy-Widom theorem for sample covariance with general population used in the Gaussian step of the proof.","marker":"Lee and Schnelli (2016)"},{"why":"Motivates the weight-fitting loss used to estimate H1 and H2 in the estimation procedure.","marker":"Ledoit and Wolf (2012)"},{"why":"Supplies the ridge-regularized likelihood-ratio baseline and the polynomial-alternative power framework reused in Section 4.","marker":"Li et al. (2020a)"},{"why":"Motivates the ridge regularization scheme for high-dimensional testing that the paper adapts to the F-matrix.","marker":"Li et al. (2020b)"}],"fun_headline_variants":["Ridge-regularized Roy's test hits Tracy-Widom even with p > n","Ridge tames singular covariance so Roy's test thrives in high dimensions","Roy's largest-root test rescued by ridge in the p > n regime","Regularization enables Roy's test for singular covariance matrices"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The theorem's promise of universality under only finite moments of the error entries rests on a strong local-law estimate for $G_\\lambda$; for non-Gaussian errors the paper gives an outline referencing prior methods rather than a complete proof, so the advertised scope depends on that outline succeeding.","fun_headline_variants_meta":{"raw":{"variants":["Ridge-regularized Roy's test hits Tracy-Widom even with p > n","Ridge tames singular covariance so Roy's test thrives in high dimensions","Roy's largest-root test rescued by ridge in the p > n regime","Regularization enables Roy's test for singular covariance matrices"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001083,"raw_usage":{"total_tokens":4554,"prompt_tokens":996,"completion_tokens":3558,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":612,"completion_tokens_details":{"reasoning_tokens":3479}},"tokens_in":612,"tokens_out":3558,"duration_ms":23496,"temperature":1.0,"reasoning_tokens":3479,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T11:24:41.108877+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate the null distribution with t-distributed errors with 4 degrees of freedom, take n2=500, p=2500, n1 in {100,250,500}, and λ in {0.5,1,1.5}, then compare the empirical 95th percentile of $p^{2/3}(\\ell_{\\max}(F_\\lambda)-\\hat\\Theta_1)/\\hat\\Theta_2$ with the 95th percentile of TW1 across many replications; a systematic exceedance beyond Monte Carlo error as p grows would contradict the finite-moment universality claim.","supporting_citations":[],"review_version":1}