{"id":"658e1ae7-e187-4683-8d66-ece07098947f","arxiv_id":"1908.07444","paper_version":4,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"For sample covariance matrices with convexly decaying population spectrum, the top eigenvalues exhibit Weibull fluctuations above a critical aspect ratio d_+ and Gaussian fluctuations below it.","lead":"This paper proves that, when the population eigenvalues have convex decay at the top of their range, the largest eigenvalues of the sample covariance matrix are controlled by the population's largest eigenvalues and converge to a Weibull distribution above a critical aspect ratio, and to a Gaussian below it. It gives the first rigorous edge fluctuation results for sample covariance matrices in this non-square-root regime.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 2.8 is conditioned on the good-configuration event Ω; the deterministic case assumes the uniform averaged convergence (2.16) without proof, so the advertised scope for deterministic Σ is not established.","rationale":"The reader identified the good-configuration event Ω, especially (2.14) and (2.16), as the weakest assumption; my reading agrees. The i.i.d. case, which is the paper's main advertised result, is supported because Appendix A proves P(Ω) → 1 with the required rates. The deterministic case, however, is not derived from any checkable condition: Assumption 2.6(i) simply asserts Ω, and (2.16) is used at a precise step in Lemma 4.4 to obtain the contraction that powers the local law and the eigenvalue comparison. This makes the deterministic claim conditional on a strong, non-explicit uniform convergence hypothesis. This does not overturn the paper's main conclusion for random i.i.d. Σ, so the appropriate verdict remains CONDITIONAL, the same as the reader's; I recommend UNCHANGED.","tokens_in":43899,"tokens_out":26296,"duration_ms":245362,"concrete_test":"Independently re-derive Lemma 4.4 starting from (4.32) while omitting the bound from (2.16); if the first term cannot be controlled by o(|1/ˆm_fc − 1/m_fc|) using only the local law and (2.14), then (2.16) is essential. If it is essential, test the natural deterministic candidate σ_i = F^{-1}((M−i+1/2)/M): compute the left-hand side of (2.16) at the z ∈ D_φ where it is maximal over a fine grid for b = 2, d = 1.5, M = 10^5, and compare with M^{-1/2+φ}. If the bound is violated, Assumption 2.6(i) fails for the most natural deterministic Σ, and the deterministic version of Theorem 2.8 has no obvious instance.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The proof of Theorem 2.8 inherits all of its quantitative control from the event Ω in Definition 2.5, which is simply postulated for deterministic Σ in Assumption 2.6(i). The specific load-bearing ingredient is condition (2.16): in Lemma 4.4, the first term on the right of (4.32) is bounded by C M^{-1/2+φ+ε/2} only when (2.16) holds, and this bound is what makes the subsequent contraction (4.40) yield |1/ˆm_fc − 1/m_fc| ≺ M^{-1/2+φ}. Condition (2.16) is a uniform-in-z∈D_φ estimate of the difference between the empirical average of σ/(1+σ m_fc) and its ν-integral. Appendix A verifies it only for i.i.d. Σ (Lemma A.3); for deterministic Σ satisfying just the ESD convergence and the spacing conditions (2.10)–(2.12), (2.16) is an extra, non-explicit hypothesis. Without (2.16), the local law in Proposition 5.1 and the eigenvalue-tracking estimate in Proposition 4.9 have no basis. Thus the abstract's claim that the results 'also hold for deterministic Σ with some additional assumptions' hides the main input: the theorem is conditional on a complex event that is not derived from any more primitive deterministic hypothesis.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies the extremal eigenvalues of sample covariance matrices Q = (Σ^{1/2}X)(Σ^{1/2}X)^* for M×N random X with independent entries and a diagonal population covariance Σ. Under the assumption that the limiting spectral density of Σ has the Jacobi form ρ_ν(t) ∝ (1−t)^b f(t) with b>1, the authors define a threshold d_+ = ∫ t^2 (1−t)^{-2} dν(t). For d>d_+, they claim that the right edge of the LSD of Q has convex decay and that the γ largest eigenvalues of Q, rescaled by M^{1/(b+1)}, converge to the corresponding order statistics of Σ, with an explicit Weibull limit for the largest eigenvalue. For d<d_+, they claim a Gaussian fluctuation at scale M^{-1/2} when the entries of Σ are i.i.d. The proof uses the deformed Marchenko–Pastur equation, a local law near the edge, fluctuation averaging, and an intermediate random object \\mhat m_fc that depends on Σ. The main theorems are stated under a 'good configuration' event Ω that encodes spacing of the top population eigenvalues, a non-resonance condition, and an averaged convergence condition.","tokens_in":44246,"tokens_out":4518,"duration_ms":47076,"significance":"If the results are correct, the paper provides a sharp phase transition for the edge behavior of sample covariance matrices with general population: above d_+ the largest eigenvalues follow the population order statistics with a Weibull limit, and below d_+ the fluctuation is Gaussian. The constants d_+, C_d, and C_ν are explicit functionals of the population law ν, and no parameter is fitted. The paper also extends the deformed-Wigner extremal-eigenvalue strategy to the sample covariance setting, which is a technically substantial step. However, the manuscript currently leaves several load-bearing estimates and verifications to analogies or external references, so the significance can be fully assessed only after these gaps are filled.","major_comments":[{"comment":"The proof of the density bound (2.19) is omitted: the text says the proof is analogous to Lemma A.4 of [18] and gives no details. This bound is load-bearing because it is used in Appendix A, Lemma A.5, to control Im m_fc, which in turn supports the verification of condition (2.14) and the estimates in §5. Without a full proof of (2.19), the convex-decay assertion that is the basis for the order-statistic and Weibull results is not established within the manuscript.","section":"§4.1, Theorem 2.7"},{"comment":"The deterministic-Σ case is advertised in the abstract as holding 'with some additional assumptions,' but the only such assumption is the un-verified event Ω of Definition 2.5, in particular the uniform averaged convergence condition (2.16). This condition is exactly what bounds the first term in (4.32), and it is also needed for Lemma 4.4 and Proposition 5.1. For i.i.d. Σ, Lemma A.3 verifies a stronger form, but for deterministic Σ satisfying only ESD convergence and spacing conditions, (2.16) is an additional non-explicit hypothesis. The theorem is therefore conditional on a hypothesis that is not derived from any more primitive deterministic condition, and the deterministic scope claimed in the abstract is not actually established.","section":"§4.4 and Assumption 2.6(i), condition (2.16)"},{"comment":"The crucial estimate |L_+ − λ_1| ≺ M^{-2/3} is obtained by asserting that the model satisfies Condition 1.1 of [3] and then applying Theorem 4.1 of [3]. No verification of Condition 1.1 is given, and the text only says that 'adapting the idea of the proof of Lemma A.4 in [18]' yields the needed bounds. This step is essential: without |L_+ − λ_1| ≺ M^{-2/3}, the Gaussian fluctuation of λ_1 cannot be identified with that of \\hat L_+, which is the point of Theorem 2.9. A complete proof must either verify Condition 1.1 of [3] explicitly or provide a self-contained derivation of the M^{-2/3} bound.","section":"§4.5, proof of Theorem 2.9"},{"comment":"Several proofs that are used directly in the main argument are deferred or omitted. The second part of Lemma 4.1 is stated to be 'analogous'; Corollary 5.11 is said to follow as in [19] without proof; Lemma 6.11, the fluctuation averaging lemma, is proved only by saying that the method of [19] or [10] applies; and Lemma A.3 in Appendix A refers to Theorem 8.2 of [19] for 'the left parts.' Since these estimates are load-bearing for the local law and eigenvalue tracking, the manuscript currently does not allow the reader to verify the central claims without consulting the cited works in detail.","section":"Lemma 4.1(4.12), Corollary 5.11, Lemma 6.11, and Appendix A"}],"minor_comments":[{"comment":"The abstract contains the undefined symbol '\\caQ' in 'the largest eigenvalue of \\caQ'; this should be Q. Also, the sentence 'in the limit M, N → ∞ with N/M → d ∈ (0,∞)' later assumes \\hat d is constant; the distinction between \\hat d and d should be clarified.","section":"Abstract and Section 1"},{"comment":"In condition (2.10), the spacing bound uses '(log M)κ_0' but the later estimate (2.17) in the i.i.d. case and Corollary 4.10 contain powers of log M; it would help to state explicitly whether the constants in Definition 2.5 are allowed to vary with n_0, since n_0 is fixed.","section":"§2.2, Definition 2.5"},{"comment":"The text refers to Figure 1 for the histograms, but the figure is not reproduced in the manuscript text; the reader cannot assess the numerical evidence. If this is an artifact of the submission format, the figure legend should at least describe the axes and the observed decay qualitatively.","section":"§2.4, numerical experiments"},{"comment":"The word 'follwing' appears in the sentence before Lemma 5.9; this is a typo. Also, in Corollary 5.11, the bound (5.91) is stated for 'γ ∈ J1,n0−1K' but the sum is over α ∈ Jn0,MK with α≠γ; the latter range already excludes γ, so the condition α≠γ is redundant but harmless.","section":"§5.3, Lemma 5.10"},{"comment":"The statements say 'sample points distributed according to ν' but do not explicitly say 'i.i.d.'; since the appendix is used only for the i.i.d. case, making this explicit would prevent confusion with the deterministic case in Assumption 2.6.","section":"Appendix A, Lemma A.1 and Lemma A.3"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses a genuinely interesting problem and contains a plausible proof strategy with explicit constants. My main concern is that the manuscript relies on several unproved or deferred estimates for its central claims, especially the density bound (2.19), the deterministic-Σ condition (2.16), and the use of Theorem 4.1 of [3] in Theorem 2.9. These are not mere presentation issues: each is load-bearing. If the authors can supply complete proofs or precise verifications of the cited conditions, the paper could be acceptable. The deterministic-Σ claim in the abstract is currently much weaker than it appears, and this should be rephrased or substantiated."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nBottom line: this paper proves a real transition. For sample covariance matrices whose population spectrum has convex decay at the right edge, the largest eigenvalue fluctuates according to a Weibull law when the aspect ratio d exceeds a threshold d_+, and Gaussian when d<d_+ (for i.i.d. population entries). That is new and settles a natural conjecture. The threshold and the Weibull constant are explicit functionals of the population law, with no fitted parameters.\n\nWhat is done well: the proof adapts the Lee–Schnelli deformed Wigner strategy to the sample-covariance setting, and the main difficulty — unbounded diagonal entries in the resolvent — is handled with a real argument. The local law, the fluctuation averaging, and the eigenvalue-tracking are laid out in detail. The paper is honest about the sources of difficulty.\n\nWhere it is soft. First, several load-bearing estimates are not proved in the text: the density bound (2.19) in Theorem 2.7, the second part of Lemma 4.1, Corollary 5.11, and parts of Appendix A are deferred to “analogous to [19]” or “omitted.” For a technical paper extending a known framework that is acceptable, but only if the analogy truly holds; a referee should verify at least one of these transfers, because the sample-covariance resolvent has unbounded diagonal entries that can break a Wigner-matrix estimate.\n\nSecond, and more materially, the deterministic-Σ scope is thinner than the abstract suggests. Assumption 2.6(i) simply postulates the entire good-configuration event Ω, including the uniform averaged convergence (2.16). For i.i.d. Σ, Appendix A shows Ω holds with high probability, so the random case is the real theorem. For a deterministic Σ, the paper gives no primitive condition that would imply (2.16); it is just assumed. The phrase “with some additional assumptions” is therefore doing a lot of work. This is not circular, but a user of the theorem needs to know that checking the assumption is essentially as hard as the proof.\n\nThe subcritical Gaussian result also leans on Theorem 4.1 of [3] to get |λ_1 – L_+| = O(N^{-2/3}); the fit to that theorem is asserted, not checked in detail. Minor, but it is another appeal.\n\nFor whom: random matrix theorists working on edge universality and deformed Marchenko–Pastur laws. It deserves a serious referee. I would accept it for review, with the request that the deferred estimates be proved or precisely stated, and that the deterministic hypothesis be clearly separated and perhaps weakened.\n\nMy recommendation: engage with it. The main i.i.d. result is likely correct and significant.","headline":"Settles the Weibull/Gaussian edge transition for sample covariance matrices with convex population decay; the main i.i.d. result is solid, while the deterministic-Σ scope is conditional on an intricate assumed event.","tokens_in":44712,"tokens_out":5193,"would_cite":true,"duration_ms":45003,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["60B20","62H10","15B52"],"pacs":[],"model":"deepseek-v4-flash","headline":"Above a critical sample-to-population ratio, the largest eigenvalues of a sample covariance matrix are governed by the population's own order statistics and converge to a Weibull law.","keywords":["sample covariance matrix","deformed Marchenko-Pastur distribution","largest eigenvalue","Weibull distribution","order statistics","phase transition","Jacobi measure","extremal eigenvalues"],"falsifier":"Run the numerical experiment at increasing M with a fixed population density (1-t)^b on [l,1], b>1, and d>d_+; collect the top eigenvalue lambda_1 over many trials and compare the empirical distribution of $M^{{1/(b+1)}}$(L_+-lambda_1) with G_{b+1}(s) using the stated C_nu. The claim fails if the fit does not stabilize to that Weibull shape and scale, or if the joint distribution of the top gamma gaps does not match the order-statistics distribution with scaling C_d; a second check is that for d just below d_+ the fluctuations remain of order $M^{{-1/2}}$ and Gaussian.","tokens_in":43720,"feed_emoji":"📊","tokens_out":8683,"duration_ms":77089,"temperature":0.7,"pith_summary":"The paper proves that, for sample covariance matrices with a general diagonal population whose limiting spectrum decays convexly at its right edge, there is a sharp threshold in the dimension ratio: if N/M tends to d above that threshold, the limiting spectrum of the sample matrix keeps the convex decay and the largest sample eigenvalues are determined by the order statistics of the population eigenvalues. In this regime the rescaled gaps $M^{{1/(b+1)}}$(L_+-lambda_gamma) converge jointly to C_d $M^{{1/(b+1)}}$(1-sigma_gamma), so the largest eigenvalue has a Weibull limit instead of the usual Tracy-Widom law. If d lies below the threshold and the population entries are i.i.d., the largest eigenvalue instead fluctuates on the $M^{{-1/2}}$ scale with a Gaussian limit. The result matters because the largest eigenvalues are the ones used for signal detection and data analysis, and the paper gives the exact asymptotic law together with a phase transition in the dimension ratio.","feed_headline":"Above a sample-size threshold, top eigenvalues follow a Weibull law","feed_subtitle":"Once the data-to-dimension ratio is high enough, the sample extremes reproduce the population extremes.","key_machinery":"The argument runs through the deformed Marchenko-Pastur equation for the Stieltjes transform, m_fc(z) = { -z + $d^{{-1}}$ int t dnu(t)/(1+t m_fc(z)) }^{-1}. The threshold d_+ is read off when the quantity H(tau)=$d^{{-1}}$ int $t^{2}$ dnu(t)/((Re tau + t)^2 + (Im tau)^2) at tau=-1 equals d_+/d<1, marking the absence of spectrum outside the bulk and pinning the right edge at L_+=1+tau_+. Near the edge, 1/m_fc(z) is approximately linear in z, 1/m_fc(z)=-1 + d/(d-d_+)(L_+-z) + O((log M)(kappa+eta)^{min{b,2}}), which converts the location of sample eigenvalues into the population order statistics. The technical core is a local law comparing the resolvent's normalized trace m(z) to a random object \\hat{m}_fc(z) defined from the empirical population spectrum, using a block linearization H(z)=[[-z I_N, X^*],[X, -$Sigma^{{-1}}$]] and fluctuation averaging to control the imaginary part of m down to scales below $N^{{-1/2}}$; because the limiting density has no square-root edge, the usual stability bound is replaced by the good-configuration assumptions on Sigma.","core_discovery":"The central claim is Theorem 2.8: under a good-configuration assumption on the top eigenvalues of Sigma, for d>d_+ the joint distribution of $M^{{1/(b+1)}}$(L_+-lambda_i) for i=1,...,gamma converges to the joint distribution of C_d $M^{{1/(b+1)}}$(1-sigma_i), where C_d=(d-d_+)/d. In particular, for i.i.d. population entries the largest eigenvalue of Q satisfies the Weibull limit G_{b+1}(s)=1-exp(-C_nu $s^{{b+1}}$/(b+1)) with C_nu=(d/(d-d_+))^{b+1} lim_{t->1} rho_nu(t)/(1-t)^b. Theorem 2.7 locates the right edge at L_+=1+tau_+ and shows that the limiting density of Q decays like kappa^b at distance kappa from the edge, and Theorem 2.9 gives a centered Gaussian limit of size $M^{{-1/2}}$ for d<d_+ when the population eigenvalues are i.i.d., with variance ($d^{2}$ M)^{-1}(int |t tau/(t+tau)|^2 dnu(t) - (int t tau/(t+tau) dnu(t))^2).","pith_inferences":["The paper does not discuss what happens when the top population gaps are far smaller than M^{-1/(b+1)}; a natural expectation is that the one-to-one order-statistics matching breaks and a spiked or BBP-type regime takes over, since the good-configuration event explicitly excludes that case.","The exponent b+1 in the Weibull law makes the effective tail index of the population edge; for fixed d>d_+, a smaller b gives heavier tails in the rescaled largest eigenvalue, which is a testable prediction for simulations.","The same peak-tracking mechanism via the imaginary part of the resolvent may transfer to other ensembles whose limiting density has non-square-root edge decay, beyond the deformed Wigner and sample covariance settings treated here.","The sharp scale change from M^{-1/2} Gaussian fluctuations below d_+ to M^{-1/(b+1)} Weibull fluctuations above d_+ suggests a genuine phase transition in the fluctuation size; checking the transition width near d_+ numerically would be a natural follow-up."],"forward_implications":["For d>d_+, the top gamma sample eigenvalues are asymptotically the same random object as the top gamma population eigenvalues, rescaled by C_d; the largest eigenvalue is Weibull with exponent b+1 rather than Tracy-Widom.","The sample spectral density has convex decay at the right edge, mu_fc(L_+-kappa) ~ kappa^b, so the local edge regime changes its behavior when the ratio d crosses d_+.","The boundary d_+ = int t^2 (1-t)^{-2} dnu(t) is explicit in the population law, so one can predict from Sigma alone which regime a dataset is in.","Under Gaussian sampling the same results hold for non-diagonal Sigma, so the finding is not an artifact of the diagonal model.","Corollary 4.10 gives quantitative rates, with errors involving M^{3phi}/M^b and (log M)^2/M^{1/(b+1)} plus the bad-configuration probability (log M)^{1+2b}/M^phi."],"supporting_citations":[{"why":"Defines the deformed Marchenko-Pastur law and the self-consistent equation for the limiting spectral distribution that the paper starts from.","marker":"[22]"},{"why":"Supplies the strategy for analyzing the self-consistent equation and locating extremal eigenvalues when the edge decay is convex rather than square-root; the local-law and fluctuation-averaging steps adapt its method.","marker":"[19]"},{"why":"Gives the local deformed semicircle law for Wigner-type matrices used as the template for the edge analysis and the linearization near the edge.","marker":"[18]"},{"why":"Provides the block linearization of Q and the resolvent identities that carry the local-law proof for real sample covariance matrices.","marker":"[20]"},{"why":"Establishes the anisotropic local-law framework whose square-root edge stability is absent here; the paper explicitly notes it cannot use that stability bound and must work around it.","marker":"[17]"},{"why":"Supplies the universality and fluctuation-averaging estimates for covariance matrices used to control imaginary parts of the resolvent.","marker":"[25]"},{"why":"Proves a Tracy-Widom limit for general population; the paper's Weibull regime is contrasted with this square-root-edge benchmark.","marker":"[6]"},{"why":"Gives the Condition 1.1 used in Theorem 2.9's proof to identify the Gaussian-dominated regime for d<d_+.","marker":"[3]"}],"fun_headline_variants":["Above sample threshold, top eigenvalues go Weibull","Weibull law for sample extremes when dimension ratio high","Threshold splits sample eigenvalue extremes: Weibull vs Gaussian","Population order statistics drive sample top eigenvalues"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole comparison rests on the good-configuration event $\\Omega$: the largest n0 eigenvalues of Sigma must be separated from each other and from the edge 1 by gaps of order $M^{{-1/(b+1)}}$ but not much larger, and the averaged sums involving Sigma must converge to the limiting integrals; if those gaps are missing, the sample eigenvalues are not matched one-to-one to population order statistics and the Weibull conclusion can fail.","fun_headline_variants_meta":{"raw":{"variants":["Above sample threshold, top eigenvalues go Weibull","Weibull law for sample extremes when dimension ratio high","Threshold splits sample eigenvalue extremes: Weibull vs Gaussian","Population order statistics drive sample top eigenvalues"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000196,"raw_usage":{"total_tokens":1419,"prompt_tokens":1065,"completion_tokens":354,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":681,"completion_tokens_details":{"reasoning_tokens":292}},"tokens_in":681,"tokens_out":354,"duration_ms":3486,"temperature":1.0,"reasoning_tokens":292,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T12:51:50.377694+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the numerical experiment at increasing M with a fixed population density (1-t)^b on [l,1], b>1, and d>d_+; collect the top eigenvalue lambda_1 over many trials and compare the empirical distribution of $M^{{1/(b+1)}}$(L_+-lambda_1) with G_{b+1}(s) using the stated C_nu. The claim fails if the fit does not stabilize to that Weibull shape and scale, or if the joint distribution of the top gamma gaps does not match the order-statistics distribution with scaling C_d; a second check is that for d just below d_+ the fluctuations remain of order $M^{{-1/2}}$ and Gaussian.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the deformed Marchenko-Pastur law and the self-consistent equation for the limiting spectral distribution that the paper starts from."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the strategy for analyzing the self-consistent equation and locating extremal eigenvalues when the edge decay is convex rather than square-root; the local-law and fluctuation-averaging steps adapt its method."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Gives the local deformed semicircle law for Wigner-type matrices used as the template for the edge analysis and the linearization near the edge."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the block linearization of Q and the resolvent identities that carry the local-law proof for real sample covariance matrices."},{"cited_title":"Knowles and J","cited_arxiv_id":null,"evidence_quote":"Establishes the anisotropic local-law framework whose square-root edge stability is absent here; the paper explicitly notes it cannot use that stability bound and must work around it."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the universality and fluctuation-averaging estimates for covariance matrices used to control imaginary parts of the resolvent."},{"cited_title":"El Karoui","cited_arxiv_id":null,"evidence_quote":"Proves a Tracy-Widom limit for general population; the paper's Weibull regime is contrasted with this square-root-edge benchmark."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Gives the Condition 1.1 used in Theorem 2.9's proof to identify the Gaussian-dominated regime for d<d_+."}],"review_version":1}