{"id":"df6b6eb8-4f2f-44e7-942a-d4606b5083ef","arxiv_id":"2507.09584","paper_version":2,"verdict":"REJECT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Non-Gaussian spiked eigenvalues get a first-order Edgeworth expansion, enabling sharper confidence intervals and spike-number estimators.","lead":"This paper derives Edgeworth corrections for the largest spiked eigenvalues of sample covariance matrices when the data are not Gaussian. The corrections lead to more accurate confidence intervals and spike-count estimates, mainly in low-dimensional settings.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 6 compares unstandardized statistics whose variances differ by O(1) when beta_z != 0; the moment-matching step is invalid, so the claimed o(n^{-1/2}) universality cannot hold.","rationale":"The reader located the gap in Theorem 6's rate. I agree that Theorem 6 is the linchpin, but the more serious problem is that the theorem compares unstandardized statistics whose limiting variances differ by \\beta_z F^2(g) whenever the fourth cumulant is nonzero. This is visible from the paper's own §A.5 variance formula and from the moment-matching step in §A.6, which cannot hold because E z^4 \\neq E y^4. Hence even the leading-order CLT universality fails, not just the Edgeworth rate. A corrected proof would need either to compare statistics after standardizing each by its own variance (which would change the Edgeworth correction and the final scaling), or to impose \\beta_z=0, which would remove the non-Gaussian effect the paper claims to capture. The simulations in Section 4 cannot rescue the proof because they use estimated coefficients and do not test the o(n^{-1/2}) rate; the theoretical claim remains unsupported. Therefore the reader's REJECT verdict stands, and no verdict adjustment is needed.","tokens_in":57668,"tokens_out":15500,"duration_ms":166941,"concrete_test":"Monte Carlo check: set n=200 and n=800, p/n=0.1, single spike l=4, V1=e1, Z entries centered Gamma with E Z^4=5 (beta_z=2) and Y standard Gaussian. Compute 10^4 replications of the two statistics in Theorem 6 (with \\tilde{\\sigma}_n containing \\pi=\\beta_z) at x=0, and estimate their variances and the Kolmogorov distance between their empirical CDFs. If the variance difference does not shrink to 0 (and the CDF distance stays bounded away from 0 as n increases), Theorem 6's o(n^{-1/2}) claim is contradicted. Equivalently, an analytic check: compare Var(\\tilde{S}_n(gn)) - Var(S_n(gn)) \\to \\beta_z F^2(g)>0 using the formula in §A.5.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central proof chain (Theorem 2, equations (1.4)-(1.6)) depends on Theorem 6, which asserts that the unstandardized non-Gaussian statistic \\tilde{S}_n(gn)+n^{-1/2}\\tilde{S}_n(gn h_n)\\rho_n^{-1}\\tilde{\\sigma}_n x and its Gaussian counterpart S_n(gn)+n^{-1/2}S_n(gn h_n)\\rho_n^{-1}\\tilde{\\sigma}_n x have distributions differing by o(n^{-1/2}). This is not merely an unproved rate; it is incompatible with the paper's own variance computation. At the end of §A.5 the authors derive Var(\\Omega(\\rho_n,Z)) = 2\\rho^2 F(g^2) + \\rho^2 \\beta_z F^2(g) and \\tilde{\\sigma}_n^2 = (2F(g^2)+\\pi F^2(g))/F^2(g^2) with \\pi = \\lim \\sum v_{tk}^4 \\beta_z. Since \\tilde{S}_n = -\\Omega/\\rho_n, the non-Gaussian statistic has variance 2F(g^2)+\\beta_z F^2(g), while the Gaussian statistic (for which \\beta_z=0) has variance 2F(g^2). These differ by \\beta_z F^2(g), an O(1) gap, so the limiting distributions are different normals. The proof in §A.6 says it 'matches the first four moments of z_k and y_k', but E z_k^4 = 3+\\beta_z and E y_k^4=3, so the swap from z_k to y_k changes the fourth cumulant; the telescoping characteristic-function argument cannot make the accumulated O(1) variance difference vanish. Thus Theorem 6 fails as stated for \\beta_z \\neq 0, and the subsequent Edgeworth expansion of Theorem 1 is not established by the proof given.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper claims to close the open problem, posed by Yang and Johnstone (2018), of first-order Edgeworth corrections for spiked eigenvalues of sample covariance matrices with non-Gaussian entries. Theorem 1 (multi-spike) and Theorem 2 (single-spike) assert sup_x |P(R_k <= x) - Phi(x) - n^{-1/2}[(1/6)kappa_{2,k}^{-3/2}kappa_{3,k}(1-x^2) - kappa_{2,k}^{-1/2}(mu(g_{nk}) + A(g_{nk}))]phi(x)| = o(n^{-1/2}) with explicit formulas for kappa_{2,k}, kappa_{3,k}, mu(g_{nk}), and A(g_{nk}). Theorem 3 provides consistent estimators of the fourth- and sixth-moment parameters beta_z, Gamma, and Delta; Theorem 4 claims O(n^{-1/2}) and o(n^{-1/2}) coverage errors for the Z-type and E-type pivots; Section 3.2 constructs a spike-number estimator; Section 4 reports extensive simulations. The proof is organized as three steps: Theorem 5 approximates R_n by a linear statistic of the non-Gaussian matrix, Theorem 6 asserts an o(n^{-1/2}) non-Gaussian/Gaussian universality via a partial generalized four moment theorem, and Theorem 7 gives the conditional Edgeworth expansion for the Gaussian statistic. The critical load-bearing step is Theorem 6, whose statement conflicts with the paper's own variance computation in Section A.5.","tokens_in":58086,"tokens_out":36561,"duration_ms":368336,"significance":"If the main results held, this would be a useful contribution: explicit, parameter-free Edgeworth corrections for spiked eigenvalues in non-Gaussian models, consistent cumulant estimators, and a credible route to confidence intervals and spike counting. The paper deserves credit for writing out the correction terms (including the cross-spike interaction A(g_{nk})), for honestly reporting in Section 4.2 that estimated Edgeworth coefficients can degrade performance, and for structuring a serious proof attempt rather than a heuristic derivation. The decisive universality claim of Theorem 6 is, however, contradicted by the paper's own variance formula in Section A.5: the non-Gaussian and Gaussian statistics compared there have asymptotic variances differing by beta_z l^{-2}, an O(1) gap, so the claimed o(n^{-1/2}) equivalence cannot hold. Theorem 6 is the bridge in the proofs of Theorems 1 and 2 (equations (1.1)-(1.3) and (1.4)-(1.6)), and its proof also depends on an unproved rate assertion (Remark 9). The main theorems, and the applications built on them (Theorem 4 and Section 3), are therefore not established by the proof given.","major_comments":[{"comment":"The stress-test concern about an O(1) variance gap is confirmed on reading the paper. Theorem 6 compares the unstandardized statistics tilde S_n(g_n) + n^{-1/2} tilde S_n(g_n h_n) rho_n^{-1} tilde sigma_n x and S_n(g_n) + n^{-1/2} S_n(g_n h_n) rho_n^{-1} tilde sigma_n x, and claims their distribution functions differ by o(n^{-1/2}). But from the identity Omega(rho_n,Z) = -rho_n tilde S_n(g_n) in Section 5.2 and the paper's own variance computation at the end of Section A.5, Var(Omega(rho_n,Z)) = 2 rho^2 F(g^2) + rho^2 beta_z F^2(g), hence Var(tilde S_n(g_n)) tends to 2F(g^2) + beta_z F^2(g), whereas the Gaussian term S_n(g_n) = n^{-1/2} sum g_n(lambda_i)(omega_i^2 - 1) has Var tending to 2F(g^2) because beta_z = 0 for Y. The gap beta_z F^2(g) = beta_z l^{-2} is O(1). Consequently the two CDFs in Theorem 6 have sup-norm distance bounded away from zero; for instance, the Gaussian event in Section 5.2 has limiting probability Phi(a x) with a = sqrt(1 + beta_z l^{-2} sigma_n^2/4) different from 1, while the non-Gaussian event has limit Phi(x). The proof in Section A.6 says it 'matches the first four moments of z_k and y_k,' but E z_k^4 = 3 + beta_z differs from E y_k^4 = 3, so the fourth moment does not match, and the telescoping characteristic-function argument cannot remove an O(1) variance difference. Since Theorem 6 is the step converting the non-Gaussian statistic to its Gaussian counterpart in equations (1.4)-(1.6) of Section A.2 and (1.1)-(1.3) of Section A.1, the proofs of Theorem 2 and Theorem 1 do not go through.","section":"Theorem 6; Section 5.2; Section A.5; Section A.6"},{"comment":"Even setting aside the variance problem, the proof of Theorem 6 does not establish the claimed rate. The final telescoping step in Section A.6 asserts o(n^{-1/2}) 'established through Lemmas 1, 2 and 3, along with Remark 10,' but Lemma 1 gives only E_k(alpha_{ki0}) = O(n^{-1/2}) for the individual conditional expectations, and summing such bounds over k would leave a contribution of order n^{1/2} without cancellation. The decisive cancellation that would close the argument is placed entirely on the unproved assertion in Remark 10 that alpha_{ki0} - alpha_{ki0y} is of order O(n^{-1}) 'demonstrated' by Jiang and Bai (2021b), with no derivation given. Remark 9, which is the actual load-bearing rate assumption of Theorem 6, states without proof that the conclusion of the partial generalized four moment theorem 'remains valid, and the asymptotic error bound can be shown to be of order o(n^{-1/2}).' The appendices repeatedly defer bounds with phrases such as 'the proofs are similar' (Lemmas 1-3 in Section A.6) and 'using the same method' (Appendix B.4), so the missing rate is not documented elsewhere in the manuscript. This is an omitted proof of a load-bearing step.","section":"Remark 9; Section A.6; Remark 10"},{"comment":"Theorem 4(2) claims a coverage error of o(n^{-1/2}) for the E-type pivot, but the proof in Section A.4 obtains the bound |u^E_n(hat rho_k, l_k) - bar F_{kn}(hat rho_k, l)| <= C_n n^{-1/2} with only the sentence 'building upon our theoretical framework established in previous sections.' This bound is essentially the uniform Edgeworth approximation of Theorem 1 itself, so Theorem 4(2) inherits the failure of Theorem 1 identified above. The argument also does not show why the post-selection conditioning on hat l_k > theta_n preserves the o(n^{-1/2}) rate rather than only the O(n^{-1/2}) rate claimed for the Z-type pivot, and the displayed probability calculation does not by itself establish the conditional claim. The confidence-interval construction in Section 3.1 and the spike-number estimator in Section 3.2 therefore rest on results that are not established.","section":"Theorem 4; Section A.4"}],"minor_comments":[{"comment":"There are numerous typos: 'eatimator' (Section 3.1), 'converagence' (Remark 6), 'varibales' (Theorem 7), 'Corolllary' (Lemma 6), and 'compansion' for 'companion' (repeatedly in Section A.1); the paper would benefit from a careful proofreading pass.","section":"Throughout"},{"comment":"Remark 5 cites 'Theorem 2.7 of Zheng et al. (2019)', but the reference list contains Zheng, Bai and Yao (2015) and Zhang, Hu and Bai (2019), and no Zheng et al. (2019); the citation should be corrected.","section":"Section 2.2; Remark 5"},{"comment":"Theorem 6 compares statistics that contain a fixed x while asserting convergence 'for any t in R'; the theorem should state explicitly that x is fixed and whether the rate is uniform in x, since the application in Theorem 5 needs uniformity in x over the real line.","section":"Theorem 6 statement"},{"comment":"In the final display of the proof of Theorem 5, the term 'n^{-1/2} tilde S_n(g_n)' appears twice with different implied coefficients, and the intermediate algebra leading to the displayed equation is difficult to verify; the authors should check for a typesetting slip and expand the derivation.","section":"Section A.5, final display"},{"comment":"The verification of condition R3 of Lemma 6 reads 'n^{1/2} integral_{|t|>epsilon} |t|^{-1} |E exp(it bar V_n^{-1/2} sum X_{ni})| dt <= n^{1/2} integral_{|t|>epsilon} |t|^{-1} delta_n^i dt = o(1)' with delta_n^i undefined; since R3 is genuinely needed for the o(n^{-1/2}) rate, this bound needs a real argument rather than an undefined symbol.","section":"Section A.7, condition R3"},{"comment":"In Tables 3-6 the Y&J-E method reports 0% estimation accuracy in several high-dimensional settings (for example Table 3, rows (80,400), (60,200), and (120,400)), and accuracy is sometimes non-monotone in n; the discussion in Section 4.2 attributes this to coefficient estimation difficulty but does not explain the mechanism or the non-monotonicity, which is important for assessing the practical claims.","section":"Section 4.2; Tables 3-6"},{"comment":"Remark 3 states that kappa_{2,k} and kappa_{3,k} are 'not the exact conditional cumulants of Z11 but rather carefully constructed approximations,' while Theorem 1 asserts an o(n^{-1/2}) expansion with these exact formulas; the paper should clarify in what sense the approximate cumulants are sufficient for the claimed exact rate, since a reader cannot tell from the text whether a remainder has been absorbed.","section":"Remark 3"},{"comment":"Equation (1.9) contains the term -(Delta + 12 hat beta_z + 6/(1 - gamma_n) - 15 hat beta_z - 21 + 8(1 + gamma_n)/(1 - gamma_n)^2)^2, whose signs on the beta_z terms are opposite to those in the definition (2.5); as written the display also appears to square a random variable (it contains hat beta_z) rather than a constant, so the formula for E(hat Delta - Delta)^2 should be checked.","section":"Section A.3, near equation (1.9)"}],"recommendation":"reject","confidential_remarks":"The referee report is intentionally decisive: Theorem 6 is falsified by the paper's own variance formula, not merely unproved, and the main theorems depend on it directly. One further consideration for the editor: the decisive rate claims (Remark 9 and Remark 10) are attributed to Jiang and Bai (2021b), a paper co-authored by the current manuscript's third author; this is not improper because the cited work is peer-reviewed, but it does mean the central novel step rests on an unevaluated extension of a result from the authors' own group. Given the structural nature of the error, a major revision would require a substantially new universality argument with variance-matched statistics, and it is unclear whether the claimed Edgeworth formula would survive; hence I recommend rejection rather than major revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper claims to resolve the non-Gaussian Edgeworth problem for spiked eigenvalues, and the explicit formulas for the cumulant corrections and the cross-spike term A(g_nk) are genuinely new. The simulations are extensive and show the Edgeworth density fits non-Gaussian data well, which is nice evidence that something along these lines is true. But the central proof chain does not hold together, and the flaw is not a missing detail; it is a contradiction with the paper's own calculations.\n\nThe load-bearing step is Theorem 6, which asserts that the unstandardized statistics \\tilde{S}_n(g_n) + n^{-1/2}\\tilde{S}_n(g_n h_n)\\rho^{-1}\\tilde{\\sigma}_n x and S_n(g_n) + n^{-1/2}S_n(g_n h_n)\\rho^{-1}\\tilde{\\sigma}_n x have distributions differing by o(n^{-1/2}). The paper itself computes in §A.5 that Var(Ω(ρ_n,Z)) = 2ρ^2 F(g^2) + ρ^2 β_z F^2(g), which implies Var(\\tilde{S}_n(g_n)) = 2F(g^2) + β_z F^2(g). The Gaussian counterpart S_n(g_n) has variance 2F(g^2). For β_z ≠ 0 these differ by an O(1) amount, so the two CDFs cannot be o(n^{-1/2})-close. The proof's moment-matching step changes the fourth cumulant (E z^4 = 3+β_z vs. E y^4 = 3) and cannot erase an O(1) variance gap. So Theorem 6 fails as stated, and with it the Edgeworth expansions of Theorems 1 and 2 are not established.\n\nThere are other soft spots that would matter even if Theorem 6 were fixed: the o(n^{-1/2}) rate in Remark 9 is asserted rather than proved, many term bounds are deferred with 'similar' arguments, and Theorem 4's post-selection claim is not really substantiated. The estimation of the moments in Section 2.2 is a useful auxiliary result, but it does not rescue the main argument.\n\nFor a reader, the value is mainly the explicit formulas and the demonstration of what a non-Gaussian Edgeworth expansion should look like. I would not cite the main theorems in their current form. The paper should not be accepted as is, but I would not desk-reject it either; the problem is important, the approach is serious, and a referee might help the authors see exactly where the standardization needs to be corrected. If it crosses your desk, send it to a referee with a specific request to check whether Theorem 6 is compatible with the variance computation in §A.5.","headline":"A serious attempt at a real open problem, but the key universality theorem contradicts the paper's own variance computation.","tokens_in":58607,"tokens_out":19383,"would_cite":false,"duration_ms":181781,"reading_group":"maybe","serious_thinker":"no","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62E20","60B20","62H25"],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper proves, under four moment and smoothness assumptions, a first-order Edgeworth expansion for the spiked eigenvalues of sample covariance matrices with non-Gaussian entries, reducing the approximation error to $o(n^{-1/2})$ and…","keywords":["Edgeworth expansion","spiked covariance model","largest eigenvalue","non-Gaussian data","partial generalized four moment theorem","confidence intervals","number of spikes","high-dimensional statistics"],"falsifier":"Compute, at increasing $n$ and for an entry distribution that matches a Gaussian only in its first four moments (for example a three-point lattice with zero mean, unit variance and zero third moment), the sup-norm distance between the empirical distribution functions of the two statistics $\\Omega_s(Z_1,Z)$ and $\\Omega_s(Z_1,Y)$ defined in Theorem 6; if this distance does not decay faster than $n^{-1/2}$, the central expansion of Theorem 1 is false.","tokens_in":57451,"feed_emoji":"📈","tokens_out":8586,"duration_ms":86467,"temperature":0.7,"pith_summary":"This paper establishes a first-order Edgeworth expansion for the distribution of the leading sample eigenvalues in a spiked covariance model when the data are not Gaussian, an extension left open by the Gaussian treatment of Yang and Johnstone (2018). The expansion adds an explicit $n^{-1/2}$ correction term built from the third cumulant of the squared entries and from the other spikes, shrinking the error of the normal approximation from $O(n^{-1/2})$ to $o(n^{-1/2})$. If the result is right, confidence intervals for population spikes can be shortened to a level of accuracy previously available only under Gaussianity, and estimation of the number of spikes becomes more reliable in low-dimensional settings. The paper also provides consistent estimators for the entry moments needed to compute the correction, and simulations show the corrected densities track the empirical ones far better than the Gaussian approximation.","feed_headline":"Non-Gaussian spiked eigenvalues get sharper limits","feed_subtitle":"A first-order Edgeworth expansion refines confidence intervals and spike-counting beyond the Gaussian theory.","key_machinery":"The argument rests on three approximations chained together. First, Theorem 5 shows the eigenvalue statistic can be replaced, up to $o(n^{-1/2})$, by a linear spectral statistic $\\Omega(\\rho_n, Z)$ that is a sum of independent contributions after conditioning on the noise spectrum. Second, Theorem 6 uses the partial generalized four-moment theorem (PG4MT) of Jiang and Bai (2021b) to assert that the distribution of this statistic is the same for genuinely non-Gaussian entries $Z$ and for Gaussian entries $Y$, again up to $o(n^{-1/2})$ — this is the step that converts the Gaussian-only analysis into a universal one. Third, Theorem 7 feeds the Gaussian version, written as $n^{-1/2}\\sum_i c_{ni}(W_i^2 - 1)$ conditionally on the eigenvalues of $n^{-1}YY'$, into the classical Edgeworth expansion for sums of independent variables (Petrov 1975; the lemma is inherited from Yang and Johnstone 2018), producing the explicit $\\Phi + n^{-1/2}p_1\\phi$ form. The cumulants $\\kappa_{2,k}$, $\\kappa_{3,k}$ and the mean shift $\\mu(g_{nk})$ enter exactly at this last step.","core_discovery":"The central discovery is a universality statement with a rate: the distribution of the normalized spiked eigenvalue statistic\n$$R_k = $n^{{1/2}}$(\\hat l_k - \\rho_{nk})/\\tilde\\sigma_{nk}$$\nis, up to an error of $o(n^{-1/2})$, independent of the entry distribution except through two low-order cumulants of $Z_{11}^2$ and the cross-spike interaction term $A(g_{nk})$. Concretely, Theorem 1 gives\n$$\\sup_x \\left|P(R_k \\le x) - \\Phi(x) - $n^{{-1/2}}$\\left[\\tfrac16 \\kappa_{2,k}^{-3/2}\\kappa_{3,k}(1-$x^{2}$) - \\kappa_{2,k}^{-1/2}(\\mu(g_{nk}) + A(g_{nk}))\\right]\\$\\varphi$(x)\\right| = o($n^{{-1/2}}$),$$\nwhere $\\kappa_{2,k},\\kappa_{3,k}$ are constructed cumulants of $\\tilde Z_{1k}^2 -1$ and $A(g_{nk})$ sums contributions from the other population spikes. The same structure resolves the open problem of Yang and Johnstone (2018) for the single-spike case and, for multiple spikes, shows that each spiked eigenvalue's Edgeworth correction depends not only on its own spike but on all spikes through $A(g_{nk})$.","pith_inferences":["The $o(n^{-1/2})$ universality suggests the same three-step strategy (linear-statistic reduction, four-moment comparison, conditional Edgeworth expansion) could be exported to other spiked ensembles — correlation matrices, F-matrices, or beta ensembles — where only Gaussian or first-order CLT results exist; the paper does not claim this.","Because the Edgeworth correction is dominated by the third cumulant of the squared entries, the method is most valuable for skewed or heavy-tailed data; for nearly Gaussian data the estimated correction adds variance without much bias reduction, which may explain the simulation cases where the Edgeworth-based method trails its Gaussian rival.","The explicit dependence of $A(g_{nk})$ on the gap $l_k - l_j$ implies that when two spikes are close, the correction is large and single-spike formulas mislead; a testable prediction is that the E-type confidence interval for the smaller spike degrades as $l_k \\to l_j$, a quality the paper does not examine.","The moment estimators rely on leave-one-out inverses and become unstable when $p$ is close to $n$; a practical extension would be to replace them by ridge-regularized inverses, mirroring the pseudo-inverse treatment already used for $p>n$ in the paper."],"forward_implications":["Confidence intervals for a population spike built from the Edgeworth-corrected E-type pivot have coverage error $o(n^{-1/2})$, one order better than the $O(n^{-1/2})$ error of the Gaussian Z-type pivot (Theorem 4).","In the multi-spike case, the correction contains the explicit interaction term $A(g_{nk}) = \\frac{l_k-1}{(l_k-1)^2-\\gamma_n}\\sum_{j\\ne k}\\frac{l_j-1}{l_k-l_j}$, so the distribution of the $k$-th spiked eigenvalue depends on all other spikes, not just on $l_k$.","A new cardinality estimator $\\hat r = \\sum_{k=1}^p I(\\hat l_k \\in C_k)$ built on Edgeworth-corrected intervals outperforms existing spike-counting methods in low-dimensional (small $n,p$) regimes across Gamma, Uniform, and Gaussian data.","The moments $\\beta_z$, $\\Gamma$ and $\\Delta$ needed to compute the correction are consistently estimated from leave-one-out inverses of $S$, making the expansion implementable without prior knowledge of the entry distribution (Theorem 3).","The authors state that a second-order Edgeworth expansion remains open because it would require a first-order approximation for the associated linear spectral statistic under non-Gaussianity."],"supporting_citations":[{"why":"Supplies the Gaussian Edgeworth expansion and the lemma (their Corollary 4) for sums of independent variables that Theorem 7 adapts.","marker":"Yang and Johnstone (2018)"},{"why":"Provides the generalized four moment theorem and the base CLT for spiked eigenvalues whose asymptotic distribution is refined here.","marker":"Jiang and Bai (2021a)"},{"why":"Provides the partial generalized four moment theorem whose asserted $o(n^{-1/2})$ rate is the load-bearing step of Theorem 6.","marker":"Jiang and Bai (2021b)"},{"why":"Classical Edgeworth expansion for sums of independent random variables invoked in the proof of Theorem 7.","marker":"Petrov (1975)"},{"why":"Supplies the Cornish-Fisher expansion used to convert the Edgeworth correction into corrected quantiles $E_\\alpha$.","marker":"Hall (2013)"},{"why":"CLT for linear spectral statistics used in truncation arguments and moment calculations throughout the proofs.","marker":"Bai and Silverstein (2004)"},{"why":"CLT for linear spectral statistics of non-Gaussian entries that yields the mean term $\\mu(g_n)$ in the expansion.","marker":"Wang and Yao (2013)"},{"why":"Provides the consistency of the estimator of $\\beta_z$ (fourth cumulant) on which Theorem 3 builds.","marker":"Zheng et al. (2019)"},{"why":"Gaussian multi-spike Edgeworth corrections and post-selection inference framework that the applications extend.","marker":"Yang (2019)"},{"why":"Stochastic decomposition for sample covariance matrices from a spiked population, used in the multi-spike proof.","marker":"Wang et al. (2014)"}],"fun_headline_variants":["Non-Gaussian spike eigenvalues get refined limits","Edgeworth corrections go non-Gaussian for spiked eigenvalues","Sharper confidence intervals for non-Gaussian spiked spectra","Beyond Gaussian: better spike detection and intervals","Non-Gaussian spiked eigenvalues: first-order Edgeworth done"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the partial generalized four moment theorem delivers a distributional error of order $o(n^{-1/2})$ for the specific linear statistics $\\Omega_s(Z_1,Z)$ versus $\\Omega_s(Z_1,Y)$; this rate is asserted in Remark 9 with the detailed cancellations deferred, and if it is invalid the non-Gaussian Edgeworth expansion is not established.","fun_headline_variants_meta":{"raw":{"variants":["Non-Gaussian spike eigenvalues get refined limits","Edgeworth corrections go non-Gaussian for spiked eigenvalues","Sharper confidence intervals for non-Gaussian spiked spectra","Beyond Gaussian: better spike detection and intervals","Non-Gaussian spiked eigenvalues: first-order Edgeworth done"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000169,"raw_usage":{"total_tokens":1260,"prompt_tokens":939,"completion_tokens":321,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":555,"completion_tokens_details":{"reasoning_tokens":259}},"tokens_in":555,"tokens_out":321,"duration_ms":3578,"temperature":1.0,"reasoning_tokens":259,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T17:52:18.576169+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute, at increasing $n$ and for an entry distribution that matches a Gaussian only in its first four moments (for example a three-point lattice with zero mean, unit variance and zero third moment), the sup-norm distance between the empirical distribution functions of the two statistics $\\Omega_s(Z_1,Z)$ and $\\Omega_s(Z_1,Y)$ defined in Theorem 6; if this distance does not decay faster than $n^{-1/2}$, the central expansion of Theorem 1 is false.","supporting_citations":[{"cited_title":"and Johnstone, I","cited_arxiv_id":null,"evidence_quote":"Supplies the Gaussian Edgeworth expansion and the lemma (their Corollary 4) for sums of independent variables that Theorem 7 adapts."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Classical Edgeworth expansion for sums of independent random variables invoked in the proof of Theorem 7."},{"cited_title":"and Silverstein, J","cited_arxiv_id":null,"evidence_quote":"CLT for linear spectral statistics used in truncation arguments and moment calculations throughout the proofs."},{"cited_title":"and Yao, J","cited_arxiv_id":null,"evidence_quote":"CLT for linear spectral statistics of non-Gaussian entries that yields the mean term $\\mu(g_n)$ in the expansion."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Gaussian multi-spike Edgeworth corrections and post-selection inference framework that the applications extend."},{"cited_title":"W., and Yao, J.-f","cited_arxiv_id":null,"evidence_quote":"Stochastic decomposition for sample covariance matrices from a spiked population, used in the multi-spike proof."}],"review_version":1}