{"id":"dc77f42b-cad1-4a01-80d9-0c1aa7d2dd04","arxiv_id":"2608.09135","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The top eigenvalues and eigenvectors of sample covariance matrices under entrywise missing data follow Gaussian limits whose parameters depend on the missingness probabilities.","lead":"A statistics paper proves central limit theorems for the top eigenvalues and eigenvectors of sample covariance matrices when some data entries are missing. The results show that missing data changes the limiting distributions in a predictable way and are used to build a test for whether two groups of variables are independent under a spiked population model.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 4.2's power analysis is for a block-diagonal (independent) alternative, not for the stated H1 of dependence, so the independence test's power claim is unsupported.","rationale":"In good faith, the main theoretical results of Sections 2 and 3 appear plausible: the proof of Theorem 2.5 follows the standard random-matrix perturbation scheme, and the covariance formulas in Theorems 2.5 and 2.9 are internally consistent with the missing-at-random zero-imputation model. The reader's weakest assumption about MCAR is a genuine limitation but is explicitly assumed, so it is not an internal flaw. The more load-bearing concern is in the application: Theorem 4.2 analyzes an alternative hypothesis that is not the stated H1. A block-diagonal covariance with extra spikes is still an independent structure under the paper's data-generating model, so the power theorem proves consistency against a different deviation, not against dependence. This directly affects the paper's advertised contribution of testing independence and leaves the empirical power results in Table 3 without theoretical support. Because the main CLT theorems may still be correct and the application section could be repaired by either restating the alternative or deriving a genuinely dependent perturbative model, the reader's CONDITIONAL verdict remains appropriate; no change is needed. The concern strengthens the justification for requiring a rigorous, correctly specified power analysis before the test is used.","tokens_in":22192,"tokens_out":15451,"duration_ms":144255,"concrete_test":"Construct the minimal dependence model with two spikes: take Σ_N = diag(α_1, α_2), Σ_S = I, and add a single covariance ρ between one coordinate of x_N and one coordinate of x_S. The resulting population covariance is not block-diagonal, so the Section 2 model and Theorem 4.1's covariance formulas do not apply directly. Simulate the statistic T_N under this model with ρ = 0.5, n = 1000, matching Table 3 settings, and record the rejection rate at α = 0.05. If the rejection rate does not approach 1 as n grows, Theorem 4.2's power claim fails for its stated H1. Alternatively, re-derive the proof of Theorem 4.2 and verify that it only uses the block-diagonal covariance structure, which would imply the same proof goes through under H0 with N+L spikes, showing the statistic cannot distinguish independence from dependence.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4.1 states H0: x_N is independent of x_S versus H1: x_N is not independent of x_S. Theorem 4.2 then says \"Under H1, i.e., when the population covariance matrix satisfies Σ = diag(Σ_{N+L}, Σ_{S−L})\". But with the independent-component model of Section 2, a block-diagonal covariance makes x_N and x_S independent: Σ^{1/2} is block-diagonal, so the two groups depend on disjoint iid coordinates. The proof's inequality eα^{[N+L]}_j > eα_j establishes power against the wrong alternative (additional large eigenvalues in the first block), not against any nonzero cross-covariance. Thus Theorem 4.2, and Theorem 4.5 which inherits its argument, do not cover the stated H1. The empirical power simulations in Table 3, which introduce pairwise covariance 0.5 between variables in x_N and x_S, are not covered by either theorem. This is an internal inconsistency in the application section, not merely a limitation of the model class.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies the asymptotic eigen-structure of sample covariance matrices under a spiked population model with entrywise missing observations. Missingness is modeled as y = b ◦ x, where b is an independent Bernoulli mask and x has an independent-component structure, so the effective population covariance is tilde Σ = PΣP + (P − P^2) ◦ Σ. Under assumptions on the effective spiked eigenvalues, the authors state an almost sure convergence result for distant spiked eigenvalues (Proposition 2.1), a central limit theorem for the spiked eigenvalues (Theorems 2.5 and 2.7), and CLTs for the associated eigenvectors (Theorems 2.8 and 2.9), with explicit limiting covariance matrices. The paper then applies these results to propose a chi-square statistic for testing independence between two groups of variables (Theorem 4.1) and a plug-in version (Theorem 4.3), supported by simulations in Section 3 and Section 4.2.","tokens_in":22321,"tokens_out":8933,"duration_ms":86803,"significance":"If the main CLT results are correct, the paper makes a useful contribution by extending spiked PCA inference to data with missing-at-random entries and by showing that the missingness mechanism changes the limiting variances and the eigenvector basis in a nontrivial way. The explicit covariance formulas in Theorems 2.5 and 2.9 are concrete and falsifiable, and the simulations in Section 3 provide a prima facie check of the Gaussian approximations. The reliance on the published machinery of Li et al. (2024) is appropriate in substance, though the paper should clearly delineate which statements are new. However, the application section has a serious internal inconsistency: the power analysis in Theorem 4.2 is conducted under a block-diagonal alternative that corresponds to independence in the Section 2 model, not the stated dependence alternative H1. In addition, the proofs of Theorem 2.7 and Theorem 4.3 are not supplied, and the core CLT lemma is only available in a supplement with a placeholder DOI. These gaps prevent the paper from being accepted in its current form; they are fixable within the manuscript's scope, so I recommend major revision.","major_comments":[{"comment":"The power analysis does not address the stated alternative H1: x_N is not independent of x_S. In the Section 2 model, x = Σ^{1/2} \\tilde x with \\tilde x having independent entries, so if Σ = diag(Σ_{N+L}, Σ_{S−L}) as assumed in the proof of Theorem 4.2, then x_N and x_S are independent, i.e., the proof is analyzing a variant of H0, not H1. The argument that some eα_j^{[N+L]} > eα_j shows only that the test can detect an increase in the number of large eigenvalues in the first block, which is a change in the marginal covariance of x_N, not a cross-covariance between x_N and x_S. The power simulation in Section 4.2 sets pairwise covariance 0.5 between variables in x_N and x_S, which is an off-diagonal perturbation not covered by Theorem 4.2. Therefore the claimed consistency of T_N and \\hat T_N against dependence is unsupported; the theorem should either be proved for a genuine cross-covariance alternative or the application should be reframed.","section":"Section 4.1, Theorem 4.2"},{"comment":"The proof of Theorem 2.7 is omitted: Section 5.1 states that the proof is similar to that of Theorem 2.5 and therefore omitted. This is load-bearing because Theorem 2.7 provides the joint CLT for the N spiked eigenvalues that underlies the chi-square statistic in Theorem 4.1. Similarly, Theorem 4.3, the plug-in version of the test statistic used in the simulations, is stated without proof. The claimed χ²(N) limit for \\hat T_N involves estimation of eα_i, ev_i, ψ, K_4, and the covariance matrix; the delta-method assertion needs a careful derivation, especially because the estimation errors enter at the √n scale. Please provide complete proofs or detailed, verifiable derivations for both results.","section":"Section 5.1 and Theorem 4.3"},{"comment":"Lemma 6.1 is the main ingredient for the CLT in Theorem 2.5, but its proof is not in the manuscript; it is relegated to the supplementary material [9], whose DOI is given as 'to be provided by the typesetter'. For a submitted manuscript, this is not a minor presentation issue: the core lemma's proof must be available in a permanent supplement or in the paper itself. Please include a complete proof or ensure that the supplement is accessible.","section":"Section 6, Lemma 6.1"}],"minor_comments":[{"comment":"Remark 2.3 says that examples are given below Theorem 2.9, but the examples actually appear in Remark 2.4, before Theorem 2.5; the cross-reference should be corrected.","section":"Remark 2.3"},{"comment":"The covariance formula for C(λ) is written using E(y_{j1} y_{s1} y_{ℓ1} y_{t1}), while the proof displays E(d_{j1}d_{s1}d_{ℓ1}d_{t1} x_{j1}x_{s1}x_{ℓ1}x_{t1}); the equivalence relies on independence of the mask D and the underlying X and should be stated explicitly.","section":"Proof of Theorem 2.5"},{"comment":"The simulated values of T_2 deviate from the theoretical values by up to about 0.03 in the Gamma and t scenarios (e.g., 0.8529 vs. 0.8825 in column (2)); the paper should comment on whether this reflects finite-sample bias or a missing term in the approximation.","section":"Table 1"},{"comment":"The simulation generates the missing probabilities as random draws (1−θ_i ∼ Uniform(0,0.5)), whereas Section 2 treats θ_i as fixed constants; please clarify whether θ_i are fixed per replication or resampled, and how this affects the reported empirical quantities.","section":"Section 3"},{"comment":"The second part of Proposition 2.1 uses the symbol α̃_k, but this quantity is never defined; it should be eα_k or a clearly defined alternative.","section":"Proposition 2.1"}],"recommendation":"major_revision","confidential_remarks":"The main theorems rely heavily on Li, Pan, Yin and Zhou (2024), which shares two co-authors with the present paper. I do not regard this as circular because [13] is published and the target results are distinct, but the paper should state more explicitly which results are new relative to [13]. The most serious issue is the mismatch between the stated H1 in the independence application and the block-diagonal alternative analyzed in Theorem 4.2; this is an internal inconsistency rather than a mere limitation of the model class, and it should be resolved before the paper is reconsidered."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. The paper contains the first CLTs for spiked eigenvalues and spiked eigenvectors when the sample covariance is formed from entrywise Bernoulli-masked data (MCAR, zero-filling), and the limiting formulas in Theorems 2.5 and 2.9 are genuinely new. The second thing: the main theorems are plausible and the proofs, while sketchy in places, follow a standard route, but the application section has a real gap—Theorem 4.3 is stated without proof, and the power analysis in Theorem 4.2 is narrower than the stated alternative hypothesis.\n\nThe new content is the effective covariance tilde Sigma = P Sigma P + (P - P^2) o Sigma, the explicit covariance kernels tau_0 and nu_0, the eigenvalue CLT with its Gaussian random matrix limit, and the eigenvector CLT with Gamma_k. The derivations follow the Bai-Yao template plus the Li et al. bulk machinery, but the covariance calculations for the masked entries require care with fourth moments of the masks, and that is done properly. The simulation results are consistent with the theorems; the Q-Q plots and the projection table look like genuine checks rather than decoration.\n\nSoft spots, in order of seriousness. (1) Theorem 4.3, the plug-in version of the test statistic, has no proof at all. It is not a one-line delta method: estimating psi(alpha_i), ev_i, and the covariance matrix Sigma_{1N} creates additional variability that needs to be tracked, and the claim that T-hat_N is asymptotically chi-square(N) is load-bearing for the application. (2) Theorem 2.7's proof is omitted as \"similar.\" For the application, the joint distribution of the eigenvalue vector requires the cross-covariance between C(psi(alpha_j)) and C(psi(alpha_i)), which is not in Theorem 2.5's single-block statement. The omitted proof is not just bookkeeping. (3) The power analysis in Theorem 4.2 tests a specific alternative: the first N+L coordinates form a spiked block, so the first L entries of x_S are fused with x_N. That is a legitimate within-model dependence, but it is not the broad \"x_N is not independent of x_S\" stated in Section 4.1. I checked the stress-test claim that block-diagonality makes x_N and x_S independent—that is wrong, because the split between x_N and x_S does not align with the blocks. Still, the empirical power simulation uses pairwise cross-covariances of 0.5 between the L variables in x_S and x_N, which is not literally covered by the theorem. The application section should either expand the theorem or soften the power claim. (4) The MCAR Bernoulli masking with zero-filling is a narrow mechanism; the authors are upfront about it, so I do not count it as a flaw, but readers should not generalize to MNAR or imputation pipelines.\n\nRecommendation: send it to a serious referee. The core eigen-structure theorems are a legitimate contribution to the spiked-model-with-missingness literature and deserve verification. The referee should ask for a proof or detailed sketch of Theorem 4.3 and a rewritten statement of Theorem 4.2's alternative. With those fixed, I would be comfortable seeing it published in a solid journal. If you work on missing-data PCA or random matrix theory, cite it.","headline":"First CLTs for spiked eigenvalues/eigenvectors under Bernoulli missingness; core theory plausible, but the application section has an unproved plug-in theorem and a narrower-than-claimed power analysis.","tokens_in":22916,"tokens_out":5480,"would_cite":true,"duration_ms":52013,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62H15","62B20","62D10"],"pacs":[],"model":"deepseek-v4-flash","headline":"Under entrywise missing observations, distant spiked sample eigenvalues and eigenvectors still have Gaussian limits, with shift and covariance formulas that differ from the complete-data case.","keywords":["central limit theorem","high-dimensional statistics","missing data","sample covariance matrix","spiked eigenvalues","spiked eigenvectors","independence test"],"falsifier":"Simulate a one-spike model with n = p = 500 under the paper's assumptions, record the empirical variance of \\sqrt{n}(\\lambda_{n,1} - \\psi(\\tilde{\\$\\alpha$}_1)) over many replications, and compare it with the variance from Theorem 2.5; then repeat the same simulation but make the Bernoulli mask depend on |x_{ij}| so that missingness is not completely at random. A systematic mismatch in the second setting, while the first matches, would confirm that the missing-completely-at-random, zero-imputation premise is the load-bearing condition.","tokens_in":21923,"feed_emoji":"📊","tokens_out":6596,"duration_ms":61853,"temperature":0.7,"pith_summary":"This paper asks what happens to principal-component analysis when a high-dimensional dataset has entries missing in a random coordinate-wise fashion, and it models missingness by multiplying each observation by an independent Bernoulli mask with zeros in unobserved entries. Its central claim is that in the spiked population model the distant sample spiked eigenvalues still converge almost surely, and after centering by a modified spike map they obey a Gaussian central limit theorem, while the normalized spiked eigenvectors are asymptotically Gaussian around the eigenvectors of a modified population covariance. The limiting parameters are not the complete-data ones: missingness changes the effective covariance, the centering map, and the variance. If correct, standard eigenvalue-based inference, including confidence intervals and tests of independent block structure, can be extended to incomplete high-dimensional data using the formulas supplied. The paper also constructs a chi-squared independence test for two blocks of variables and shows it has asymptotic power against the modeled dependence alternative.","feed_headline":"Missing data reshapes the Gaussian limits of spiked PCA","feed_subtitle":"New formulas give exact eigenvalue and eigenvector fluctuations, enabling inference on incomplete high-dimensional data.","key_machinery":"The machinery is the effective covariance matrix \\tilde{\\Sigma} = P\\Sigma P + (P - $P^{2}$) \\circ \\Sigma, which converts the Bernoulli-masked problem into a no-missing problem with a modified population, together with the spike-to-eigenvalue map \\psi(\\tilde{\\$\\alpha$}) = \\tilde{\\$\\alpha$}(1 + c \\int t/(\\tilde{\\$\\alpha$} - t) H(dt)) whose derivative sign separates distant spikes from closed ones. The Gaussian limit is carried by an N × N symmetric random matrix C(\\psi(\\tilde{\\$\\alpha$}_k)) whose covariance, specified in Theorem 2.5, is expressed through \\tau_0, \\nu_0 and mixed fourth moments such as E(y_{j1} y_{s1} y_{\\ell 1} y_{t1}). The proofs condition on the resolvent of the bulk block, extract the deterministic part [1 + c f_1(\\$\\lambda$)] \\tilde{\\Sigma}, apply a central limit theorem for random quadratic forms, and use an eigenvector perturbation resolvent bound; the application then turns the joint normal vector from Theorem 2.7 into a chi-square statistic.","core_discovery":"Under a block-diagonal spiked population, the sample covariance matrix built from y = b ◦ x behaves spectrally as if the population covariance were \\tilde{\\Sigma} = P \\Sigma P + (P - $P^{2}$) \\circ \\Sigma, where P = diag(\\theta_i) and \\circ is the Hadamard product. For a distant spike \\tilde{\\$\\alpha$}_k with \\psi'(\\tilde{\\$\\alpha$}_k) > 0, where \\psi(\\tilde{\\$\\alpha$}) = \\tilde{\\$\\alpha$}(1 + c \\int t/(\\tilde{\\$\\alpha$} - t) H(dt)), Proposition 2.1 gives \\lambda_{n,j} \\to \\psi(\\tilde{\\$\\alpha$}_k) almost surely. Theorem 2.5 states that \\sqrt{n}(\\lambda_{n,j} - \\psi(\\tilde{\\$\\alpha$}_k)) converges in distribution to the eigenvalues of an explicit Gaussian random matrix whose covariance involves fourth-order moments of the observed entries, the effective covariance \\tilde{\\Sigma}, and the spectral functionals \\tau_0 and \\nu_0; Theorem 2.9 gives \\sqrt{n}(a_k - \\tilde{e}v_k) \\Rightarrow N(0, \\Gamma_k) for the normalized eigenvectors. The paper's key message is that normality survives missingness, but the shift and the spread are governed by the missing rates and the effective spectrum, not by the original spectrum.","pith_inferences":["Because the proof reduces the masked problem to an effective covariance, the same Gaussian-limit structure should survive under more general independent-coordinate masks, such as deterministic missing proportions or heterogeneous Bernoulli rates, with only the effective covariance and fourth-moment terms changed; the paper does not pursue this generalization.","The variance formulas suggest a testable extension: comparing the empirical variance of \\sqrt{n}(\\lambda_{n,1} - \\psi(\\tilde{\\alpha}_1)) across different missing rates could be used to estimate the missingness probabilities, although inversion is not developed here.","For practitioners, the paper implies that applying classical PCA after zero-filling and using complete-data CLT formulas will misstate uncertainty whenever missingness is non-negligible; the corrected formulas provide the benchmark that simpler approaches would need to match.","The framework deliberately excludes missing-not-at-random mechanisms and non-zero imputation; adapting the limiting formulas to those settings is a natural next step, and the covariance formulas here give a concrete reference point for such extensions."],"forward_implications":["Confidence intervals and hypothesis tests for spiked eigenvalues can be built from the new centering \\psi(\\tilde{\\alpha}_k) and the covariance in Theorem 2.5 instead of the complete-data formulas.","The independence test statistic T_N is asymptotically \\chi^2(N) under the null and has power tending to 1 against the modeled dependence alternative, giving a practical tool for block independence in incomplete data.","Eigenvector-based inference, such as subspace estimation, must use the effective eigenvectors \\tilde{e}v_k and the covariance \\Gamma_k from Theorem 2.9; the plain sample eigenvector is biased by a factor governed by 1/(1 + c f_2(\\psi(\\tilde{\\alpha}_k)) \\tilde{\\alpha}_k).","Uniform missingness does not simply rescale the spectrum: it can shift a spike below the phase-transition boundary and turn a spiked sample eigenvalue into a bulk eigenvalue.","The delta-method estimator \\hat{T}_N with plug-in estimates of \\psi and \\tilde{\\alpha} retains the chi-square limit, allowing the test to be applied without knowing the population parameters."],"supporting_citations":[{"why":"Supplies the spectral analysis of Gram matrices with missing observations, including the almost sure convergence of the bulk empirical spectral distribution and the resolvent central limit theorem that the proofs invoke.","marker":"[13]"},{"why":"Provides the generalized spiked population model and the distant-spike classification through \\psi'(\\tilde{\\alpha}) > 0 that the paper extends to missing data.","marker":"[4]"},{"why":"Gives the central limit theorem for spiked eigenvalues in the complete-data case and the perturbative determinant method that Theorem 2.5 adapts.","marker":"[3]"},{"why":"Supplies the eigenvector perturbation bound used to derive the eigenvector central limit theorem in Theorem 2.9.","marker":"[12]"},{"why":"Supplies the estimation formulas for \\psi' and f_2 used in the plug-in statistic of the paper's independence test application.","marker":"[17]"}],"fun_headline_variants":["Missing data warps the Gaussian limits of spiked PCA","Missing observations change spiked PCA's asymptotic covariance","Under missing data, spiked PCA's Gaussian limits depend on missing rates","Missing data twists the Gaussian limits of spiked PCA"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole analysis rests on the assumption that each entry is missing completely at random and independently of the data, with missing entries recorded as zero; if missingness is correlated with the value that would have been observed, or with other coordinates, or if a different imputation is used, the stated limits no longer apply.","fun_headline_variants_meta":{"raw":{"variants":["Missing data warps the Gaussian limits of spiked PCA","Missing observations change spiked PCA's asymptotic covariance","Under missing data, spiked PCA's Gaussian limits depend on missing rates","Missing data twists the Gaussian limits of spiked PCA"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000479,"raw_usage":{"total_tokens":2371,"prompt_tokens":941,"completion_tokens":1430,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":557,"completion_tokens_details":{"reasoning_tokens":1363}},"tokens_in":557,"tokens_out":1430,"duration_ms":11561,"temperature":1.0,"reasoning_tokens":1363,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T22:48:03.134778+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate a one-spike model with n = p = 500 under the paper's assumptions, record the empirical variance of \\sqrt{n}(\\lambda_{n,1} - \\psi(\\tilde{\\$\\alpha$}_1)) over many replications, and compare it with the variance from Theorem 2.5; then repeat the same simulation but make the Bernoulli mask depend on |x_{ij}| so that missingness is not completely at random. A systematic mismatch in the second setting, while the first matches, would confirm that the missing-completely-at-random, zero-imputation premise is the load-bearing condition.","supporting_citations":[{"cited_title":"and ZHOU, W","cited_arxiv_id":null,"evidence_quote":"Supplies the spectral analysis of Gram matrices with missing observations, including the almost sure convergence of the bulk empirical spectral distribution and the resolvent central limit theorem that the proofs invoke."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the generalized spiked population model and the distant-spike classification through \\psi'(\\tilde{\\alpha}) > 0 that the paper extends to missing data."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Gives the central limit theorem for spiked eigenvalues in the complete-data case and the perturbative determinant method that Theorem 2.5 adapts."},{"cited_title":"and ZHONG, P","cited_arxiv_id":null,"evidence_quote":"Supplies the estimation formulas for \\psi' and f_2 used in the plug-in statistic of the paper's independence test application."}],"review_version":1}