{"id":"026cadf2-b807-41f8-a064-becfa15a5c8d","arxiv_id":"2412.10753","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Bayesian inverse-Wishart inference for spiked covariance matrices is extended with eigenvalue bias corrections and a BIC-based spike count that is consistent when p exceeds n.","lead":"This paper builds a Bayesian method for high-dimensional principal component analysis, using an inverse-Wishart prior to estimate spiked eigenvalues, eigenvectors, and the number of spikes. It adds bias corrections for the eigenvalues and proves that the corrected posterior converges to the right answers even when there are more variables than observations.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The p>n BIC-for-marginal-likelihood step in Eq. (21) is unquantified; Theorem 4.1 proves consistency only of a BIC-weighted pseudo-posterior, not of the exact Bayesian posterior over K.","rationale":"The reader's weakest-assumption analysis identifies Eq. (21) as the fragile bridge from BIC to the exact posterior, and I concur. The eigenvalue contraction results (Theorems 3.3-3.6) are supported by a substantial proof apparatus and appear internally coherent, so the central quantitative eigenstructure claim is not the most exposed point. The K-selection theorem, by contrast, requires an approximation of the marginal likelihood whose error is never bounded. The paper consistently uses the phrase 'BIC-based posterior approximation' when proving concentration, but then concludes 'the posterior is consistent: pi(K=K0|X_n)->1' without qualification. That slippage is consequential: if the BIC approximation error is not negligible, the uncertainty quantification reported for K in the real-data application is not the posterior uncertainty of the stated Bayesian model. I therefore agree with the reader's CONDITIONAL verdict; the concern does not warrant rejection because the pseudo-posterior consistency result is still a useful model-selection theorem, but it does not license the 'fully Bayesian' interpretation in p>n. Other issues, such as the apparent sign problem in Eq. (14), are secondary for the K-selection headline and do not change this conclusion.","tokens_in":52073,"tokens_out":13187,"duration_ms":133126,"concrete_test":"Derive the Laplace remainder for the profiled likelihood of Section 4.1: write log p(X_n|K) = -BIC_K/2 + R_K and bound R_K under the conditions of Theorem 4.1. A minimal concrete check is to compute the information-determinant contribution: with d_K ~ pK parameters, Laplace's method contributes -(1/2) log det I(theta_hat), which is of order pK log n; verify whether the BIC penalty d_K log n absorbs this term exactly and whether all remaining terms in R_K are o(1). If R_K is not o(1), Theorem 4.1 describes only the pseudo-posterior (21), and exact Bayesian K-posterior consistency is not established.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4 replaces the marginal likelihood p(X_n|K) with exp(-BIC_K/2) in (21), citing Kass and Raftery (1995). That Laplace approximation requires a fixed-dimensional parameter space; here d_K ~ pK grows with n and p can exceed n. Moreover, the maximized profile likelihood in Section 4.1 uses a pooled bulk eigenvalue estimator and is not obtained by integrating over a prior on K and Sigma, so it is not a marginal likelihood in the Bayesian sense. No theorem bounds |log p(X_n|K) + BIC_K/2|, and the paper's own Remark 4.2 cites Bai et al. (2018) showing that classical BIC needs modification in the p>n regime. Consequently, Theorem 4.1 establishes only that the exponentiated-BIC weights in (21) concentrate on K0; the later assertion pi(K=K0|X_n)->1 for the exact Bayesian posterior is unsupported. Since the advertised contributions include honest uncertainty quantification for K and fully Bayesian model selection in p>n, this missing approximation bound is the load-bearing weakness of the K-selection claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies Bayesian inference in the spiked covariance model under high-dimensional settings, placing an inverse-Wishart prior on the covariance matrix to obtain posterior distributions for spiked eigenvalues and eigenvectors. It proposes two bias-correction methods for the posterior eigenvalues—a prior degree-of-freedom calibration rule and a post-hoc multiplicative correction—and derives asymptotic contraction rates for the posterior eigenstructure. For the number of spikes K, it develops a BIC-type approximation to the marginal likelihood and proves consistency of the resulting posterior over K in the p>n regime. The theoretical results are supplemented with simulation studies and an S&P 500 real-data application, with claims of computational efficiency due to independent posterior sampling.","tokens_in":52326,"tokens_out":7086,"duration_ms":62909,"significance":"If the results are correct, the paper provides a useful Bayesian complement to frequentist high-dimensional PCA, with posterior contraction rates for spiked eigenvalues and eigenvectors, a minimax-optimal eigenvector rate in the single-spike case, and a new consistency result for selecting K when p>n. The computational scheme—independent inverse-Wishart draws—is a practical strength, and the detailed proofs and simulation evidence are valuable. However, the paper's central claims rest on several technical points that currently appear unsound or unverified, particularly the sign of the posterior expectation in Eq. (14), the applicability of Theorem 3.4 to the proposed prior calibration rule, and the validity of the BIC approximation as a marginal likelihood in the p>n regime.","major_comments":[{"comment":"Eq. (14) asserts that the posterior expectation of ||[Omega21]_k||^2 equals (n - nu_n - 2p - 2) * hat_lambda_k * (sum_{l=K+1}^p hat_lambda_l) / ((n + nu_n - 2p - 1)(n + nu_n - 2p - 4)). For the hyperparameter values used throughout the simulations, nu_n = 2p + 2, the numerator is n - 4p - 4, which is negative when p > n/4, while the denominator is positive for n > 4. Since ||[Omega21]_k||^2 is a nonnegative random variable, Eq. (14) cannot be correct as stated. This is not a peripheral typo: the derivation of the bias-correction factor gamma_tilde_1 in Eq. (8) explicitly replaces ||[Omega21]_k||^2 in gamma_1 by this posterior expectation, and the text then approximates (n - nu_n - 2p - 2)/(n - nu_n - 2p - 1) ~ 1, an approximation that is not meaningful when the numerator and denominator have opposite signs. The formula likely needs a sign change (possibly n + nu_n - 2p - 2), but as written the theoretical justification of both bias-correction strategies in the p > n regime is invalid.","section":"Section 3.2, Eq. (14)"},{"comment":"Theorem 3.4 assumes nu_n - 2p = o(n) for the inverse-Wishart prior and then claims that the prior calibration rule (6) satisfies this condition. This is not established and appears false in the high-dimensional settings targeted by the paper. For example, with n = 100 and p = 500, Eq. (6) gives nu_n approximately O(1) while 2p = 1000, so nu_n - 2p = Omega(p), which is not o(n). Consequently, the posterior contraction claim (17) is not proven for the prior calibration method when p > n. Additionally, the denominator in (6) is a random quantity that can be negative or near zero for some data configurations; the paper provides no condition ensuring that the rule is well-defined. The proof of Theorem 3.4 needs to either verify the o(n) condition for (6) under the assumptions of the theorem or state a separate set of assumptions under which prior calibration achieves the claimed rate.","section":"Section 3.3, Theorem 3.4"},{"comment":"The paper replaces the marginal likelihood p(X_n | K) with exp(-BIC_K/2), citing Kass and Raftery (1995). This Laplace-type approximation is not justified in the p > n regime studied here: the number of free parameters d_K ~ pK grows with p and n, the maximized log-likelihood in Section 4.1 is a plug-in profile likelihood rather than an integrated marginal likelihood, and no theorem bounds the approximation error |log p(X_n|K) + BIC_K/2|. Remark 4.2 itself acknowledges that classical BIC requires modification for p > n, citing Bai et al. (2018). As a result, Theorem 4.1 and the subsequent display pi(K=K0|X_n) -> 1 establish consistency of the BIC-weighted pseudo-posterior in Eq. (21), not of the exact Bayesian posterior pi(K|X_n). Since the paper advertises fully Bayesian model selection and honest uncertainty quantification for K in high dimensions, this missing approximation bound is load-bearing and needs to be addressed, either by proving the approximation under the stated spike conditions or by weakening the claims to those for a BIC-based criterion.","section":"Section 4, Eq. (21)"}],"minor_comments":[{"comment":"The theorem defines epsilon_n^2 = K^3/n + (lambda_{0,K+1}/lambda_{0,k})(p/n vee 1), but then states the contraction bound with epsilon_n. To avoid ambiguity, the statement should explicitly define epsilon_n as the square root of the displayed quantity.","section":"Theorem 3.4 statement"},{"comment":"The notation for the correction term c-hat is used as c-hat p in Eq. (6), while it is defined earlier as c-hat with a denominator p - K - pK/n; the notation should be harmonized to avoid confusion.","section":"Eq. (6)"},{"comment":"In the displayed derivation for the post-hoc correction, the symbol 'nu_2' appears in the denominator (e.g., in the bound on gamma_2/gamma_tilde_1 - 1) where 'nu_n' is intended; please correct this typo.","section":"Appendix C, proof of Theorem 3.4"},{"comment":"For Setting 1 with n=100 and p=1000, the proposed method has ACC=0.88 and A VG=2.89, showing noticeable underestimation of K in the most challenging scenario; the paper mentions this but does not discuss whether the underestimation is due to the BIC approximation or the prior on K, which would be informative for users.","section":"Section 5.2, Table 3"},{"comment":"The parameter count d_K = pK - K(K+1)/2 + K + 1 includes a term -K(K+1)/2 for the orthogonality constraints of the eigenvector matrix; it would be helpful to state explicitly that this uses the standard dimension of the Grassmann manifold, since not all readers will immediately see the counting.","section":"Section 4.1"}],"recommendation":"major_revision","confidential_remarks":"The negative posterior expectation in Eq. (14) is a red flag that suggests a sign error in the derivation; if it is a typo, the correction must be carried through the entire bias-correction derivation, and the claim that nu_n from (6) satisfies nu_n - 2p = o(n) needs to be verified or replaced with appropriate conditions. The BIC-for-marginal-likelihood issue is more than a technicality: the paper's headline contribution of 'fully Bayesian model selection in p>n' is currently not supported by the theorems. I would advise the editor that these issues are fixable in principle, but they require substantial revision rather than copy-editing."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should look at this paper if you care about Bayesian PCA in high dimensions. The core eigenvalue and eigenvector results are worth taking seriously, but the dimension-selection claim is over-sold.\n\nWhat's actually new: they give a clean perturbation framework for inverse-Wishart posterior eigenvalues, propose two bias-correction strategies (prior calibration and post-hoc multiplication), and prove posterior contraction rates for spiked eigenvalues and eigenvectors, including minimax optimality in the single-spike case. Those theorems look substantive and the proof structure is careful. The simulations are consistent with the theory, and the computational advantage of independent posterior samples is real.\n\nThe soft spots are in two places. First, the BIC step in Section 4. The paper replaces the marginal likelihood with exp(-BIC/2) in (21), cites Kass and Raftery, and then proves that the BIC-based weights concentrate on K0. But the number of parameters d_K grows with pK, p can exceed n, and the maximized profile likelihood is not a prior-integrated marginal likelihood. No theorem bounds the approximation error. So Theorem 4.1 establishes consistency of a BIC-weighted pseudo-posterior, not the exact Bayesian posterior. The title and abstract claim \"fully Bayesian model selection\" and uncertainty quantification for K, and that is not supported. At minimum the claims need to be reframed or the approximation error needs a theorem.\n\nSecond, Eq. (14) gives the posterior expectation of ||[Ω21]_k||^2 as proportional to (n - ν_n - 2p - 2). For the hyperparameters used in the simulations (ν_n = 2p + 2, n = 100, p = 500), this is n - 4p - 4 < 0, which would make a squared norm negative in expectation. That cannot be right as written. It may be a typo—the γ̃1 formula in (8) has the opposite sign—but the paper should be corrected. The prior calibration rule is explicitly designed to match the Wang-Fan debiased estimator; that is a defensible frequentist-assisted calibration, but the text should acknowledge it rather than present it as purely Bayesian.\n\nThe central eigenvalue/eigenvector contraction is plausible and the paper deserves a serious referee. The fixes are clear: correct Eq. (14), and either prove a BIC approximation bound under their asymptotics or rewrite the K-selection claims as consistency of the BIC posterior. I would engage with it and push for revision.","headline":"Solid Bayesian spiked-covariance results with a load-bearing gap: the BIC-based posterior for K is a pseudo-posterior, and Eq. (14) contains a sign error that makes a posterior expectation negative in the simulation settings.","tokens_in":52825,"tokens_out":4228,"would_cite":true,"duration_ms":39071,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62H25","62F15","62H12"],"pacs":[],"model":"deepseek-v4-flash","headline":"A Bayesian inverse-Wishart posterior, corrected for high-dimensional eigenvalue inflation, recovers spiked eigenvalues and the number of spikes consistently even when p exceeds n.","keywords":["spiked covariance model","Bayesian principal component analysis","inverse-Wishart prior","eigenvalue bias correction","posterior contraction","BIC model selection","high-dimensional covariance estimation","number of spikes"],"falsifier":"Simulate $n=100$, $p=1000$, $K_0=3$ under the paper's first design, compute the exact marginal likelihood for $K=2,3,4$ by numerical integration over the inverse-Wishart posterior, and compare the argmax with the maximizer of $\\exp(-\\mathrm{BIC}_K/2)$. A systematic ranking difference at a frequency that does not vanish would show that Theorem 4.1 establishes consistency for a pseudo-posterior, not the exact Bayesian posterior.","tokens_in":51886,"feed_emoji":"📊","tokens_out":9034,"duration_ms":75046,"temperature":0.7,"pith_summary":"The paper sets out to show that a fully Bayesian treatment of the spiked covariance model can estimate spiked eigenvalues, eigenvectors, and the unknown number of spikes with honest uncertainty even when the dimension $p$ exceeds the sample size $n$. The route is an inverse-Wishart prior on the covariance matrix, whose conjugacy makes posterior sampling cheap, followed by two bias corrections that undo the high-dimensional inflation of posterior eigenvalues. The paper proves that the corrected posterior contracts around the true spiked eigenvalues at rate $\\sqrt{K^3/n + (\\lambda_{0,K+1}/\\lambda_{0,k})(p/n \\vee 1)}$ and that the BIC-based posterior for the number of spikes concentrates on the true $K_0$, with $\\pi(K=K_0 \\mid X_n) \\to 1$, in the $p>n$ regime. A sympathetic reader would care because this supplies a practical, checkable Bayesian alternative to frequentist PCA dimension selection, one that yields credible intervals rather than only point estimates.","feed_headline":"Bayesian correction recovers PCA spikes when p > n","feed_subtitle":"Inverse-Wishart posterior plus two corrections gives consistent eigenvalues and spike counts with honest uncertainty.","key_machinery":"The engine is the conjugate inverse-Wishart posterior, which remains full-rank when $p>n$ and yields independent posterior draws. The argument is carried by an eigenvalue perturbation expansion (Theorem 3.1) that writes the $k$-th eigenvalue of a random positive definite matrix as the $k$-th eigenvalue of a $K\\times K$ principal submatrix times a factor driven by the squared norm of the off-diagonal block; this factor is what inflates the posterior eigenvalues, and formulas (6) and (7) are constructed to cancel it. For the spike count, the machinery is a BIC pseudo-posterior $\\pi(K\\mid X_n)\\propto \\exp(-\\mathrm{BIC}_K/2)\\pi(K)$, where $\\mathrm{BIC}_K$ uses the maximized plug-in likelihood under a flat-bulk approximation of the $p-K$ trailing eigenvalues.","core_discovery":"On its own terms, the paper establishes that if the observations are Gaussian with a spiked covariance matrix and the prior is inverse-Wishart, the conjugate posterior $\\Sigma \\mid X_n \\sim \\mathrm{IW}_p(A_n + nS_n, \\nu_n + n)$ is a valid inferential engine for spiked eigenstructure once its eigenvalue bias is understood. The raw posterior inflates the $k$-th leading eigenvalue by a multiplicative factor that the paper identifies through an eigenvalue perturbation theorem; the prior-calibration rule (6) and the post-hoc rule (7) cancel this factor. Under the paper's conditions, the corrected posterior contracts at rate $\\epsilon_n = \\sqrt{K^3/n + (\\lambda_{0,K+1}/\\lambda_{0,k})(p/n \\vee 1)}$ for eigenvalues and at rate $B_k K/n + \\lambda_{0,K+1}/(n\\lambda_{0,k})$ for the top $k$ eigenvectors, with the single-spike eigenvector rate attaining the minimax lower bound. For the number of spikes, the paper approximates the marginal likelihood by $\\exp(-\\mathrm{BIC}_K/2)$ and proves posterior consistency, $\\pi(K=K_0 \\mid X_n) \\to 1$, under a prior supported on $K=o(n)$, including when $p>n$.","pith_inferences":["Editorial inference: the post-hoc correction factor in (7) should transfer to any inverse-Wishart posterior with a different scale matrix, so it could serve as a general debiasing template for full-rank Bayesian covariance posteriors; the paper does not explore that transfer.","Editorial inference: the consistency theorem treats the BIC pseudo-posterior, not the exact marginal-likelihood posterior; replacing $\\exp(-\\mathrm{BIC}_K/2)$ with a proper Laplace approximation could change the finite-sample ranking of $K$ values, and the paper does not bound that difference.","Editorial inference: the rise in posterior entropy of $K$ around financial crises, visible in the real-data example, suggests the posterior spread over spike counts could be formalized as a real-time structural-break detector.","Editorial inference: a direct test of the BIC approximation would compute the exact marginal likelihood for $K=2,3,4$ at moderate $p,n$ by numerical integration over the inverse-Wishart posterior and compare its mode with the $\\exp(-\\mathrm{BIC}_K/2)$ maximizer."],"forward_implications":["If Theorem 3.4 holds, posterior credible intervals for spiked eigenvalues keep nominal coverage in high dimensions whenever $\\lambda_{0,k}$ grows fast enough relative to $p/n$, a regime where naive inverse-Wishart intervals fail.","If Theorem 4.1 holds, the BIC posterior supplies a consistency guarantee for a Bayesian PCA dimension-selection method in the $p>n$ regime, with a full posterior over the number of components rather than a single chosen number.","Because the posterior draws are independent, the method bypasses MCMC convergence diagnostics and parallelizes across cores, a practical corollary of the conjugacy that the paper emphasizes.","In the single-spike model, the posterior eigenvector contraction rate matches the minimax lower bound, so no estimator can do uniformly better in that benchmark case."],"supporting_citations":[{"why":"Supplies the sample-eigenvalue debiasing relation $\\lambda_k(S_n)\\approx \\lambda_{0,k} + \\hat{c}p/n$ that the post-hoc correction factor $\\gamma_2$ in (9) is built on.","marker":"Wang and Fan (2017)"},{"why":"Provides the spiked-eigenvalue limiting laws and edge conditions (their Assumptions 2, 7, 8) that the paper adapts as Assumptions 1–3 for the posterior contraction analysis.","marker":"Cai et al. (2020)"},{"why":"Defines the single-spiked model and the consistency threshold for sample eigenvectors that the paper's minimax comparison targets.","marker":"Johnstone and Lu (2009)"},{"why":"Justifies the BIC approximation $\\pi(K\\mid X_n)\\propto\\exp(-\\mathrm{BIC}_K/2)\\pi(K)$ used for the posterior over the number of spikes.","marker":"Kass and Raftery (1995)"},{"why":"Supplies the concentration of the averaged trailing sample eigenvalues (their Lemma 7) used in the BIC gap and in the $\\hat{c}$ estimate.","marker":"Yata and Aoshima (2012)"},{"why":"Supplies the Wishart spectral-norm tail bounds used to control deviations of the scaled inverse-Wishart posterior from the identity.","marker":"Lee and Lee (2018)"},{"why":"Establishes the earlier $p/n<1$ consistency threshold for classical BIC in high-dimensional PCA, which the paper's $p>n$ result extends (Remark 4.2).","marker":"Bai et al. (2018)"}],"fun_headline_variants":["Bayesian bias fix recovers PCA spikes when p > n","Spike count and eigenstructure from corrected Bayes","Unbiased Bayesian spiked PCA in high dimensions","Bayes corrects eigenvalue bias, proves spike consistency","High-dimensional Bayesian spike detection, bias-free"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that replacing the marginal likelihood with a BIC approximation whose penalty grows with the number of variables $p$ does not change the posterior over the number of spikes; if that approximation is poor in the $p>n$ regime, the consistency proof applies to a modified pseudo-posterior, not the exact Bayesian one.","fun_headline_variants_meta":{"raw":{"variants":["Bayesian bias fix recovers PCA spikes when p > n","Spike count and eigenstructure from corrected Bayes","Unbiased Bayesian spiked PCA in high dimensions","Bayes corrects eigenvalue bias, proves spike consistency","High-dimensional Bayesian spike detection, bias-free"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000253,"raw_usage":{"total_tokens":1607,"prompt_tokens":1033,"completion_tokens":574,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":649,"completion_tokens_details":{"reasoning_tokens":500}},"tokens_in":649,"tokens_out":574,"duration_ms":5596,"temperature":1.0,"reasoning_tokens":500,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T15:39:48.098373+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate $n=100$, $p=1000$, $K_0=3$ under the paper's first design, compute the exact marginal likelihood for $K=2,3,4$ by numerical integration over the inverse-Wishart posterior, and compare the argmax with the maximizer of $\\exp(-\\mathrm{BIC}_K/2)$. A systematic ranking difference at a frequency that does not vanish would show that Theorem 4.1 establishes consistency for a pseudo-posterior, not the exact Bayesian posterior.","supporting_citations":[],"review_version":1}