{"id":"bbaff768-01f8-4ff8-b919-3f8f1e603994","arxiv_id":"2501.19280","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"For sparse sample covariance matrices X^T X with nonzero-mean weights, the typical largest eigenvalue and top eigenvector component distribution are obtained from replica-based recursive distributional equations solved by population dynamics.","lead":"This paper uses the replica method to compute the average largest eigenvalue and the density of top eigenvector components for large sparse covariance matrices built from a data matrix with few nonzero entries. The results, solved numerically with a population dynamics algorithm, match direct diagonalization and recover known dense-limit formulas.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The replica-symmetric ansatz in Sec. 4.3 is the load-bearing unproven step: without a replicon stability check, Eqs. (69)-(89) may describe a metastable saddle point rather than the true top eigenpair.","rationale":"The reader's verdict is CONDITIONAL, and this stress-test supports that status. The replica algebra from the RS ansatz to Eq. (69) is internally consistent, and the dense-limit recovery (Sec. 7) together with the diagonalisation comparisons (Figs. 2-4) are genuine positive evidence. However, all of those checks use the same RS equations, so they cannot certify the ansatz itself. The most damaging scenario is not a minor algebraic slip but a wrong saddle-point manifold: if cross-replica terms are relevant, the action (37) is missing the true extremum, and both lambda and T(u) could be incorrect in regimes not yet tested. The paper's own admission in Sec. 4.3 that the ansatz is not the most general possible, and its remark that the rotationally invariant ansatz fails for top eigenpair statistics, make this the clearest load-bearing gap. Section 6's population-dynamics stability criterion is also asserted rather than proved, but it is secondary: it is a proposed way to solve the RS equations, whereas the RS ansatz determines which equations are being solved. A replicon calculation is the standard concrete test for local RS stability; it is heavy but well defined. Until such a check, or an equivalent 1RSB computation, is performed, the paper should remain conditional rather than being accepted as a rigorous derivation. Hence the verdict is UNCHANGED relative to the reader's CONDITIONAL assessment.","tokens_in":25282,"tokens_out":16946,"duration_ms":182458,"concrete_test":"Perform a replicon stability analysis of the RS saddle point. Linearize the replicated action (37) around the solution of (69) with respect to a perturbation of the order parameters phi and psi that breaks replica symmetry, for example an off-diagonal Gaussian component, and compute the corresponding Hessian eigenvalue in the n -> 0 limit. If the replicon eigenvalue is negative for the parameters of Figs. 3-4, or for a sweep including signed p(K) and q near the isolation threshold, the RS ansatz is locally unstable and Eqs. (69)-(89) cannot be trusted. If it is positive in those regimes, the central claim is supported against this specific instability.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4.3 restricts the saddle point to a superposition of Gaussians with no cross-replica terms (Eqs. 31-35), explicitly noting that this is \"not the most general possible\". All central results—the fixed-point system (69), the identification <lambda_1> = lambda (Eq. 68), and the eigenvector density (89)—follow from this ansatz. No stability analysis with respect to replica-symmetry-breaking perturbations is provided. For sparse-matrix problems, the top eigenpair is precisely the regime where the more restrictive rotationally invariant ansatz fails, as the paper itself states in Sec. 4.3, so the exclusion of cross-terms is not a harmless technicality. If an RSB saddle point has lower action, the population-dynamics solution of (69) is the stationary point of the wrong functional, and the two numerical examples in Figs. 3-4, while encouraging, do not rule out failures in untested regimes such as small q near the detachment transition or signed weight distributions. This is the most load-bearing unproven assumption in the paper.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops a replica-method calculation of the statistics of the largest eigenvalue and the associated eigenvector for diluted Wishart matrices J = X^T X, where X is an N x M sparse matrix with independent nonzero weights drawn from p(K) and with bounded row/column degrees R and C. The average largest eigenvalue is expressed through a Lagrange parameter lambda that is identified from the stability threshold of a population-dynamics algorithm, and the full system of recursive distributional equations is given in Eq. (69). The density of the top eigenvector components is then obtained as Eq. (89). The authors validate the approach against direct numerical diagonalisation for two weight distributions (delta and uniform) over a range of sparsity parameters q, and they show that a suitable dense limit recovers the known noncentral Wishart result, with an additional conjecture on the fluctuation scale of the eigenvector components.","tokens_in":25541,"tokens_out":4748,"duration_ms":50802,"significance":"If the derivation is correct, this is a valuable contribution to the sparse random matrix literature, where analytical results for extreme eigenpair statistics are scarce. The paper provides explicit, numerically solvable equations, a self-contained algorithm, and nontrivial benchmarks: the recovery of the noncentral Wishart result in the dense limit, the agreement with diagonalisation in Figs. 3-4 for two qualitatively different weight distributions, and the scaling study in Fig. 2 that quantifies finite-size corrections. The conjecture in Eqs. (120)-(121) is clearly flagged as unproven and is a useful stimulus for further work. The main limitation is that the central results rest on the permutation-symmetric, cross-term-free replica ansatz, whose stability is not analysed.","major_comments":[{"comment":"The replica-symmetric ansatz restricts the saddle point to superpositions of Gaussians with no cross-replica terms, and the paper itself notes that this is 'not the most general possible'. All subsequent central results—the fixed-point system (69), the identification <lambda_1> = lambda (Eq. (68)), and the eigenvector density (89)—follow from this ansatz, yet no stability analysis with respect to replica-symmetry-breaking perturbations is provided. Because the top-eigenpair problem is precisely the regime where the more restrictive rotationally invariant ansatz is known to fail (as stated in Sec. 4.3), the absence of a replicon check is a load-bearing gap. The two numerical examples are encouraging but do not cover small q near the detachment transition or signed weight distributions, where an RSB saddle point could plausibly have lower action. I recommend that the authors either perform a local stability analysis around the permutation-symmetric saddle point or provide a clear argument (or numerical evidence in previously untested regimes) that cross-replica terms cannot affect the extremal action.","section":"Sec. 6, stability criterion after Eq. (69)"},{"comment":"The population dynamics algorithm terminates by asserting that the h and mu populations diverge for lambda < <lambda_1> and shrink to zero for lambda > <lambda_1>, with the critical lambda giving the desired eigenvalue (Eq. (68)). This stability criterion is essential: it is what converts the underdetermined system (69) into a predictive value of <lambda_1>. However, the criterion is not proved in the paper but deferred to [22]. Since the present manuscript aims to be self-contained and the eigenvalue prediction depends entirely on this criterion, I ask for a more explicit derivation or, at minimum, a systematic numerical demonstration of the claimed divergence/vanish dichotomy beyond the examples in Fig. 1.","section":"Sec. 4.3 and discussion after Eq. (69)"},{"comment":"The replacement of the microcanonical bounded-degree model by the canonical Bernoulli model with a truncated Poisson degree distribution is invoked as a shortcut and justified by reference to Appendix B of [22]. The truncation is then used to enforce the row/column bounds R and C, which are essential for the O(1) behaviour of <lambda_1> (Appendix A). Since the equivalence of the truncated and microcanonical models is a non-trivial technical step and is load-bearing for the claim that the equations describe the bounded-degree ensemble, the paper should state precisely which parts of the argument are carried over from [22] and which are assumed, or include a self-contained justification in an appendix.","section":"Sec. 5, derivation of Eq. (89)"}],"minor_comments":[{"comment":"In the Introduction, 'No table examples include...' should read 'Notable examples include...'.","section":"Sec. 1"},{"comment":"In the displayed system (69), the third equation uses a sum with an upper limit that is not explicitly indicated as the truncation bound R, although the text explains that p_{alpha q}(s) is the truncated Poisson distribution; please make the notation uniform in the displayed equations.","section":"Eq. (69)"},{"comment":"The initialisation of lambda to a 'large' value uses the upper bound from Appendix A, but the algorithm's step (ix) decreases lambda by Delta; it would be useful to state explicitly how the target error tolerance Delta relates to the final uncertainty quoted as ±Delta/2.","section":"Sec. 6, algorithm step (i)"},{"comment":"The sentence 'This result indeed follows directly from evaluating \\bar{\\mu}' is terse; the connection between the normalisation condition and the value of \\bar{\\mu} would benefit from one additional equation or a short explanation.","section":"Sec. 7, after Eq. (117)"},{"comment":"The paper states that its analysis does not account for finite-size corrections, but Fig. 2 shows that such corrections are ~4% at the smallest size; a brief comment on the expected scaling of these corrections with N and M would be informative.","section":"Sec. 8"},{"comment":"In Eq. (B.4), the approximation sign is used when replacing the product by an exponential; for clarity, the condition q << sqrt(NM) and the fact that the correction is exponentially small in the large-N,M limit could be stated explicitly.","section":"Appendix B"}],"recommendation":"major_revision","confidential_remarks":"The paper is a technically dense extension of the authors' previous line of work, and it relies on several non-rigorous steps (replica ansatz, population-dynamics stability criterion, Poisson truncation shortcut) that are partly deferred to refs. [22-24]. The numerical validation in the tested regimes is convincing, and the dense-limit reduction is a strong consistency check. The main obstacle is the missing stability analysis of the replica-symmetric saddle point; this is a standard concern in sparse-matrix replica calculations and, given the paper's own admission that the ansatz is not the most general, it should be addressed before publication. I do not believe rejection is warranted, but a revision that adds a replicon check or at least a detailed discussion and additional tests near the detachment transition is necessary."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is the first analytic handle on the top eigenpair of sparse covariance matrices of the form X^T X with nonzero-mean weights, and it looks right. Equations (69) and (89) are genuinely new as far as I can tell—the density of top eigenvector components (89) in particular. The derivation extends the Susca-Vivo-Kuhn replica framework from sparse adjacency matrices to the bipartite noncentral case, which is a real step forward. The paper earns its keep with two checks: the dense limit recovers the noncentral Wishart result <λ1>=<K~>^2 with the correct 1/q scaling, and the population-dynamics numerics match direct diagonalisation very well for both δK,1 and uniform weights. I also like the honest conjecture about Gaussian fluctuations in the dense double-scaling limit, with supporting numerics.\n\nThe soft spots are the usual ones for this style of replica calculation, and they are not fatal. The RS ansatz in Sec. 4.3 is explicitly not the most general possible, and there is no replicon-style stability check; the population-dynamics stability criterion in Sec. 6 is asserted rather than derived. Since the top eigenpair is exactly where rotationally invariant ansätze fail for sparse matrices, the exclusion of cross-replica terms is a real assumption, not a harmless technicality. The two numerical examples are encouraging but do not probe small-q near the detachment transition or signed weight distributions, so the boundary of validity is unknown. Also, key derivations are deferred to earlier papers (Appendix E of [24], Appendix B of [22]), and no code or data are provided, which makes reproduction harder than it should be. These are limitations, not fatal flaws; the dense-limit and numerics give me reasonable confidence that the central result is right.\n\nWho should read this: people working on spectral theory of sparse random matrices, and anyone using sparse PCA or covariance estimation in the high-dimensional regime. It deserves a serious referee. I would send it out, with requests for a replicon check or at least more comprehensive numerics, plus release of the population-dynamics code. The core result is worth publishing even if the rigor is not complete.","headline":"A legitimate first for top eigenpairs of diluted Wishart matrices, likely correct; the replica-symmetric ansatz is the main unproven step but doesn't sink it.","tokens_in":26040,"tokens_out":2950,"would_cite":true,"duration_ms":27868,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["60B20","15B52","82B44"],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper shows that the average largest eigenvalue of a diluted Wishart matrix equals the critical parameter at which the population dynamics reaches stability, and that this parameter plus the stable populations yield the density of the…","keywords":["diluted Wishart matrices","replica method","population dynamics","largest eigenvalue statistics","eigenvector component density","sparse random matrices","noncentral Wishart ensemble","spectral gap"],"falsifier":"Run direct numerical diagonalisation for a bounded-support, nonzero-mean weight distribution not considered in the paper, such as a two-point distribution $p(K)=\\frac12(\\delta_{K,a}+\\delta_{K,b})$, and compare the measured top eigenvalue and top eigenvector component histogram against the predictions of Eqs. (69) and (89); a systematic discrepancy beyond finite-size fluctuations would refute the central claim.","tokens_in":25091,"feed_emoji":"📊","tokens_out":9973,"duration_ms":84135,"temperature":0.7,"pith_summary":"This paper claims that the statistics of the top eigenpair of a diluted Wishart matrix $\\mathbf{J} = \\mathbf{X}^T \\mathbf{X}$ become tractable when the data matrix $\\mathbf{X}$ is sparse with bounded row and column degrees and its nonzero weights have a nonzero mean that produces an isolated largest eigenvalue. Working in the limit of large $N,M$ with fixed ratio, the authors reformulate the largest eigenvalue as a zero-temperature optimisation on the sphere and evaluate the disorder average with the replica method. The central claim is that the average largest eigenvalue $\\langle\\lambda_1\\rangle$ equals the critical Lagrange parameter $\\lambda$ that stabilises a population-dynamics algorithm solving a closed system of recursive distributional equations, and that the stable populations then give the density of the top eigenvector's components through a separate integral formula. If correct, this supplies analytical control over the top eigenpair of sparse covariance matrices, a regime where existing results were mostly confined to the average spectral density. The paper verifies the formulas against direct diagonalisation for two weight distributions and shows that the dense limit recovers the known noncentral Wishart answer.","feed_headline":"Sparse Wishart top eigenpair equals a stability threshold","feed_subtitle":"Population dynamics pinpoints the top eigenpair of sparse covariance matrices, matching diagonalisation.","key_machinery":"The engine of the calculation is a replica-symmetric ansatz that represents the replicated order parameters as superpositions of an uncountable set of Gaussians with nonzero mean, parametrised by two probability densities $\\pi(\\omega,h)$ and $\\rho(\\sigma,\\mu)$ (and their conjugates). A Hubbard-Stratonovich transformation turns the disorder-averaged replicated partition function into a functional integral, and a saddle-point evaluation in the limits $\\beta\\to\\infty$ and $n\\to0$ produces the closed system of recursive distributional equations (69), together with an integral constraint that fixes the Lagrange parameter $\\lambda$. These equations are solved iteratively by a population-dynamics algorithm, in which $\\lambda$ is tuned until the first moments of the $h$ and $\\mu$ populations neither explode nor vanish; that critical $\\lambda$ is the average largest eigenvalue, and the converged populations feed the formula (89) for the top eigenvector component density.","core_discovery":"The paper's central claim is that for a diluted Wishart matrix $\\mathbf{J} = \\mathbf{X}^T \\mathbf{X}$ with bounded row/column degrees and a nonzero-mean weight distribution $p(K)$ that generates a spectral gap, the typical largest eigenvalue is $\\langle\\lambda_1\\rangle = \\lambda$, where $\\lambda$ is the unique parameter for which the recursive distributional equations (69) admit a nontrivial stable solution: the auxiliary populations $h$ and $\\mu$ diverge when $\\lambda$ is below the true top eigenvalue and shrink to zero when it is above. Once $\\lambda$ and the stable densities $\\pi(\\omega,h)$ and $\\rho(\\sigma,\\mu)$ are obtained, the density of the top eigenvector's components is given by the integral formula (89). The paper confirms these predictions numerically for $p(K)=\\delta_{K,1}$ and a uniform $K\\in(0,1)$, and in the dense limit recovers the noncentral Wishart result $\\langle\\lambda_1\\rangle=\\langle\\tilde{K}\\rangle^2$ together with full localisation of the top eigenvector.","pith_inferences":["If the replica-symmetric equations hold, the same framework could be extended to track the overlap between top eigenvectors of two coupled sparse covariance matrices by adding an external field to the replicated action, which would open a route to principal-component retrieval problems in the sparse regime.","The paper's stability criterion suggests a practical diagnostic that the top eigenvalue can be located by monitoring the first moment of the $h$-population alone, which may be usable as a cheap estimator in settings where full diagonalisation is infeasible.","The paper leaves the non-gapped regime open; a natural test is to add a zero-mean weight component and check where the formulas begin to fail, which would mark the sparse analogue of the BBP transition.","The dense-limit conjecture of Gaussian fluctuations, Eq. (120), is directly testable at finite $M$; if verified, it would provide a quantitative interpolation between the sparse localised regime and the dense localised regime."],"forward_implications":["The average largest eigenvalue of a diluted Wishart matrix can be obtained from a one-parameter stability search in population dynamics, without diagonalising the matrix.","The same computation yields the full density of top eigenvector components, giving quantitative access to localisation and component statistics in the sparse regime.","Because the equations are valid for any connectivity distribution, the method applies directly to matrices with hard caps on row and column degrees, as used in the numerical checks.","In the dense limit, the formulas reduce to the known noncentral Wishart results: the top eigenvalue tends to $\\langle\\tilde{K}\\rangle^2$ and the top eigenvector components concentrate at $u=1$, with the paper's conjectured Gaussian fluctuations forming a testable bridge between sparse and dense behaviour."],"supporting_citations":[{"why":"Introduces the functional order parameters for the spectral density of $\\mathbf{X}^T\\mathbf{X}$ and the $\\alpha\\leftrightarrow 1/\\alpha$ duality that structure the replica calculation.","marker":"[21]"},{"why":"Supplies the method for the typical largest eigenvalue of sparse weighted graphs, the population-dynamics stability criterion, and the justification for replacing the Poisson degree distribution with any connectivity distribution.","marker":"[22]"},{"why":"Extends the population-dynamics approach to the top eigenpair of sparse graphs, informing the derivation of the eigenvector component density.","marker":"[23]"},{"why":"Provides the technical details of the replica-symmetric transformation, including the appendix used to integrate out the replicated vectors.","marker":"[24]"},{"why":"Establishes the representation of replica order parameters as continuous superpositions of Gaussians, which is the core ansatz of the calculation.","marker":"[39]"},{"why":"Introduces the population dynamics algorithm used to solve the recursive distributional equations numerically.","marker":"[67]"},{"why":"Gives the noncentral Wishart largest eigenvalue formula that the paper recovers in the dense limit.","marker":"[83]"}],"fun_headline_variants":["Top eigenpair of diluted Wishart set by population dynamics","Sparse Wishart top eigenvalue: a self-consistent threshold","Replica method pins down top eigenpair of sparse covariance","Diluted Wishart: top eigenvector density from stable populations","Top eigenpair of sparse Wishart from replica population equations"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the replica-symmetric ansatz (superpositions of Gaussians with no cross-replica terms) together with the population-dynamics stability criterion correctly encode the top eigenpair statistics; if either of these unproven steps fails, the equations (69) and (89) would not describe the true largest eigenvalue and eigenvector.","fun_headline_variants_meta":{"raw":{"variants":["Top eigenpair of diluted Wishart set by population dynamics","Sparse Wishart top eigenvalue: a self-consistent threshold","Replica method pins down top eigenpair of sparse covariance","Diluted Wishart: top eigenvector density from stable populations","Top eigenpair of sparse Wishart from replica population equations"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000476,"raw_usage":{"total_tokens":2352,"prompt_tokens":925,"completion_tokens":1427,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":541,"completion_tokens_details":{"reasoning_tokens":1343}},"tokens_in":541,"tokens_out":1427,"duration_ms":10589,"temperature":1.0,"reasoning_tokens":1343,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T20:41:40.636417+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run direct numerical diagonalisation for a bounded-support, nonzero-mean weight distribution not considered in the paper, such as a two-point distribution $p(K)=\\frac12(\\delta_{K,a}+\\delta_{K,b})$, and compare the measured top eigenvalue and top eigenvector component histogram against the predictions of Eqs. (69) and (89); a systematic discrepancy beyond finite-size fluctuations would refute the central claim.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the functional order parameters for the spectral density of $\\mathbf{X}^T\\mathbf{X}$ and the $\\alpha\\leftrightarrow 1/\\alpha$ duality that structure the replica calculation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the method for the typical largest eigenvalue of sparse weighted graphs, the population-dynamics stability criterion, and the justification for replacing the Poisson degree distribution with any connectivity distribution."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Extends the population-dynamics approach to the top eigenpair of sparse graphs, informing the derivation of the eigenvector component density."},{"cited_title":"Monasson and D","cited_arxiv_id":null,"evidence_quote":"Provides the technical details of the replica-symmetric transformation, including the appendix used to integrate out the replicated vectors."},{"cited_title":"Katzav and I","cited_arxiv_id":null,"evidence_quote":"Establishes the representation of replica order parameters as continuous superpositions of Gaussians, which is the core ansatz of the calculation."},{"cited_title":"Krivelevich and B","cited_arxiv_id":null,"evidence_quote":"Gives the noncentral Wishart largest eigenvalue formula that the paper recovers in the dense limit."}],"review_version":1}