{"id":"93d52ed9-d39f-48a4-93bf-b63f48a96789","arxiv_id":"2607.25486","paper_version":1,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"low","formal_verification":"none","parameter_count":2,"one_line_summary":"Under a shared latent scale, the high-dimensional Spearman matrix has a generalized Marchenko–Pastur limit determined by the law of the conditional rank-score variance, not the classical MP law.","lead":"Spearman rank correlation matrices from data with a common hidden scale do not converge to the usual Marchenko–Pastur spectrum; their eigenvalue law is a generalized Marchenko–Pastur curve set by the distribution of an effective rank variance. The paper proves this and gives explicit formulas and approximations for heavy-tailed data such as Student distributions.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No significant objection identified: symmetry assumption is explicit and generalized by Theorem 4.1; proof of Theorem 2.3 appears complete.","rationale":"The reader's verdict is ACCEPT with moderate confidence. My stress-test concurs. The central theorem is well-supported and the only major caveat—the symmetry assumption—is explicitly stated and generalized in Theorem 4.1. I agree with the reader's identification of symmetry as the weakest assumption, but I do not see it as a correctness risk: the theorem is explicitly conditional, and the broader latent-variable framework covers the asymmetric case. The proof of Theorem 2.3 is complete: conditioning on σ gives independent oracle scores with variance one, the Lindeberg condition is verified using s(σ)>0 a.s., the ECDF approximation is justified by DKW, and the rank normalization is shown to differ from the ECDF matrix by a vanishing perturbation. The numerical fits are clearly disclosed as approximations. A concrete test of the asymmetric case would further validate Theorem 4.1 and confirm the boundary of Theorem 2.3, but this does not undermine the paper's central contribution. Therefore the verdict remains unchanged.","tokens_in":16877,"tokens_out":23385,"duration_ms":210186,"concrete_test":"Run a simulation with N=400, T=800 (q=0.5), σ_t i.i.d. log-normal(0,1), ξ asymmetric (e.g., ξ = Exp(1)-1), compute the ESD of R_N, and compare to the theoretical density from Theorem 4.1 (μ_ψ) and Theorem 2.3 (μ_s). If the empirical spectrum matches Theorem 4.1, the paper's generalization is validated and the symmetry assumption is confirmed necessary.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim, Theorem 2.3, is conditional on ξ being symmetric about 0 (Assumption (ii)). This is the most load-bearing assumption: it makes the standardized rank transform g odd, so E[g(σξ)|σ]=0 and the oracle scores are centered, leading to the LSD equation with μ_s=Law(s(σ)). If ξ is asymmetric, Theorem 2.3's formula with μ_s is not valid; the bulk is instead governed by μ_ψ=Law(Var[g(σξ)]) and an outlier can appear, as stated in Theorem 4.1. The paper is fully transparent about this assumption, and provides the broader latent-variable framework in Section 4. The proof of Theorem 2.3 follows standard lines (conditioning on σ, Lindeberg condition, DKW inequality, rank perturbation) and I found no gap. The numerical section is clearly presented as approximations. Hence the symmetry assumption is a scope limitation, not a flaw, and it does not change the accept recommendation.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies the limiting spectral distribution (LSD) of the sample Spearman rank correlation matrix for high-dimensional scale-mixture data x_t = σ_t ξ_t, where σ_t is a common latent scale and ξ_t has i.i.d. symmetric coordinates. Under the proportional-growth regime N/T→q and s(σ)>0 a.s., Theorem 2.3 proves that the empirical spectral distribution converges a.s. to a generalized Marchenko–Pastur law whose Stieltjes transform solves (3) with μ_s = Law(E[g(σξ)^2|σ]); the classical MP law is recovered only when s(σ) is deterministic. Theorem 4.1 extends this to a latent-variable model H(L,ε) and predicts the bulk from μ_ψ = Law(Var[g(H(L,ε))|L]), with a possible finite-rank outlier when φ(L)=E[g(H(L,ε))|L]≠0. Sections 3–4 give solvable examples (Rademacher scale-mixture, beta-copula ansatz, one-cut approximation, one-factor Gaussian, regime-switching, time-varying tail-thickness) and numerical illustrations.","tokens_in":17163,"tokens_out":33248,"duration_ms":289567,"significance":"If correct, the main result provides a clean and explicit demonstration that rank-based Spearman matrices are not insensitive to a common latent scale even though marginal magnitudes are discarded; this is of interest in robust multivariate statistics and random matrix theory. The proof of Theorem 2.3 is rigorous and follows a clean three-step strategy: oracle scores, DKW, rank perturbation, built on a generalized Marchenko–Pastur lemma. The paper is transparent about the symmetry assumption (it is explicit, and Theorem 4.1 covers the asymmetric case) and labels all numerical fits as approximations. The solvable examples, in particular the Rademacher case (Proposition 3.1) and the hypergeometric R-transform for beta-copulas (Proposition 3.3), are useful benchmarks. The numerical section is honest: the beta and one-cut densities are fitted to simulated effective rank variances, not presented as parameter-free predictions.","major_comments":[],"minor_comments":[{"comment":"Lemma 6.2 cites 'Corollary A.42 of [2]', but [2] is a journal article (Statistica Sinica, 2008) rather than a book with appendices. Please correct the reference (likely Corollary A.42 of [1], the Bai–Silverstein book) or supply the appropriate source.","section":"§6.4 (Lemma 6.2)"},{"comment":"Before defining w_{t,n}=A_{n,t}/√ψ_t, the proof should explicitly restrict to the almost-sure event {ψ_t>0 for all t} (analogous to Ω0 in Proposition 5.1). As written, the definition requires positivity of every ψ_t without stating this event; the later bound w²≤12/ψ_t also depends on it.","section":"§5.2 (Theorem 4.1 proof)"},{"comment":"The fitting procedure for the polynomial coefficients (β_k) and multiplicities (m0,m1) is not described; since Fig. 4 uses these fits, please state how the parameters were selected (e.g., least squares on the simulated density) or label the curves explicitly as illustrative.","section":"§3.2.2 (one-cut approximation)"},{"comment":"The abstract states that coordinates are 'pairwise uncorrelated in both the Pearson and Spearman sense' without mentioning the moment condition; the Pearson part requires finite second moments, as acknowledged in §2.1. Add a parenthetical clarification in the abstract.","section":"Abstract / §2.1"},{"comment":"The claim that 'the fitted parameter appears to grow linearly with ν' is made without a regression line or error bars in the figure. Please either add these or soften the wording to avoid overinterpretation.","section":"§3.2.1, Fig. 3"}],"recommendation":"minor_revision","confidential_remarks":"This is a solid contribution. The main theorem is correct and well proved; the secondary theorem's proof is a somewhat compressed adaptation and the numerical fitting details need small clarifications. The issues are local and do not affect the central claims, so I recommend acceptance after a minor revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague—\n\nThe thing to know: this is a legitimate, well-executed paper that extends the known LSD machinery for rank-based correlation matrices to the scale-mixture case, and it makes the result concrete with solvable examples. It is not a revolution—the proof route is the conditional generalized MP approach from Wu-Wang and Bai-Zhou, plus DKW—but the explicit identification of the effective rank variance s(σ) and the solvable families (binary-spin beta law, beta-copula) are genuinely new and useful.\n\nWhat's good: Theorem 2.3 is proved cleanly. The three-step argument (oracle scores → ECDF scores → exact rank normalization) is standard but the bounds are checked. The symmetry assumption on ξ is stated clearly, and the paper honestly notes that without it Theorem 2.3's formula is not correct; Theorem 4.1 generalizes the bulk to μ_ψ and flags a possible outlier. The binary-spin result—that s(σ)/3 is always Beta(1/2,1) independent of σ's distribution—is a neat, surprising fact. The beta-copula and one-cut examples are honest approximations, labeled as such.\n\nSoft spots, in proportion: (1) Theorem 4.1's proof is a sketch, adapting the same argument; for a paper whose abstract highlights it as a second main theorem, I'd want the details filled in or at least a clear statement that it is a sketched generalization. (2) The numerical agreement in Figs. 1–2 is partly manufactured: the effective rank variance distribution is fitted to simulations, then the spectral curve is predicted from that fit. That's a sensible workflow, and the paper is transparent about it, but it's a weaker demonstration than a parameter-free prediction. (3) The symmetry assumption is load-bearing: it makes the oracle scores conditionally centered. The paper addresses the asymmetric case via Theorem 4.1, so this is a scope limitation, not a hidden flaw. (4) Novelty is incremental: the fixed-point equation is the same generalized MP form from [27] and [2]. But the explicit computation of s(σ) for the scale-mixture model is the contribution, and it's useful.\n\nBottom line: this is a serious, honest paper. A good referee should engage with the proof of Theorem 4.1 and ask for more detail there, and should check the numerical fitting disclosure. It deserves peer review, not a desk reject. I'd cite it for the binary-spin and beta-copula results. Bring it to reading group if you want to see the technique cleanly applied; otherwise a quick skim suffices.","headline":"A solid, honest extension of generalized MP theory to Spearman spectra under scale mixtures—the explicit solvable examples are the real contribution, and the symmetry caveat is properly generalized.","tokens_in":17615,"tokens_out":1813,"would_cite":true,"duration_ms":17412,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["60B20","62H20"],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper proves that, under a scale-mixture model, the eigenvalue spectrum of a Spearman rank correlation matrix converges to a generalized Marchenko-Pastur law governed by the distribution of a latent scale variable, so rank-based correla","keywords":["Spearman correlation matrices","scale-mixture model","limiting spectral distribution","generalized Marchenko-Pastur law","effective rank variance","latent variable dependence","heavy-tailed data","random matrix theory"],"falsifier":"For large N,T with symmetric ξ and non-deterministic σ, compute the empirical eigenvalue density of R_N and compare its second moment with 1 + q + q Var(s(σ)), where s(σ) = E[g(σξ)^2] is estimated from the data; a mismatch beyond finite-size error would contradict Corollary 2.6.","tokens_in":16811,"feed_emoji":"📈","tokens_out":4614,"duration_ms":50550,"temperature":0.7,"pith_summary":"The paper establishes a high-dimensional spectral law for Spearman rank correlation matrices when data come from a scale mixture: each observation is a common scalar scale times an independent symmetric noise vector. It proves that, as dimension and sample size grow at the same rate, the eigenvalue distribution of the sample Spearman matrix converges almost surely to a deterministic generalized Marchenko-Pastur law determined by the distribution of an effective rank variance. The classical Marchenko-Pastur law appears only when the scale variable is constant, so the rank transformation does not remove latent-scale dependence. This gives a concrete handle on heavy-tailed, factor-driven data in which ranking is used for robustness, since the paper also provides solvable examples and an extension to general latent variables.","feed_headline":"Spearman spectra keep the fingerprint of latent scale","feed_subtitle":"A shared volatility or hidden regime changes the bulk spectrum even after ranks discard all magnitudes.","key_machinery":"The central object is the effective rank variance s(σ) = E[g(σξ)^2], where g is the standardized population rank transform; its law μ_s enters a fixed-point equation tying the Stieltjes transform m(z) of the limiting spectrum to z = 1/m(z) + ∫ x/(1 − q x m(z)) dμ_s(x). The proof works by replacing ranks with oracle scores, conditioning on the scale variables, applying a generalized Marchenko-Pastur theorem under a Lindeberg condition, then showing the difference from exact ranks is a vanishing perturbation.","core_discovery":"Under the scale-mixture model x_t = σ_t ξ_t with symmetric i.i.d. coordinates ξ, and in the proportional regime N/T → q, the empirical spectral distribution of the sample Spearman rank correlation matrix R_N converges almost surely to F_{q,μ_s}, a deterministic probability measure whose Stieltjes transform m satisfies z = 1/m(z) + ∫_{[0,3]} x/(1 − q x m(z)) dμ_s(x), where s(σ) = E[g(σξ)^2] is the effective rank variance, g is the standardized population rank transform, and μ_s is the law of s(σ). The classical Marchenko-Pastur law is recovered only when μ_s is a point mass at 1. The paper also proves a broader latent-variable version: if coordinates are conditionally independent given a late","pith_inferences":["A reader could use the second-moment excess q Var(s(σ)) as a nonparametric diagnostic for a common latent scale: estimate the spectrum of a Spearman matrix and compare its variance with the classical Marchenko-Pastur benchmark.","The beta-copula structure appearing in the solvable examples suggests that for many scale-mixture laws the effective rank variance may be approximately beta-distributed; testing this on other heavy-tailed families would map how universal the approximation is.","The extension to asymmetric directional components implies a phase transition: as the conditional mean grows, the bulk law changes and a finite-rank outlier appears; one could simulate asymmetric scale mixtures to locate that transition.","Because the formula depends only on the law of g(σξ)^2, a similar fixed-point equation should hold for other bounded score functions, e.g., normal-score transforms, with the same conditional-variance substitution."],"forward_implications":["The limiting spectrum of the Spearman matrix is computable in closed form whenever the law of s(σ) is known; the paper derives explicit R-transforms for binary-spin and beta-copula models.","If σ is deterministic, the LSD is exactly the classical Marchenko-Pastur law; any non-constant σ increases the spectral variance by q Var(s(σ)).","The same rank-spectrum mechanism extends to any latent-variable dependence where coordinates are conditionally independent given a latent variable, including regime-switching and one-factor Gaussian models.","For latent-variable models with a nonzero conditional mean, the bulk LSD is governed by ψ(L), while the mean component may produce one large outlying eigenvalue without affecting the bulk.","Because Spearman correlations are invariant under increasing transformations, the result covers transelliptical families and heavy-tailed Student data."],"fun_headline_variants":["Rank spectra betray latent scale in heavy-tailed data","Scale mixture leaves a fingerprint on Spearman spectra","Spearman spectra exposed: hidden scale survives rank transform","Effective rank variance shapes rank correlation spectra","Shared volatility imprints on Spearman bulk spectrum"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The directional shocks that get multiplied by the common scale variable are assumed symmetric around zero; if they have a nonzero mean, the simple formula for the bulk spectrum no longer holds without extra terms.","fun_headline_variants_meta":{"raw":{"variants":["Rank spectra betray latent scale in heavy-tailed data","Scale mixture leaves a fingerprint on Spearman spectra","Spearman spectra exposed: hidden scale survives rank transform","Effective rank variance shapes rank correlation spectra","Shared volatility imprints on Spearman bulk spectrum"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000336,"raw_usage":{"total_tokens":1710,"prompt_tokens":767,"completion_tokens":943,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":511,"completion_tokens_details":{"reasoning_tokens":873}},"tokens_in":511,"tokens_out":943,"duration_ms":8240,"temperature":1.0,"reasoning_tokens":873,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T02:16:57.348158+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"For large N,T with symmetric ξ and non-deterministic σ, compute the empirical eigenvalue density of R_N and compare its second moment with 1 + q + q Var(s(σ)), where s(σ) = E[g(σξ)^2] is estimated from the data; a mismatch beyond finite-size error would contradict Corollary 2.6.","supporting_citations":[],"review_version":1}