REVIEW 3 major objections 5 minor 1 cited by
Large dimensional Spearman's rank correlation matrices: The central limit theorem and its applications
T0 review · 3 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read The paper proves a central limit theorem for linear spectral statistics of large Spearman rank correlation matrices, covering analytic functions and yielding the first CLT for Hoeffding's improved Spearman matrix.
desk verdict A technically strong CLT paper that generalizes Bao et al. (2015) to analytic LSS and adds the first order-3 U-statistic CLT, but the main theorem depends on two unproven lemmas imported from an unpublished Bao preprint. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the Gram matrix $g_n = \frac{1}{p}\sum_j s_j s_j^{\top}$ of the standardized rank vectors, whose nonzero eigenvalues coincide with those of the Spearman matrix scaled by $p/n$. Because ranking makes the columns dependent, the paper treats them as a sample from a population with covariance $\Sigma = \frac{n}{n-1}(I_n - \frac{1}{n}\mathbf{1}\mathbf{1}^{\top})$ and uses the Stieltjes transform together with a martingale decomposition. The load-bearing identity is the covariance formula for quadratic forms in one column, $\operatorname{cov}(s^{\top}As, s^{\top}Bs)=2\operatorname{tr}(AB) - \frac{6}{5}\operatorname{tr}(A\circ B) - \frac{4}{5n}\operatorname{tr}(A)\operatorname{tr}(B)+O(\|A\|\|B\|)$, whose three leading terms produce the three corrections in the Gaussian mean and covariance. A strong edge-rigidity estimate, imported from an unpublished source, is used to keep the extreme eigenvalues inside the integration contour.
What would settle it
Simulate $n\times p$ doubly independent absolutely continuous data with $p/n$ fixed and examine whether $P(\lambda_1(\rho_n)>\eta_r)$ and $P(\lambda_{\min\{n,p\}}(\rho_n)\le\eta_l)$ decay faster than any power of $n$ for $\eta$ outside the Marchenko-Pastur support; alternatively, verify Proposition 2.3 of Bao (2019b) on which Lemma 6.1 is based. A data-generating mechanism for which the edge tails are only polynomially small would show that Lemma 6.1, and hence the proof of Theorems 3.1–3.3, does not hold as stated.
Extended reading notes
Core claim
The paper's central claim is Theorem 3.2: under doubly independent, absolutely continuous entries and $p/n\to y\in(0,\infty)$, the process $\{T(f)\}$ over analytic $f$ converges weakly to a Gaussian process with mean $EZ_f = -\frac{1}{2\pi i}\oint f(z)\frac{\mu(z/y)}{y}\,dz$ and covariance $\operatorname{cov}(Z_f,Z_g) = -\frac{1}{4\pi^2}\oint\!\oint f(z_1)g(z_2)\frac{\sigma(z_1/y,z_2/y)}{y^2}\,dz_1 dz_2$. The mean and covariance are built from the Stieltjes transform $s(z)$ of the limiting Marchenko-Pastur law and from three covariance corrections specific to rank vectors: the standard $2\operatorname{tr}(AB)$ term, a Hadamard-product term $-\frac{6}{5}\operatorname{tr}(A\circ B)$, and a new trace-product term $-\frac{4}{5n}\operatorname{tr}(A)\operatorname{tr}(B)$. Theorem 3.3 shows the same covariance, with an extra mean term, for Hoeffding's improved Spearman matrix, obtained by controlling the difference between the classical and improved rank matrices.
Load-bearing premise
The proof requires a very strong concentration bound for the largest and smallest eigenvalues of the Spearman matrix (tails of order $o(n^{-m})$ for every $m$), which is cited to an unpublished preprint and not proved in the paper; if that bound fails, the Gaussian-process convergence is not established.
Editorial extensions
If this is right
- Theorem 4.1 gives explicit CLTs for $\log|\rho_n|$ and $\operatorname{tr}(\rho_n^k)$ and for their improved-Spearman analogues; for example, $\log|\rho_n| + (n-p)\log(1-y_n)+p$ is asymptotically normal with mean $\frac{3}{2}\log(1-y)+2y$ and variance $-2\log(1-y)-2y$.
- The improved Spearman CLT is the first CLT for linear spectral statistics of a standard U-statistic of order 3 in random matrix theory.
- The resulting tests $L_{\rho,2}$, $L_{\rho,\log}$, $L_{\tilde{\rho},2}$, and $L_{\tilde{\rho},\log}$ have empirical sizes close to the nominal 5% level under normal, Cauchy, and mixed distributions, whereas Pearson-correlation tests fail completely under heavy tails.
- The polynomial results reproduce the earlier results of Bao et al. (2015), confirming that the Stieltjes-transform route and the moment and cumulant route agree where both apply.
Reading between the lines
- The covariance identity with coefficients $6/5$ and $4/5$ suggests an exchangeable-rank universality class: any rank-based Gram matrix built from vectors uniform on permutations may have LSS CLTs with the same Gaussian covariance, differing only in mean shifts; Kendall's matrix, being a U-statistic of order 2, is a natural test case for this pattern.
- Remark 3.4 suggests replacing the ratio $p/n$ by $p/(n-1)$ in the centering; a direct simulation comparing empirical sizes of the proposed tests under both centering choices would show whether the extra mean term $\mu_3$ is practically removable, a testable extension.
- If the imported edge-rigidity bound were replaced by a proven edge bound under weaker moment assumptions, the CLT could extend to heavier-tailed or discrete data; numerical checks of the tail probability $P(\lambda_1(\rho_n)>\eta_r)$ could indicate how strong an assumption is really needed.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper establishes a central limit theorem for linear spectral statistics (LSS) of Spearman's rank correlation matrices in the proportional-growth regime p/n -> y. The proof represents the rank vectors as columns with a permutation-uniform distribution, forms the associated Gram matrix, and applies a Bai--Silverstein martingale decomposition to the Stieltjes transform. The resulting Gaussian-process limit for analytic test functions generalizes the polynomial LSS result of Bao et al. (2015). The same framework is applied to Hoeffding's improved Spearman matrix, a U-statistic of order 3; the paper shows that replacing the classical Spearman matrix by the improved version changes only the asymptotic mean. These results are used to construct four tests of independence, with simulations for normal, Cauchy, and mixed distributions.
Significance. If the theorems are correct, the paper fills a natural gap: it upgrades polynomial LSS CLTs for Spearman matrices to analytic LSS, and it provides what appears to be the first CLT for LSS of a matrix that is a U-statistic of order 3. The covariance computation is nontrivial because the rank vectors are not independent and have a singular covariance; the three-term covariance formula in Lemma A.1 is explicit, and the polynomial limit is checked against the known result of Bao et al. (2015). The proposed tests are clearly motivated and the simulations support their usefulness under heavy tails. The main reservations are verifiability: two quantitative estimates that carry the proof are imported from an unpublished preprint, and one further approximation lemma is delegated to prior papers by reference only.
major comments (3)
- [Section 6.1, Lemma 6.1 and Remark 6.1] The edge-rigidity bound (19) is load-bearing: it is used to justify the contour-integral representation (18), the truncation error estimate (20), and the uniform moment bound (32) on which the tightness of M_n(z) depends. However, Lemma 6.1 is not proved in the manuscript; Remark 6.1 asserts it follows directly from Proposition 2.3 of Bao (2019b), which is listed as an unpublished preprint. Consequently Theorems 3.1--3.3 are conditional on a result that is neither proved here nor available in a published refereed form. I ask the authors to provide a complete proof of Lemma 6.1 in the appendix, or to replace the reference by a published version, and to state explicitly whether Proposition 2.3 of Bao (2019b) requires assumptions beyond 'doubly independent and absolutely continuous entries' (for example, moment or smoothness conditions). The theorem statements should be amended if the imported result has extra hypotheses.
- [Appendix, Lemma A.2] The concentration inequality (55), used at multiple points in Steps 1 and 3 of Lemma 6.2 (e.g., to prove (23) and to justify the limits (27)--(30)), is imported from Proposition 2.1 of the same unpublished Bao (2019b). Since the bound is needed in the exact form stated, with exponent n^{-q/2+delta} and operator norms, a proof or a precise published citation is required. Without this bound, the martingale-difference verification of finite-dimensional convergence and the explicit covariance formula in Theorem 3.1 do not close. This is a separate load-bearing input from Lemma 6.1.
- [Section 6.3, Lemma 6.4] The proof of Lemma 6.4 consists of a one-sentence reference to Wu and Wang (2022) and Li et al. (2023), without theorem or equation numbers, and those papers do not appear to state the Frobenius-norm bounds (50)--(51) in this form. Lemma 6.4 is essential for replacing ~rho_n and K_n by U_n and V_n in the derivation of Lemma 6.3; the bound E ||~rho_n - U_n||_F^2 = o(p) is nontrivial and should not be delegated without precise pointers. Please include a complete proof or state exactly where in the cited papers these estimates are proved.
minor comments (5)
- [Section 3.2, Remark 3.2] The sentence 'they utilized ... and proposed a two-step comparison approach' contains a duplicated 'for for any positive integer k'; please correct the typo.
- [Section 3.1, after equation (6)] The notation for the Stieltjes transform of rho_n/y_n is not visually distinguished from that of g_n in the displayed equations; please introduce separate symbols or a clear subscript so the two transforms can be told apart.
- [Section 3.3, Theorem 3.3] The condition 'as y_n -> y' should be stated as 'as p/n -> y' for consistency with Theorem 3.2, since y_n = p/n in the paper's notation.
- [Section 6.1, Lemma 6.1] The notation lambda_min{n,p}(rho_n) should be defined explicitly: if interpreted as the smallest eigenvalue of rho_n, the statement is false for p > n because rho_n then has p-n zero eigenvalues; presumably the index min(n,p) in the descending order (i.e., the smallest nonzero eigenvalue when p > n) is intended. Please clarify.
- [Section 4 and Section 5] There are several typographical errors, including 'matirx', 'neighborhoog', 'rouine', and 'concerntration'; these should be corrected in a revision.
Circularity Check
No significant circularity: the derivation is self-contained apart from external probabilistic inputs that are not fitted to the target result.
full rationale
The paper's derivation chain is not circular under the stated criteria. Theorem 3.1 is proven from first principles for the Gram matrix of a centered rank vector: the asymptotic mean and covariance are computed explicitly from exact covariance identities for quadratic forms (Lemma A.1) and concentration bounds (Lemma A.2), followed by a martingale CLT and tightness argument. The covariance formula 2tr(AB) - (6/5)tr(A∘B) - (4/(5n))tr(A)tr(B) is derived combinatorially from the uniform-permutation distribution of the ranks, not fitted to any target statistic. Theorem 3.2 is a change of variable from the Gram matrix to Spearman's matrix, and the polynomial consistency check against Bao et al. (2015) is a check, not an input. Theorem 3.3 is obtained by explicitly evaluating the difference between the improved Spearman matrix and the classical Spearman matrix via Hoeffding's decomposition; the additional mean term is computed from that difference, not chosen to match simulations. The test statistics in Section 4 are derived from the CLTs with explicit centering constants. The only load-bearing external inputs are Lemma A.2, which rests on Proposition 2.1 of Bao (2019b), and Lemma 6.1, which rests on Proposition 2.3 of Bao (2019b). These are external citations to a preprint by a different author, not self-citations, and they are not fitted to or equivalent to the target results; whether that preprint is fully verified is a correctness or verifiability concern, not circularity. No fitted parameter is renamed as a prediction, no ansatz is smuggled in through self-citation, and no known result is merely relabeled. The central claims therefore do not reduce to their own inputs by construction.
Assumptions & free parameters
assumptions (6)
- domain assumption Entries of X are doubly independent and absolutely continuous with respect to Lebesgue measure.
- standard math The LSD of Spearman's correlation matrix is the Marchenko-Pastur law F_y and the corresponding Stieltjes transform equation (6).
- domain assumption Edge rigidity for extreme eigenvalues of Spearman's matrix as in Lemma 6.1.
- domain assumption Hoeffding decomposition A_ij = A_i + A_j + epsilon_ij with uncorrelated remainders.
- standard math Martingale CLT (Theorem 35.12 of Billingsley) applies to the conditional-expectation martingale differences.
- domain assumption Uniform boundedness of the spectral norm of Kendall's matrix K_n almost surely.
Cite this review
Pith. "Pith review of Large dimensional Spearman's rank correlation matrices: The central limit theorem and its applications." pith.science (2026). https://pith.science/paper/V5Y4HKQE
@misc{pith2026241115861,
author = {Pith},
title = {Pith review of: Large dimensional Spearman's rank correlation matrices: The central limit theorem and its applications},
year = {2026},
howpublished = {\url{https://pith.science/paper/V5Y4HKQE}},
note = {Machine review of arXiv:2411.15861}
}
read the original abstract
This paper is concerned with Spearman's correlation matrices under large dimensional regime, in which the data dimension diverges to infinity proportionally with the sample size. We establish the central limit theorem for the linear spectral statistics of Spearman's correlation matrices, which extends the results of [\emph{Ann. Statist.} 43(2015) 2588--2623]. We also study the improved Spearman's correlation matrices [\emph{Ann. Math. Statist} 19(1948) 293--325] which is a standard U-statistic of order 3. As applications, we propose three new test statistics for large dimensional independent test and numerical studies demonstrate the applicability of our proposed methods.
Forward citations
Cited by 1 Pith paper
-
Spectra of high-dimensional Spearman correlation matrices under scale-mixture dependence
Under a shared latent scale, the high-dimensional Spearman matrix has a generalized Marchenko–Pastur limit determined by the law of the conditional rank-score variance, not the classical MP law.
Reference graph
Works this paper leans on
-
[1]
G. W. Anderson and O. Zeitouni. A CLT for regularized sample covariance matrices. Annals of Statistics, 36 0 (6): 0 2553--2576, 2008
work page 2008
- [2]
- [3]
- [4]
- [5]
-
[6]
Z. Bai, D. Jiang, J. Yao, and S. Zheng. Corrections to lrt on large dimensional covariance matrix by rmt. Annals of Statistics, 37 0 (6B): 0 3822--3840, 2009
work page 2009
-
[7]
A. S. Bandeira, A. Lodhia, and P. Rigollet. Mar c enko- P astur law for K endall’s tau. Electronic Communications in Probability, 22: 0 32, 2017
work page 2017
-
[8]
Z. Bao. Tracy--widom limit for K endall’s tau. Annals of Statistics, 47 0 (6): 0 3504--3532, 2019 a
work page 2019
Show all 45 references
-
[9]
Z. Bao. Tracy--widom limit for S pearman’s rho. preprint, 2019 b
2019
-
[10]
Z. Bao, G. Pan, and W. Zhou. Tracy- W idom law for the extreme eigenvalues of sample correlation matrices. Electronic Journal of Probability, 17 0 (88): 0 1--32, 2012
2012
-
[11]
Bao, L.-C
Z. Bao, L.-C. Lin, G. Pan, and W. Zhou. Spectral statistics of large dimensional S pearman’s rank correlation matrix and its application. Annals of Statistics, 43 0 (6): 0 2588--2623, 2015
2015
-
[12]
Billingsley
P. Billingsley. Convergence of probability measures. John Wiley & Sons, 2013
2013
-
[13]
Billingsley
P. Billingsley. Probability and measure. John Wiley & Sons, 2017
2017
-
[14]
S. X. Chen, L.-X. Zhang, and P.-S. Zhong. Tests for high-dimensional covariance matrices. Journal of the American Statistical Association, 105 0 (490): 0 810--819, 2010
2010
-
[15]
Dobriban and S
E. Dobriban and S. Wager. High-dimensional asymptotics of prediction: Ridge regression and classification. Annals of Statistics, 46 0 (1): 0 247--279, 2018
2018
-
[16]
El Karoui
N. El Karoui. Concentration of measure and spectra of random matrices: Applications to correlation matrices, elliptical distributions and beyond. The Annals of Applied Probability, 19 0 (6): 0 2362--2405, 2009
2009
-
[17]
J. Gao, X. Han, G. Pan, and Y. Yang. High dimensional correlation matrices: The central limit theorem and its applications. Journal of the Royal Statistical Society, Series B, 79 0 (3): 0 677--693, 2017
2017
-
[18]
F. Han, S. Chen, and H. Liu. Distribution-free tests of independence in high dimensions. Biometrika, 104 0 (4): 0 813--828, 2017
2017
-
[19]
Hastie, A
T. Hastie, A. Montanari, S. Rosset, and R. J. Tibshirani. Surprises in high-dimensional ridgeless least squares interpolation. Annals of Statistics, 50 0 (2): 0 949, 2022
2022
-
[20]
Heiny and N
J. Heiny and N. Parolya. Log determinant of large correlation matrices under infinite fourth moment . Annales de l'Institut Henri Poincaré, Probabilités et Statistiques, 60 0 (2): 0 1048 -- 1076, 2024
2024
-
[21]
Heiny and J
J. Heiny and J. Yao. Limiting distributions for eigenvalues of sample correlation matrices from heavy-tailed populations. Annals of Statistics, 50 0 (6): 0 3249--3280, 2022
2022
-
[22]
Hoeffding
W. Hoeffding. A class of statistics with asymptotically normal distribution. Annals of Mathematical Statistics, 19 0 (3): 0 293 -- 325, 1948
1948
-
[23]
T. Jiang. The asymptotic distributions of the largest entries of sample correlation matrices. Annals of Applied Probability, 14 0 (2): 0 865--880, 2004 a
2004
-
[24]
T. Jiang. The limiting distributions of eigenvalues of sample correlation matrices. Sankhy \=a : The Indian Journal of Statistics , pages 35--48, 2004 b
2004
-
[25]
T. Jiang. Determinant of sample correlation matrix with application. Annals of Applied Probability, 29 0 (3): 0 1356--1397, 2019
2019
-
[26]
Leung and M
D. Leung and M. Drton. Testing independence in high dimensions with sums of rank correlations. Annals of Statistics, 46 0 (1): 0 280--307, 2018
2018
-
[27]
Z. Li, Q. Wang, and R. Li. Central limit theorem for linear spectral statistics of large dimensional K endall’s rank correlation matrices and its applications. Annals of Statistics, 49 0 (3): 0 1569--1593, 2021
2021
-
[28]
Z. Li, C. Wang, and Q. Wang. On eigenvalues of a high-dimensional K endall’s rank correlation matrix with dependence. Science China Mathematics, 66 0 (11): 0 2615--2640, 2023
2023
-
[29]
Lytova and L
A. Lytova and L. Pastur. Central limit theorem for linear eigenvalue statistics of random matrices with independent entries. Annals of Probability, 37 0 (5): 0 1778--1840, 2009
2009
-
[30]
Mar c enko and L
V. Mar c enko and L. Pastur. Distribution of eigenvalues for some sets of random matrices. Sbornik: Mathematics, 1 0 (4): 0 457--483, 1967
1967
-
[31]
Mestre and P
X. Mestre and P. Vallet. Correlation tests and linear spectral statistics of the sample correlation matrix. IEEE Transactions on Information Theory, 63 0 (7): 0 4585--4618, 2017
2017
-
[32]
G. Pan. Comparison between two types of large sample covariance matrices. Annales de l'IHP Probabilit \'e s et statistiques , 50 0 (2): 0 655--677, 2014
2014
-
[33]
Pan and W
G. Pan and W. Zhou. Central limit theorem for signal-to-interference ratio of reduced rank linear receiver. Annals of Applied Probability, 18 0 (3): 0 1232--1270, 2008
2008
-
[34]
Parolya, J
N. Parolya, J. Heiny, and D. Kurowicka. Logarithmic law of large random correlation matrices. Bernoulli, 30 0 (1): 0 346--370, 2024
2024
-
[35]
Paul and A
D. Paul and A. Aue. Random matrix theory in statistics: A review. Journal of Statistical Planning and Inference, 150: 0 1--29, 2014
2014
-
[36]
N. S. Pillai and J. Yin. Edge universality of correlation matrices. Annals of Statistics, 40 0 (3): 0 1737--1763, 2012
2012
-
[37]
Wang and B
C. Wang and B. Jiang. On the dimension effect of regularized linear discriminant analysis. Electronic Journal of Statistics, 12 0 (2): 0 2709--2742, 2018
2018
-
[38]
C. Wang, J. Yang, B. Miao, and L. Cao. Identity tests for high dimensional data using rmt. Journal of Multivariate Analysis, 118: 0 128--137, 2013
2013
-
[39]
H. Wang, B. Liu, L. Feng, and Y. Ma. Rank-based max-sum tests for mutual independence of high-dimensional random vectors. Journal of Econometrics, 238 0 (1): 0 105578, 2024
2024
-
[40]
Wang and J
Q. Wang and J. Yao. On the sphericity test with large-dimensional observations. Electronic Journal of Statistics, 7: 0 2164--2192, 2013
2013
-
[41]
Wu and C
Z. Wu and C. Wang. Limiting spectral distribution of large dimensional spearman’s rank correlation matrices. Journal of Multivariate Analysis, 191: 0 105011, 2022
2022
-
[42]
J. Yao, S. Zheng, and Z. Bai. Sample covariance matrices and high-dimensional data analysis. Cambridge University Press, 2015
2015
-
[43]
Zheng, Z
S. Zheng, Z. Bai, and J. Yao. Substitution principle for CLT of linear spectral statistics of high-dimensional sample covariance matrices with applications to hypothesis testing. Annals of Statistics, 43 0 (2): 0 546--591, 2015
2015
-
[44]
Zheng, G
S. Zheng, G. Cheng, J. Guo, and H. Zhu. Test for high dimensional correlation matrices. Annals of Statistics, 47 0 (5): 0 2887, 2019
2019
-
[45]
W. Zhou. Asymptotic distribution of the largest off-diagonal entry of correlation matrices. Transactions of the American Mathematical Society, 359 0 (11): 0 5345--5363, 2007
2007
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.