REVIEW 5 minor 1 cited by
Distribution of singular values in large sample cross-covariance matrices
T0 review · 0 major / 5 minor · reviewed 2026-08-08 · deepseek-v4-flash
Pith's one-line read This paper derives the full singular-value distribution of the sample cross-covariance matrix of two independent Gaussian datasets and shows it is governed by a cubic equation for the Stieltjes transform, with band edges given by simple…
desk verdict Honest, correct extension of product-of-Wishart results to explicit cross-covariance singular-value bounds; worth a serious referee, mainly for the practical edge formulas. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is the S-transform of free probability: for asymptotically free random matrices $\mathbf{A}$ and $\mathbf{B}$, the S-transform of their product is multiplicative, $S_{\mathbf{AB}}(t) = S_{\mathbf{A}}(t)S_{\mathbf{B}}(t)$. The paper writes the squared cross-covariance, up to scaling, as a product of two normalized Wishart matrices, uses the known S-transform of each, and converts the result into a cubic equation for the Stieltjes transform $h(z)$. The eigenvalue density follows from the imaginary part of $h$, and the band edges follow from the zeros of the discriminant of the cubic equation.
What would settle it
Simulate the model with $T=1000$, $N_X=N_Y=2000$ (so $p_X=p_Y=0.5$) but draw the entries from a heavy-tailed distribution such as Student-$t$ with 5 degrees of freedom instead of Gaussians, and compare the empirical singular-value density and band edges with Eqs. (15) and (20); a mismatch would show the result depends essentially on Gaussianity rather than on the free-product structure. A second check: introduce a shared low-rank signal into $\mathbf{X}$ and $\mathbf{Y}$ and see whether the top singular value separates from the predicted $\gamma_+$ as the sample size grows.
Extended reading notes
Core claim
For two independent Gaussian i.i.d. matrices $\tilde{\mathbf{X}}$ and $\tilde{\mathbf{Y}}$ of dimensions $T\times N_X$ and $T\times N_Y$, the normalized empirical cross-covariance matrix $\mathbf{C} = \tfrac{1}{T}\tilde{\mathbf{Y}}^\top \tilde{\mathbf{X}}$ has nonzero singular values whose density follows from the imaginary part of the solution of the cubic equation (Eq. 15) with coefficients given in Eqs. (16)-(19). The band edges are the zeros of the discriminant of this cubic; in simplifying limits they have closed forms: for equal aspect ratios $p_X=p_Y<1$ the edges are given by Eq. (20), in the severely undersampled limit $p_X=p_Y\to 0$ by Eq. (21), when one dataset is much higher dimensional by Eq. (22), when both are severely undersampled by Eq. (23), and in the well-sampled case $T>N_X,N_Y$ by Eq. (25) with the lower edge at zero. In all cases the typical scale of the singular values is $\sqrt{N_XN_Y}/T$, and the result extends the Marchenko-Pastur spectrum of self-covariances to cross-covariances.
Load-bearing premise
Everything rests on the entries of the two data matrices being independent, identically distributed Gaussian draws with known variances, so that the normalized self-covariances are asymptotically free and the S-transform multiplication rule applies; if the data are non-Gaussian, the two datasets are dependent, or the three sizes are not simultaneously large with fixed ratios, the cubic equation and the edge formulas need not hold.
Editorial extensions
If this is right
- With the band edges known, singular-value outliers in a cross-covariance of uncorrelated data can be screened: any singular value above $\gamma_+$ is a candidate for real shared signal.
- The noise floor scales as $\sqrt{N_X N_Y}/T$, so doubling the sample size shrinks spurious singular values by a factor $\sqrt{2}$ relative to the dimensions.
- Signal detection is possible when $T < N_X, N_Y$, a regime where self-covariances cannot be inverted and CCA-style whitening methods fail.
- In the well-sampled regime the cross-covariance band edge is smaller than the whitened cross-correlation edge (same scaling, different prefactor), suggesting that whitening with estimated covariances injects extra noise.
- A strong shared signal that is invisible in each dataset's own covariance spectrum can be detected in the cross-covariance when one dataset is well sampled enough to compensate for the other's poor sampling.
Reading between the lines
- The same cubic-equation machinery should extend to more than two datasets: because the S-transform of a product of $M$ Wishart matrices is known, the cross-covariance band for a chain of $M$ datasets should follow from a higher-degree polynomial, though the paper does not explore this.
- A testable extension is to check robustness to non-Gaussian entries: free-probability universality suggests the cubic equation should hold for a wide class of i.i.d. entries with finite fourth moments, but the paper only proves the Gaussian case, so simulating heavy-tailed entries at the same aspect ratios would test this directly.
- The explicit upper edge $\gamma_+$ could seed a hypothesis test for shared signal: compare the largest observed singular value against $\gamma_+$ computed from estimated aspect ratios, although the distribution of the largest singular value beyond the bulk edge is not derived here.
- The asymmetric band in the $p_X, p_Y \ll 1$ limit might be exploited to estimate the aspect ratio of one dataset from the observed spectrum of the other, given the scaling of the edges.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper considers two Gaussian i.i.d. matrices X (T×N_X) and Y (T×N_Y) and studies the singular-value spectrum of the normalized empirical cross-covariance matrix C=(1/T)\tilde Y^\top \tilde X. Using the S-transform product rule for asymptotically free Wishart matrices, it derives a cubic equation (Eq. 15) for the Stieltjes transform of H=(1/(σ_X^2 σ_Y^2 T^2)) X X^\top Y Y^\top, whose imaginary part yields the eigenvalue density of C^\top C and hence the singular-value density of C. The paper then derives simplified closed-form band edges in several limits (Eqs. 20–25), confirms the formulas numerically for T=1000 with 500 independent realizations, and discusses implications for detecting shared signals when one or both dimensions exceed the sample size.
Significance. If correct, the paper supplies a parameter-free null distribution for singular values of an unwhitened sample cross-covariance matrix, extending the Marchenko–Pastur result and complementing earlier work on whitened cross-correlations. The S-transform derivation is standard and algebraically consistent, and the numerical simulations in Figs. 1–3 support the density and the edge formulas. The simplified edge formulas are useful for practitioners who want to calibrate significance of cross-correlations in high-dimensional data. The paper honestly acknowledges the equivalence of its auxiliary equation to earlier results (Refs. [11,14]) and the heuristic nature of the signal-detection discussion, which is an explicit strength rather than a hidden assumption.
minor comments (5)
- [Section II, Eqs. (13)–(14)] The notation δ(z) in the Stieltjes-transform relations is incorrect as written: a unit point mass at eigenvalue zero contributes 1/z to the Stieltjes transform, not δ(z). This matters especially for p_X>1 or p_Y>1, where the coefficient 1−p_X is negative and the intended cancellation with the zero-mass part of h_T is otherwise invisible. Please replace δ(z) by 1/z throughout these equations.
- [Appendix A, Eq. (A3)] The displayed definition of h(z) is missing the matrix inverse: it should read h(z)=lim_{T→∞} (1/T) E[Tr(zI−H)^{-1}], not Tr(zI−H).
- [Appendix B, Eqs. (B10)–(B19)] The derivations of the simplified edge formulas are asymptotic expansions of discriminants with statements such as “this can only happen if…” and “the contribution of higher-order terms will be negligible,” but no error bounds or rigorous justification of the root selection are given. Since Eqs. (20)–(25) are presented as results, please state explicitly that these are formal asymptotic expansions, and either provide control of the remainder terms or mark the formulas as asymptotically verified by the numerical experiments.
- [Section III, Eq. (15)] The paper does not specify how the physical root of the cubic equation is selected when computing the density from the imaginary part of h. For reproducibility, please state the branch rule (for example, the root whose imaginary part is positive for z in the lower half-plane) or describe the numerical root-selection procedure used for the semi-analytic curves.
- [Section II, paragraph after Eq. (3)] The statement that sampling fluctuations in estimating the scalar variances σ_X^2 and σ_Y^2 are negligible compared with spectral fluctuations is asserted but not analyzed. This is acceptable for the theorem, which is stated for the model with known variances, but because the paper motivates statistical significance of cross-correlations, it would be helpful to state explicitly that this is an additional modeling assumption and to indicate the size of the effect it can have on the band edges.
Circularity Check
No significant circularity: the cubic spectrum and edge bounds follow algebraically from textbook S-transform rules; the only self-citations are contextual and not load-bearing.
full rationale
The central derivation is self-contained. The matrix H is expressed as a scalar multiple of the free product of two normalized Wishart matrices (Eq. A1), the S-transform of a Wishart matrix is quoted from the external textbook Ref. [2], and the free-product rule (Eq. A7) with the scalar rule (Eq. A8) are applied. Inverting the T-transform then yields exactly the cubic equation for the Stieltjes transform, Eq. (A13)/(15). No parameter is fitted to data or to simulations, and no later quantity is defined in terms of the target spectrum. The band-edge formulas in Section III and Appendix B are obtained by setting the discriminant of this cubic to zero and solving in stated asymptotic limits; these are algebraic reductions, not empirical fits. The paper explicitly acknowledges that the cubic can be mapped onto Refs. [11,14], so the admitted equivalence affects novelty, not circularity. The only self-citations are Ref. [10] in Section II, used to justify neglecting scalar-variance estimation fluctuations, and Ref. [17] in the Discussion, cited for a numerical observation about signal detection; neither enters the S-transform algebra or the derivation of the spectral bounds. The numerical simulations are independent checks, not fitting steps. Thus no circular step is exhibited; the score of 2 reflects only the presence of minor non-load-bearing self-citations, not any reduction of the central result to its inputs.
Assumptions & free parameters
assumptions (4)
- domain assumption Independent Gaussian i.i.d. entries with known variances (Eq. 1)
- standard math WX and WY are asymptotically free, so the S-transform of their product is the product of the individual S-transforms (Eq. A7)
- standard math The S-transform of a Wishart matrix with aspect ratio p is 1/(1 + p t) (Eq. A10)
- standard math The support boundaries of the spectral density are found where the discriminant of the cubic for the Stieltjes transform vanishes (Appendix B)
Cite this review
Pith. "Pith review of Distribution of singular values in large sample cross-covariance matrices." pith.science (2026). https://pith.science/paper/RRVHR4EP
@misc{pith2026250205254,
author = {Pith},
title = {Pith review of: Distribution of singular values in large sample cross-covariance matrices},
year = {2026},
howpublished = {\url{https://pith.science/paper/RRVHR4EP}},
note = {Machine review of arXiv:2502.05254}
}
abstract
For two large matrices ${\mathbf X}$ and ${\mathbf Y}$ with Gaussian i.i.d.\ entries and dimensions $T\times N_X$ and $T\times N_Y$, respectively, we derive the probability distribution of the singular values of $\mathbf{X}^T \mathbf{Y}$ in different parameter regimes. This extends the Marchenko-Pastur result for the distribution of eigenvalues of empirical sample covariance matrices to singular values of empirical cross-covariances. Our results will help to establish statistical significance of cross-correlations in many data-science applications.
Figures
Forward citations
Cited by 1 Pith paper
-
Accurate Estimation of Mutual Information in High Dimensional Data
Neural MI estimators can become reliable in low-latent-dimension settings with a protocol of max-test early stopping, subsampling extrapolation, and probabilistic critics.
Reference graph
Works this paper leans on
-
[1]
Simplified solutions for pX = pY For pX = pY, the cubic equation for the Stieltjes trans- form, Eq. (A13), reduces to: h3z2pX 2 + h2z (pX (1 − pX ) +pX (1 − pX )) + h (1 − pX )(1 − pX ) − zpX 2 + p2 X = 0, (B1) and the discriminant (Eq. A15) simplifies to D = (4p4 X − 12p5 X + 12p6 X − 4p7 X )z3 + (−8p6 X − 20p7 X + p8 X )z4 + 4p8 X z5. (B2) Solving Eq. (...
-
[2]
Simplified solutions for pX < 1, pY ≪ pX For pY = αpX under the condition α → 0, the cubic equation for the Stieltjes transform Eq. (A13) reduces to: αh3z2p2 X + h2zpX (α(1 − pX ) + (1− αpX )) + h (1 − pX )(1 − αpX ) − zαp2 X + αp2 X = 0. (B9) The discriminant of Eq. (B9) is calculated using Eq. (A15). We then organize this discriminant as a poly- nomial ...
-
[3]
(A13) reduces to: αh3z2p2 X + h2zpX (α + 1) +h 1 − zαp2 X + αp2 X = 0
Simplified solutions for pX , pY ≪ 1 For pY = αpX under the conditionpX → 0 and α <1, the cubic equation for the Stieltjes transform Eq. (A13) reduces to: αh3z2p2 X + h2zpX (α + 1) +h 1 − zαp2 X + αp2 X = 0. (B15) The discriminant of Eq. (B15) is calculated using Eq. (A15). Written as a polynomial inz, it is D = 4z5α4p8 X + z4(α2p6 X − 10α3p6 X + α4p6 X −...
- [4]
-
[5]
Potters and J.-P
M. Potters and J.-P. Bouchaud,A First Course in Ran- dom Matrix Theory: For Physicists, Engineers and Data Scientists (Cambridge University Press, 2020)
2020
-
[6]
J.-P. Bouchaud, L. Laloux, M. A. Miceli, and M. Potters, The European Physical Journal B55, 201 (2007)
work page 2007
-
[7]
F. Benaych-Georges, J.-P. Bouchaud, and M. Potters, The Annals of Applied Probability33, 1295 (2023)
work page 2023
-
[8]
N. Firoozye, V. Tan, and S. Zohren, Journal of Banking & Finance154, 106952 (2023)
work page 2023
Show all 21 references
-
[9]
HOTELLING, Biometrika 28, 321 (1936), https://academic.oup.com/biomet/article-pdf/28/3- 4/321/586830/28-3-4-321.pdf
H. HOTELLING, Biometrika 28, 321 (1936), https://academic.oup.com/biomet/article-pdf/28/3- 4/321/586830/28-3-4-321.pdf
1936
-
[10]
Wold, Multivariate analysis , 391 (1966)
H. Wold, Multivariate analysis , 391 (1966)
1966
-
[11]
W. You, Z. Yang, and G. Ji, Computers in Biology and Medicine 174, 108434 (2024)
2024
-
[12]
K.-A. L. Cao, D. Rossouw, C. Robert-Granié, and P. Besse, Statistical Applications in Genetics and Molec- ular Biology7, doi:10.2202/1544-6115.1390 (2008)
2008
-
[13]
Fleig and I
P. Fleig and I. Nemenman, Phys. Rev. E106, 014102 (2022)
2022
-
[14]
J. W. Rocks and P. Mehta, Phys. Rev. E106, 025304 (2022)
2022
-
[15]
Burda, A
Z. Burda, A. Jarosz, G. Livan, M. A. Nowak, and A. Swiech, Phys. Rev. E82, 061114 (2010)
2010
-
[16]
P. J. Forrester, Journal of Physics A: Mathematical and Theoretical 47, 345202 (2014)
2014
-
[17]
Dupic and I
T. Dupic and I. P. Castillo, Spectral density of products of wishart dilute random matrices. part i: the dense case (2014), arXiv:1401.7802 [cond-mat.dis-nn]
2014 arXiv
-
[18]
Benigni and S
L. Benigni and S. Péché, Electronic Journal of Probabil- ity 26, 1 (2021)
2021
-
[19]
Pennington and P
J. Pennington and P. Worah, in Advances in Neu- ral Information Processing Systems , Vol. 30, edited by I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fer- gus, S. Vishwanathan, and R. Garnett (Curran Asso- ciates, Inc., 2017)
2017
-
[20]
Abdelaleem, A
E. Abdelaleem, A. Roman, K. M. Martini, and I. Ne- menman, Transactions on Machine Learning Research (2024)
2024
-
[21]
Voiculescu, in Operator Algebras and their Con- nections with Topology and Ergodic Theory , edited by H
D. Voiculescu, in Operator Algebras and their Con- nections with Topology and Ergodic Theory , edited by H. Araki, C. C. Moore, Ş.-V. Stratila, and D.-V. Voiculescu (Springer Berlin Heidelberg, Berlin, Heidel- berg, 1985) pp. 556–588
1985
Reviewed August 8, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.