Pith. sign in

REVIEW 4 major objections 4 minor 27 references

A New Two-Sample Test for Covariance Matrices in High Dimensions: U-Statistics Meet Leading Eigenvalues

T0 review · 4 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read A new two-sample test for high-dimensional covariance matrices combines a Frobenius-norm U-statistic with leading sample eigenvalues; the paper proves the two components are asymptotically independent under the null hypothesis.

desk verdict Solid new joint CLT and a practical hybrid test; the main proofs have omitted steps that a referee should push to have filled in. read the letter →

arxiv 2506.06550 v1 pith:SH2PZRSN submitted 2025-06-06 math.ST math.PRstat.TH

classification math.STmath.PRstat.TH MSC 15A1860F0562H15
keywords two-sampletestcovariancematriceshigh-dimensionalstatisticsspikedmodelleadingeigenvaluesU-statisticsp-valuecombinationcentrallimittheorem
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper proposes a single two-sample test for covariance matrices that works when the dimension grows proportionally to the sample size. The test combines two detectors: a U-statistic estimator of the squared Frobenius norm of the difference between the population covariance matrices, which sees dense changes, and the difference of leading sample eigenvalues, which sees sparse changes. The paper's main theoretical result is a joint central limit theorem showing these two components are asymptotically independent and standard normal under the null hypothesis, provided the leading population eigenvalues are supercritical. That independence justifies combining the two p-values through a classical p-value combination method, yielding a test statistic with an asymptotic chi-squared distribution with 4 degrees of freedom, so the level can be controlled asymptotically. A sympathetic reader would care because the test is consistent against both dense and sparse alternatives while existing methods tend to be specialized to one type.

What carries the argument

The central object is the pair (T_{n,1}, T_{n,2}). T_{n,1} is a U-statistic estimator of the trace of the squared difference of the two population covariance matrices, built from sums over distinct indices, standardized by a consistent variance estimator. T_{n,2} is the scaled difference of the two leading sample eigenvalues. The argument rests on identifying supercritical eigenvalues, meaning population eigenvalues strong enough to separate from the bulk spectrum, and on a martingale representation of both statistics; the joint central limit theorem follows once the cross-covariance terms are shown to vanish. For the multi-eigenvalue extension, the covariance matrix of the normalized leading eigenvalue differences is estimated from sample eigenvector components and kurtosis estimators.

What would settle it

Simulate, say, p equal to 1000 and n equal to 100, with i.i.d. standard normal entries and two samples drawn from the same covariance matrix that has one leading eigenvalue well above the bulk, so the paper's assumptions hold. Repeat many times and estimate the joint distribution of (T_{n,1}, T_{n,2}); if the empirical covariance matrix is not close to the identity or the marginals deviate systematically from standard normal, Theorem 2.1 would be contradicted.

Watch

Extended reading notes

Core claim

The paper establishes that, under the null hypothesis Sigma_n^(1) equals Sigma_n^(2) and under assumptions (A1)-(A4), the vector (T_{n,1}, T_{n,2}) consisting of the centered and scaled Frobenius-norm U-statistic and the scaled leading-eigenvalue difference converges in distribution to N_2(0, I_2); in particular, T_{n,1} and T_{n,2} are asymptotically independent. The proof works by replacing the leading sample eigenvalues with associated random quadratic forms, reducing to the centered-data case, and then applying a martingale central limit theorem whose key step is showing that the cross-covariance terms vanish. Consequently, the combined statistic T_{n,FC} converges to a chi-squared distribution with 4 degrees of freedom under the null hypothesis, and the test is consistent whenever either the squared Frobenius difference of the population covariance matrices diverges or the scaled leading-eigenvalue drift diverges. The same structure extends to K leading eigenvalues through a generalized statistic and a multivariate normal limit for the eigenvalue differences.

Load-bearing premise

The load-bearing premise is that both populations have at least one leading eigenvalue strong enough to be supercritical, meaning separated from the bulk spectrum, so that the sample leading eigenvalue is asymptotically Gaussian; without such an outlier, the null distribution of the eigenvalue component is not established.

Editorial extensions

If this is right

  • If the joint central limit theorem is correct, the combined test statistic T_{n,FC} has an asymptotic chi-squared null distribution with 4 degrees of freedom, so critical values are available without simulation under the null hypothesis.
  • The same reasoning yields a multi-eigenvalue version whose combined statistic again has an asymptotic chi-squared distribution with 4 degrees of freedom, provided the leading population eigenvalues are supercritical and separated.
  • Under alternatives with a diverging Frobenius difference or a diverging scaled leading-eigenvalue drift, the rejection probability tends to 1, so the test is consistent for both dense and sparse signal classes.
  • The asymptotic independence means the two detectors can be reported separately and combined without requiring a user to choose which type of alternative to optimize for.
  • The variance estimators are ratio-consistent under the null and under the alternative, so the level guarantee and the consistency result hold simultaneously.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A natural stress test the authors did not run: set both populations to have no supercritical eigenvalue and compare the empirical rejection rates of the combined test to the nominal level; the theory here predicts breakdown because the leading-eigenvalue component has no proven central limit theorem, but quantifying the distortion would map the method's applicability boundary.
  • The independence mechanism may extend to other pairs of statistics whose first component is a smooth spectral U-statistic and whose second is a spiked eigenvalue; if the vanishing cross-covariance argument generalizes, the same p-value combination strategy could equip other high-dimensional detectors.
  • In practice, estimating kurtosis and eigenvector alignments from the sample introduces finite-sample sensitivity, so the chi-squared approximation likely degrades when the estimated spike variances are near the boundary of the supercritical regime, a region the paper's appendix probes only for eigenvalue multiplicity violations.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The manuscript proposes a two-sample test for equality of p-dimensional covariance matrices in the high-dimensional regime p/n -> y in (0,infinity). The test statistic T_{n,FC} combines, via Fisher's method, a Frobenius-norm U-statistic of Li and Chen (2012) with a leading-eigenvalue statistic of Zhang et al. (2022). The central theoretical result (Theorem 2.1) states that under the null hypothesis and the spiked-eigenvalue assumptions (A1)-(A4), the two components are asymptotically standard normal and independent, so T_{n,FC} is asymptotically chi-squared with 4 degrees of freedom. A multi-spike extension (Theorems 2.3-2.4) is also provided, and power consistency is claimed for dense alternatives (large Frobenius difference) and sparse alternatives (large gap in leading eigenvalues). The proofs are based on martingale central limit theorems. However, several load-bearing steps are omitted: Lemma 4.5 (the high-probability event) is stated without proof, the cross-term bounds C3-C5 in Lemma 4.6 are declared analogous and not proved, and the consistency of the kurtosis estimator in the p/n in (0,infinity) regime is asserted without proof. The simulation study compares the new test with existing methods under normal, t7, and Laplace data.

Significance. If the missing proofs can be supplied, the result is a useful contribution: it addresses a question left open in Zhang et al. (2022) by establishing asymptotic independence between the Frobenius-norm U-statistic and the leading eigenvalues, and it yields a simple Fisher-combined test that is sensitive to both dense and sparse alternatives. The multi-eigenvalue extension is natural, and the paper is transparent about which results are imported from Li and Chen (2012) and Zhang et al. (2022) and about the spiked-eigenvalue scope. The main reservations are completeness: the central independence claim is not closed within the manuscript, and the kurtosis consistency extension is asserted without proof. These issues are fixable in revision.

major comments (4)
  1. [Section 4.5 (Lemma 4.5)] Lemma 4.5, the high-probability event B_n, is stated but its proof is omitted with the explanation that it is 'similar to proof of Lemma 4.2 in Zhang et al. (2022), which can be found in the supplement to their paper.' This event is used before truncation and in the martingale representation: the indicator functions in (4.10) and the bounds in (4.23) depend on it. Since the central theorem rests on the martingale CLT for the truncated variables, the proof of Lemma 4.5 must be included in the manuscript or in a supplement, not merely referenced.
  2. [Section 4.5.1 (Lemma 4.6, Eq. (4.18))] The off-diagonal covariance between the Frobenius statistic and the eigenvalue statistic is written in (4.18) as a sum of five cross terms C1,n through C5,n, and asymptotic independence requires each term to be o_P(1). The proof establishes only C1,n and C2,n for j=1; C3,n, C4,n, and C5,n are declared 'analogous' and omitted. In particular, C5,n contains the cross-sample matrix G^(1,2)_{n,n} built from the first sample, so it is not an immediate consequence of the independence of the two samples; it is precisely the term coupling the second-sample eigenvalue statistic with the first-sample U-statistic. The proof of Lemma 4.6 must be completed for all five terms and for all j=1,...,K.
  3. [Appendix A.2 (Lemma A.4(d))] The consistency of the kurtosis estimator \hat\gamma_{4,n} is stated for p/n in (0,infinity), while the cited result of Dörnemann and Dette (2024) is only for p/n in (0,1). The sentence 'A closer examination of the proof reveals that the consistency of this estimator remains valid in the case p/n in (0,infinity)' is an unproved extension. Since Lemma 4.1 uses \hat\gamma_{4,n} to estimate the normalization of both components of the test statistic, this is load-bearing. Please provide the proof or a precise reference to a theorem covering the full p/n in (0,infinity) regime.
  4. [Section 4.3 (proof of Lemma 4.2)] The reduction from the leading sample eigenvalues to the quadratic forms Q^(i)_{j,n} is made by citing 'identity (S.3.9)' and 'similar arguments to (S.4.4)' in the supplement to Zhang et al. (2022), without stating those identities in the paper. This reduction is an essential step in the proof of Theorem 2.3. Please state the identities explicitly and either prove them or provide a self-contained derivation in the appendix.
minor comments (4)
  1. [Section 3 (figure captions)] The acronym 'LDD' is used in Figures 1 and 2 without being defined; specify that it denotes the new test (2.2), and similarly define LDD(1) and LDD(3) in Figures 3-5.
  2. [Section 4.5 and 4.5.1] There are several typos that impede reading, including 'Due the the CLTs' at the start of the proof of Lemma 4.4 and 'there exists there exists' in the proof of Lemma 4.6; please correct them throughout.
  3. [Appendix A.2] The estimator \hat\sigma^2_{n,1} is defined as 4/n^2 (B^(1)_n+B^(2)_n)^2, whereas in Section 2.1 it is described as an estimator of Var(B^(1)_n+B^(2)_n-2C_n); state explicitly that this is the estimator under H0 and how the cross term C_n enters under H0.
  4. [Section 3 and Remark 2.1] The simulations with t7-distributed data do not satisfy Assumption (A2), which requires a finite eighth moment; please state clearly that these results are robustness checks outside the theorem's assumptions, rather than presenting them on the same footing as the normal and Laplace cases.

Circularity Check

0 steps flagged · score 2.0 of 10

No substantive circularity; the joint CLT is derived by martingale arguments, and the only self-citation is ancillary.

full rationale

The paper's central claim is Theorem 2.1/2.3: asymptotic independence of the Frobenius-norm U-statistic and the leading spiked eigenvalues. The proof chain (Sections 4.3-4.5) is: Lemma 4.1 (variance-estimator consistency), Lemma 4.2 (reduction to quadratic forms), Lemma 4.3 (centering reduction), and Lemma 4.4 (martingale CLT). Independence is not assumed or built into the definitions; it is established by showing that the conditional-covariance cross terms in (4.17) are oP(1), with C1,n and C2,n bounded in Lemma 4.6 and C3,n-C5,n declared analogous. No parameter is fitted to data and then renamed as a prediction; the identity covariance block in Theorem 2.3 is a derived limit, not an input. The paper does import marginal CLTs and technical identities from Li and Chen (2012) and Zhang et al. (2022), and it explicitly omits the proof of Lemma 4.5 (the high-probability event B_n) and the C3,n-C5,n proofs; these are completeness and verifiability gaps, not circularity. The Dörnemann and Dette (2024) kurtosis estimator is a self-citation involving two of the three authors, but it is an ancillary implementation detail for variance estimation and does not feed back into the asymptotic-independence derivation. The supercritical-eigenvalue assumption (A4)/(A5) restricts the scope of the test but is not a circular reuse of the conclusion. Overall, no derivation step reduces to its own input, so the circularity is not substantive.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

The paper introduces no free parameters: all quantities in the test statistic are either data dependent estimators or user-specified nominal levels. The central claim rests on a set of standard random matrix assumptions, the most restrictive of which is the supercritical spike assumption (A4)/(A5). Without that assumption the leading eigenvalue component lacks a Gaussian null limit.

assumptions (5)
  • domain assumption (A1): p/n converges to y in (0, infinity) as n goes to infinity.
    This fixes the high-dimensional proportional regime that defines the paper's scope.
  • domain assumption (A2): the entries of z_j^(i) are i.i.d. with zero mean, unit variance, and finite eighth moment.
    This strengthens the fourth-moment assumption in Zhang et al. (2022) to accommodate the U-statistic CLT for the Frobenius norm.
  • domain assumption (A3): the population covariance matrices admit nonrandom limiting spectral distributions and their smallest eigenvalues are bounded away from zero.
    This technical condition provides spectral regularity and is used in variance bounds and Stieltjes transform arguments.
  • domain assumption (A4)/(A5): the leading K population eigenvalues are supercritical, do not depend on n, and are well separated from the bulk and from each other.
    This is the load-bearing spiked model assumption: without supercritical spikes, the leading sample eigenvalues do not satisfy the required Gaussian CLT and the eigenvalue test component is not justified.
  • domain assumption (A6): limits of certain fourth-moment sums of eigenvector entries exist.
    This ensures the existence of the asymptotic variances and covariances of the spiked sample eigenvalues.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A New Two-Sample Test for Covariance Matrices in High Dimensions: U-Statistics Meet Leading Eigenvalues." pith.science (2026). https://pith.science/paper/SH2PZRSN

@misc{pith2026250606550,
  author       = {Pith},
  title        = {Pith review of: A New Two-Sample Test for Covariance Matrices in High Dimensions: U-Statistics Meet Leading Eigenvalues},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/SH2PZRSN}},
  note         = {Machine review of arXiv:2506.06550}
}
read the original abstract

We propose a two-sample test for covariance matrices in the high-dimensional regime, where the dimension diverges proportionally to the sample size. Our hybrid test combines a Frobenius-norm-based statistic as considered in Li and Chen (2012) with the leading eigenvalue approach proposed in Zhang et al. (2022), making it sensitive to both dense and sparse alternatives. The two statistics are combined via Fisher's method, leveraging our key theoretical result: a joint central limit theorem showing the asymptotic independence of the leading eigenvalues of the sample covariance matrix and an estimator of the Frobenius norm of the difference of the two population covariance matrices, under suitable signal conditions. The level of the test can be controlled asymptotically, and we show consistency against certain types of both sparse and dense alternatives. A comprehensive numerical study confirms the favorable performance of our method compared to existing approaches.

Figures

Figures reproduced from arXiv: 2506.06550 by the authors.

Figure 1
Figure 1. Simulated rejection probabilities of the test (2.2) (solid line, LDD), the test of Zhang et al. (2022) (dashed line, ZZPZ) and the test of Li and Chen (2012) (dotted line, LC) in model (3.1). Left panels: (p, n) = (500, 100). Right panels: (p, n) = (1000, 100). First row: z (i) 1,1 ∼ N (0, 1), second row: z (i) 1,1 ∼ t7/ q 7/5, third row: z (i) 1,1 ∼ Laplace(0, 1/ √ 2). 13 [PITH_FULL_IMAGE:figures/full_fig_p013_1.png] view at source ↗
Figure 2
Figure 2. Simulated rejection probabilities of the test (2.2) (solid line, LDD), the test of Zhang et al. (2022) (dashed line, ZZPZ) and the test of Li and Chen (2012) (dotted line, LC) in model (3.2). Left panels: (p, n) = (500, 100). Right panels: (p, n) = (1000, 100). First row: z (i) 1,1 ∼ N (0, 1), second row: z (i) 1,1 ∼ t7/ q 7/5, third row: z (i) 1,1 ∼ Laplace(0, 1/ √ 2). 15 [PITH_FULL_IMAGE:figures/full_fig_p015_2.png] view at source ↗
Figure 3
Figure 3. Simulated rejection probabilities of the test (2.2) (solid line, LDD(1)), the test of (2.9) (dot dashed line, LDD(3)), the test of Zhang et al. (2022) (dashed line, ZZPZ) and the test of Li and Chen (2012) (dotted line, LC) in model (3.3). Left panels: (p, n) = (500, 100). Right panels: (p, n) = (1000, 100). First row: z (i) 1,1 ∼ N (0, 1), second row: z (i) 1,1 ∼ t7/ q 7/5, third row: z (i) 1,1 ∼ Laplace(0, 1/ √ 2)… view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Simulated rejection probabilities of the test (2.2) (solid line, LDD(1)), the test of [PITH_FULL_IMAGE:figures/full_fig_p019_4.png]
Figure 5
Figure 5. Figure 5: Simulated rejection probabilities of the test (2.2) (solid line, LDD(1)), the test of (2.9) (dot dashed line, LDD(3)), the test of Zhang et al. (2022) (dashed line, ZZPZ) and the test of Li and Chen (2012) (dotted line, LC) in model (A.1). Left panels: (p, n) = (500, 1…

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

27 extracted references · 25 canonical work pages

  1. [1]

    Anderson, T. W. (1984). An Introduction to Multivariate Statistical Analysis . Wiley Series in Probability and Mathematical Statistics: Probability and Mathematical Statistics. John Wiley & Sons, Inc., New York, second edition

  2. [2]

    and Ding, X

    Bai, Z. and Ding, X. (2012). Estimation of spiked eigenvalues in spiked models. Random Matrices: Theory and Applications , 1(2):1150011

  3. [3]

    and Saranadasa, H

    Bai, Z. and Saranadasa, H. (1996). Effect of high dimension: by an example of a two sample problem. Statistica Sinica , pages 311--329

  4. [4]

    and Silverstein, J

    Bai, Z. and Silverstein, J. W. (2010). Spectral analysis of large dimensional random matrices , volume 20. Springer

  5. [5]

    and Yao, J.-f

    Bai, Z. and Yao, J.-f. (2008). Central limit theorems for eigenvalues in a spiked population model. Annales de l'IHP Probabilit \'e s et statistiques , 44(3):447--474

  6. [6]

    and Yin, Y.-Q

    Bai, Z.-D. and Yin, Y.-Q. (2008). Limit of the smallest eigenvalue of a large dimensional sample covariance matrix. In Advances In Statistics , pages 108--127. World Scientific

  7. [7]

    B., and P \'e ch \'e , S

    Baik, J., Arous, G. B., and P \'e ch \'e , S. (2005). Phase transition of the largest eigenvalue for nonnull complex sample covariance matrices . The Annals of Probability , 33(5):1643 -- 1697

  8. [8]

    L., and Xia, Y

    Cai, T., Liu, W. L., and Xia, Y. (2013). Two-sample covariance matrix testing and support recovery in high-dimensional and sparse settings. Journal of the American Statistical Association , 108(501):265--277

Show all 27 references
  1. [9]

    and D \"o rnemann, N

    Dette, H. and D \"o rnemann, N. (2020). Likelihood ratio tests for many groups in high dimensions. Journal of Multivariate Analysis , 178:104605

  2. [10]

    Ding, X., Hu, Y., and Wang, Z. (2024). Two sample test for covariance matrices in ultra-high dimension. Journal of the American Statistical Association , pages 1--12

  3. [11]

    and Dette, H

    D \"o rnemann, N. and Dette, H. (2024). Detecting change points of covariance matrices in high dimensions. arXiv preprint arXiv:2409.15588

  4. [12]

    Dörnemann, N. (2022). Likelihood ratio tests under model misspecifications in high dimensions. arXiv preprint arXiv:2203.05423

  5. [13]

    Fisher, R. A. (1950). Statistical Methods for Research Workers . Oliver and Boyd, London, 11th edition

  6. [14]

    and Heyde, C

    Hall, P. and Heyde, C. (1980). Martingale Limit Theory and its Application . Academic Press

  7. [15]

    and Yang, F

    Jiang, T. and Yang, F. (2013). Central limit theorems for classical likelihood ratio tests for high-dimensional normal distributions. Annals of Statistics , 41:2029--2074

  8. [16]

    and Chen, S

    Li, J. and Chen, S. X. (2012). Two sample tests for high-dimensional covariance matrices . The Annals of Statistics , 40(2):908 -- 940

  9. [17]

    Li, Z., Han, F., and Yao, J. (2020). Asymptotic joint distribution of extreme eigenvalues and trace of large sample covariance matrix in a generalized spiked population model . The Annals of Statistics , 48(6):3138 -- 3160

  10. [18]

    Littell, R. C. and Folks, J. L. (1971). Asymptotic optimality of fisher’s method of combining independent tests. Journal of the American Statistical Association , 66(336):802--806

  11. [19]

    Liu, Z., Hu, J., Bai, Z., and Song, H. (2023). A clt for the lss of large-dimensional sample covariance matrices with diverging spikes. The Annals of Statistics , 51(5):2246--2271

  12. [20]

    Mirsky, L. (1975). A trace inequality of John von Neumann . Monatshefte für Mathematik , 79:303--306

  13. [21]

    Paul, D. (2007). Asymptotics of sample eigenstructure for a large dimensional spiked covariance model. Statistica Sinica , pages 1617--1642

  14. [22]

    Qi, Y., Wang, F., and Zhang, L. (2019). Limiting distributions of likelihood ratio test for independence of components for high-dimensional normal vectors. Annals of the Institute of Statistical Mathematics , 71(4):911--946

  15. [23]

    S., Yanagihara, H., and Kubokawa, T

    Srivastava, M. S., Yanagihara, H., and Kubokawa, T. (2014). Tests for covariance matrices in high dimension with less sample size. Journal of Multivariate Analysis , 130:289--309

  16. [24]

    and Pan, G

    Yang, Q. and Pan, G. (2017). Weighted statistic in detecting faint and sparse alternatives for high-dimensional covariance matrices. Journal of the American Statistical Association , 112(517):188--200

  17. [25]

    Q., Bai, Z., and Krishnaiah, P

    Yin, Y. Q., Bai, Z., and Krishnaiah, P. R. (1988). On the limit of the largest eigenvalue of the large dimensional sample covariance matrix. Probability Theory and Related Fields , 78:509--521

  18. [26]

    Zhang, Z., Zheng, S., Pan, G., and Zhong, P.-S. (2022). Asymptotic independence of spiked eigenvalues and linear spectral statistics for large sample covariance matrices. The Annals of Statistics , 50(4):2205--2230

  19. [27]

    Zheng, S., Lin, R., Guo, J., and Yin, G. (2020). Testing homogeneity of high-dimensional covariance matrices. Statistica Sinica , 1:35--53

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.