REVIEW 3 major objections 5 minor 44 references
Asymptotic analysis of high-dimensional uniformity tests under heavy-tailed alternatives
T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read The packing test is the one that catches heavy-tailed alternatives on high-dimensional spheres, while the Rayleigh and Bingham tests fail.
desk verdict Real new results on packing-test consistency and asymptotic independence of three uniformity statistics; the Fisher test's size guarantee depends on an imported lemma that needs a genuine transfer proof. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the $\varepsilon$-good event for a vector $X$: the condition $\max_i|X_i|/\|X\|_2 \ge 1-\varepsilon$, meaning one coordinate dominates the norm. Under an $\alpha$-spherical law this event occurs with positive limiting probability $C_{\alpha,\varepsilon}$, and conditionally on it the argmax coordinate is uniform over $\{1,\dots,p\}$; therefore with $p=o(n^2)$ two independent $\varepsilon$-good vectors share the same dominant coordinate, forcing their inner product close to $1$. The independence result is carried by an inclusion-exclusion decomposition of the extreme event $\{P_n\ge z\}$, a leave-one-out and martingale argument for the joint normality of $R_n$ and $B_n$, and a combinatorial bound on the higher-order inclusion-exclusion terms. For the FvML corollary, the machinery is the LAN expansion together with Le Cam's third lemma, which lets the Gumbel null law of $P_n$ pass to the alternative without a Chen-Stein Poisson approximation.
What would settle it
Compute the quantity $N(n,k)=\sum_{I_1<\cdots<I_k}P(C_{I_1}\cdots C_{I_k})$ directly for $X_i\sim\mathrm{Unif}(\mathbb{S}^{p-1})$ with $p/(\log n)^2\to\infty$: if it is not eventually bounded by a constant multiple of $1/k!$, the imported bound in Lemma 9 fails and Theorem 4 collapses. A finite check is also possible: simulate the joint law of $(R_n,B_n,P_n)$ under uniformity for large $p$ and compare the empirical Fisher-test size with the nominal $\beta$; systematic inflation would signal that the independence assumption is not transferring.
Extended reading notes
Core claim
On the paper's own terms, the central finding is a separation of detection capabilities. For symmetric $\alpha$-spherical data with $p/n^2\to 0$, Theorem 1 shows $R_n\xrightarrow{d}N(0,1)$ exactly as under uniformity, so the Rayleigh test's asymptotic power equals its nominal size. For $p/n\to\gamma\in(0,\infty)$, Theorem 2 shows $\sqrt{n}\,B_n/p\xrightarrow{d}N(0,(2-\alpha)^2/(8\gamma))$, which puts the Bingham test's rejection probability at $1/2$ regardless of nominal level. Theorem 3 shows the packing statistic's core quantity, $\max_{i<j}(X_i^\top X_j)^2$, converges to $1$ in probability under $\alpha$-spherical data with $p=o(n^2)$, while under uniformity it is of order $4\log n/p$; hence $P_n$ is consistent. Theorem 4 then states that under uniformity and $p/(\log n)^2\to\infty$, $(R_n,B_n,P_n)$ converges in distribution to the product of two independent standard normals and a Gumbel variable, and the resulting Fisher combination test $\phi_n^F$ has asymptotic size $\beta$ and inherits the optimality of whichever component test detects a given alternative.
Load-bearing premise
The proof of Theorem 4 borrows a combinatorial bound for inclusion-exclusion terms that was originally proved for max-sum statistics in panel data, and the paper does not demonstrate that the bound's hypotheses are satisfied for uniform data on the hypersphere; if that transfer fails, the claimed asymptotic independence of $P_n$ from $(R_n,B_n)$ and the Fisher test's size guarantee are not established.
Editorial extensions
If this is right
- The packing test becomes the first of the three high-dimensional uniformity tests proven consistent against a class of alternatives, namely $\alpha$-spherical heavy-tailed projections under $p=o(n^2)$.
- The Rayleigh and Bingham tests should not be used alone for data that may be heavy-tailed: their asymptotic powers are respectively the nominal level and $1/2$ under the stated regimes.
- The Fisher combination test $\phi_n^F$ inherits the known minimax optimality of the Rayleigh test for FvML alternatives and of the Bingham test for Watson alternatives, while gaining consistency against heavy-tailed alternatives.
- Under FvML alternatives at the optimal rate $\kappa_n=\tau n^{3/4}/\sqrt{n}$, the packing statistic has the same Gumbel limit as under uniformity, so it contributes no power at that threshold.
- The law of large numbers for the largest off-diagonal entry of a sample correlation matrix with regularly varying entries stands as a separate heavy-tailed result.
Reading between the lines
- Because Lemma 9 is imported from a panel-data max-sum setting without a written verification of its hypotheses on the hypersphere, the asymptotic-independence theorem is, as written, conditional on that transfer; checking the bound $N(n,k)$ directly under $\mathrm{Unif}(\mathbb{S}^{p-1})$ would complete the argument.
- The $\varepsilon$-good and argmax mechanism suggests the packing test's consistency may extend to any heavy-tailed or sparse model in which a positive fraction of vectors have a dominant coordinate, since only the collision of two dominant indices is required; testing this for sparse projections would be a natural next experiment.
- The Le Cam third lemma route to non-null extreme-value distributions may generalize to other extreme-value statistics whenever a LAN expansion holds, avoiding Poisson approximation; this seems to be the paper's intended methodological lesson.
- Because the Fisher combination threshold is calibrated from asymptotic independence, its practical size may suffer in the same moderate-$p$ regime where the packing test's size is distorted; the simulations only cover $p\le 120$.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies uniformity testing on the p-dimensional sphere when p and n both diverge, focusing on a new class of alternatives: projections of i.i.d. heavy-tailed coordinates (α-spherical distributions). Theorem 1 shows the Rayleigh test remains asymptotically N(0,1) under such alternatives when p/n^2 → 0; Theorem 2 shows the Bingham test has asymptotic power 1/2 when p/n → γ. Theorem 3 claims the packing test is consistent: the largest squared inner product converges to 1 in probability when p = o(n^2). Theorem 4 claims that under uniformity and p/(log n)^2 → ∞, the Rayleigh, Bingham, and packing statistics are asymptotically independent, leading to a Fisher combination test with controlled size and power against any alternative detected by one of the three tests. The paper also sketches a Le Cam third lemma argument for the non-null distribution of the packing test under FvML alternatives.
Significance. The paper addresses an open problem: consistency of the packing test. It identifies a broad class of heavy-tailed alternatives for which the two classical U-statistic tests fail while an extreme-value statistic succeeds, and it proposes a combination procedure. If the gaps identified below are fixed—especially the transfer of Lemma 9 from [11] to the hypersphere setting—the results would be a valuable contribution to high-dimensional directional statistics. The martingale CLT computations for Theorems 1–2 are careful, and the moment estimates (Lemmas 4–7) are explicit. The paper also offers a genuinely different route to non-null extreme-value distributions via Le Cam's third lemma, which is of independent interest. However, as written, the proof of Theorem 4 is not self-contained, and Theorem 3's statement is stronger than what is proved.
major comments (3)
- [Section 2.3 / Section 6.3, Theorem 3] Theorem 3 states max_{1≤i<j≤n} X_i^T X_j → 1 in probability, but the proof in Section 6.3 establishes only max_{1≤i<j≤n} |X_i^T X_j| ≥ 1−4ε+2ε² with high probability. Since the packing statistic Pn in (4) depends on (X_i^T X_j)^2, the squared version of the theorem would suffice for consistency; the signed-maximum statement is stronger and is not proved. Please restate the theorem or add an argument that a positive inner product close to 1 occurs with high probability.
- [Section 7, proof of (20); Lemma 9 in Section 8] Lemma 9 asserts lim_{k→∞} limsup_{n→∞} N(n,k) = 0 with the factorial bound (42), citing 'Lemma 35 in [11] with p=T and n=N' and equation (S.82). No verification is provided that the hypotheses of that panel-data lemma—which concerns max-sum statistics of cross-sectional sums—are satisfied by the exceedance events C_I defined in (29) for inner products of uniform vectors on the hypersphere. The dependence among C_I events is governed by Gram-matrix constraints and is not obviously of the same form. Because N(n,k) is used in inequality (32) to pass from inclusion-exclusion bounds to the asymptotic independence (20), Theorem 4 and the size control of the Fisher test in (6) are not established as written. Please provide a self-contained proof or a detailed transfer argument.
- [Section 6.3, proof of Theorem 3] The proof conditions on the event An that at least ⌊nC_{α,ε}⌋/4 vectors are ε-good, then invokes Lemma 2 to assert that the indices i(X_{i_j}) of these selected vectors are i.i.d. uniform on {1,…,p}. Lemma 2 is stated conditional on all X_i being ε-good, not merely on a subset of size K. The needed statement—that conditional on being ε-good, each argmax index is uniform and independent across observations—is plausible but is not what Lemma 2 says. Please correct the lemma or the invocation.
minor comments (5)
- [Section 8, Lemma 10] A^-_{n,ε} is defined as {Rn ≥ x−ε, Bn ≥ y−ε}; the subsequent sandwich argument requires A^-_{n,ε} ⊂ An, so the inequalities should be ≤. Please correct the definition.
- [Section 4, Tables 1 and 2] The first column of both tables is labeled 'n=80, p=40', while the text states that the scenarios are (80,100), (100,100), and (100,120). Please reconcile the table labels with the text.
- [Section 2.1] The section ends with the incomplete sentence 'As a result, the'; this should be completed or removed.
- [Section 6.3] There are typos in the proof: 'argmax_{1≤≤n}' should be 'argmax_{1≤i≤p}', and 'max_{1≤i<≤p}' should be 'max_{1≤i<j≤n}' in several displays.
- [Section 8, Lemma 9] The citation 'Lemma 35 in [11] with p=T and n=N' is cryptic; please spell out the mapping between the notation of [11] and the present setting so that the claimed transfer can be checked.
Circularity Check
No significant circularity: Theorem 4's asymptotic independence is derived from a martingale CLT and an independently published inclusion-exclusion bound, not from a fitted or definitionally equivalent input.
full rationale
The main claims (Theorems 1-3) are derived within the paper from stated alpha-spherical assumptions: martingale CLT with explicit moment bounds from Lemmas 5-7, and the epsilon-good argument with Lemma 3 from Darling's theorem. These are forward derivations, not fits renamed as predictions. Theorem 4's new content is asymptotic independence; its proof is largely self-contained (verification of (18) and the inclusion-exclusion framework), and the one imported ingredient, the factorial-decay estimate N(n,k) <= C/k!, is quoted as an already proved Lemma 35 of Feng-Jiang-Liu-Xiong [11] with the variable identification p=T, n=N and bound (S.82). A citation whose authors overlap with the present paper does not by itself constitute circularity, and a published, parameter-free result quoted with its own proof is treated as independent support. The legitimate concern is that the paper gives no argument that the panel-data hypotheses of Lemma 35 transfer to the hypersphere exceedance events C_I in (29); if they do not, inequality (32) and the size guarantee of phi_F^n would fail. That is a correctness or rigor gap, not a circular reduction: no equation of the paper equals its conclusion by construction, and no fitted value is relabeled as a predicted quantity. The secondary display issue in Lemma 10 (sandwich definition using '>=' where a shrinking event is needed) is a typographical slip and likewise not circular. Accordingly the circularity score is 0.
Assumptions & free parameters
assumptions (6)
- domain assumption Coordinates are i.i.d. with regularly varying tail of index alpha in (0,2), projected onto the sphere (Definition 1).
- domain assumption Symmetric alpha-spherical distributions for Theorems 1 and 2.
- standard math Martingale central limit theorem (Hall and Heyde, [16]) applied to Rayleigh, Bingham, and their linear combination.
- standard math Gumbel null distribution for the packing statistic from Cai-Fan-Jiang [5], requiring p/(log n)^2 -> infinity.
- domain assumption Lemma 9, imported from [11] as 'content of Lemma 35', supplies Poisson-type inclusion-exclusion bounds for the maximum.
- standard math LAN expansion for FvML alternatives from Cutting-Paindaveine-Verdebout [7].
Cite this review
Pith. "Pith review of Asymptotic analysis of high-dimensional uniformity tests under heavy-tailed alternatives." pith.science (2026). https://pith.science/paper/YDOQUF2V
@misc{pith2026250600393,
author = {Pith},
title = {Pith review of: Asymptotic analysis of high-dimensional uniformity tests under heavy-tailed alternatives},
year = {2026},
howpublished = {\url{https://pith.science/paper/YDOQUF2V}},
note = {Machine review of arXiv:2506.00393}
}
abstract
We study the high-dimensional uniformity testing problem, which involves testing whether the underlying distribution is the uniform distribution, given $n$ data points on the $p$-dimensional unit hypersphere. While this problem has been extensively studied in scenarios with fixed $p$, only three testing procedures are known in high-dimensional settings: the Rayleigh test \cite{Cutting-P-V}, the Bingham test \cite{Cutting-P-V2}, and the packing test \cite{Jiang13}. Most existing research focuses on the former two tests, and the consistency of the packing test remains open. We show that under certain classes of alternatives involving projections of heavy-tailed distributions, the Rayleigh test is asymptotically blind, and the Bingham test has asymptotic power equivalent to random guessing. In contrast, we show theoretically that the packing test is powerful against such alternatives, and empirically that its size suffers from severe distortion due to the slow convergence nature of extreme-value statistics. By exploiting the asymptotic independence of these three tests, we then propose a new test based on Fisher's combination technique that combines their strengths. The new test is shown to enjoy all the optimality properties of each individual test, and unlike the packing test, it maintains excellent type-I error control.
Reference graph
Works this paper leans on
-
[11]
Banerjee and J
A. Banerjee and J. Ghosh. Frequency sensitive competitive learning for scalable balanced clustering on high-dimensional hyperspheres. IEEE T. Neural Network. , 15:702–719, 2004
2004
-
[1]
Rayleigh test in [7] and [27] . This test can be formulated in terms of a U-statistic of the data points with the inner product kernel, i.e. Rn := √2p n X 1≤i<j≤n X ⊤ i Xj. (2)
-
[2]
Bingham test in [8, 37] and [27] . This test is also based on a U-statistic of the data points, but with a quadratic inner product kernel, i.e. Bn := p n X 1≤i<j≤n h X ⊤ i Xj 2 − 1 p i . (3)
-
[3]
Packing test in [5]. This test is based on the smallest angle, i.e. Pn := p · max 1≤i<j≤n X ⊤ i Xj 2 − 4 logn + log logn. (4) It is known that the Rayleigh test Rn and the Bingham test Bn enjoy a doubly robust property: under the null hypothesis and the single assumption min {n, p} → ∞, both Rn and Bn converge in distribution to the standard normal distri...
-
[4]
E X ⊤ 1 X2 X1 = 0 almost surely
-
[5]
E h X ⊤ 1 X2 2 X1 i = 1/p almost surely
-
[6]
E h X ⊤ 1 X2 · X ⊤ 1 X3 X2, X3 i = p−1 · X ⊤ 2 X3 almost surely. Proof of Lemma 5 . For the first claim, write E X ⊤ 1 X2 X1 = pX k=1 X1kE X2k X1 = pX k=1 X1kE (X2k) = 0. The second equality in the expression below holds almost surely so the first statement follows. To prove the second statement, write E h X ⊤ 1 X2 2 X1 i = pX k=1 X 2 1kE X 2 2k X1 + 2 X ...
-
[7]
Then, we have lim n→∞ Var 1 n n−1X i=2 i−1X k=1 h4 (Xk) ! = 0 and lim n→∞ Var 1 n n−1X i=2 X 1≤u̸=v≤i−1 L(Xu, Xv) ! = 0. Proof of Lemma 7 . Let us start with the first statement. Note that since h4(Xi)’s are i.i.d., we have Var 1 n n−1X i=2 i−1X k=1 h4 (Xk) ! = 1 n2 Var n−1X i=1 (n − i) · h4(Xi) ! = 1 n2 · O(n3) · Var (h4(X1)) . It suffices to show that n...
Show all 44 references
-
[8]
Albisetti, F
I. Albisetti, F. Balabdaoui, and H. Holzmann. Testing for spherical and elliptical symmetry. Journal of Multivariate Analysis , 180:104667, 2020
2020
-
[9]
Bak and D.J
J. Bak and D.J. Newman. Complex Analysis. Springer New York, 2010
2010
-
[10]
Banerjee, I
A. Banerjee, I. Dhillon, J. Ghosh, and S. Sra. Generative model-based clustering of direc- tional data. Proceedings of the ninth ACM SIGKDD international conference on Knowledge discovery and data mining , pages 19–28, 2003
2003
-
[12]
T. Cai, J. Fan, and T. Jiang. Distributions of angles in random packing on spheres. J. Mach. Learn. Res., 14:1837–1864, 2013
2013
-
[13]
Chan and P
Y-B. Chan and P. Hall. Robust nearest-neighbor methods for classifying high-dimensional data. Ann. Statist. 37(6A): 3186-3203. , 37(6A):3186–3203, 2009
2009
-
[14]
Cutting, D
C. Cutting, D. Paindaveine, and T. Verdebout. Testing uniformity on high-dimensional spheres against contiguous rotationally symmetric alternatives. Ann. Stat. , 45:1024–1058, 2017. 38
2017
-
[15]
Cutting, D
C. Cutting, D. Paindaveine, and T. Verdebout. Testing uniformity on high-dimensional spheres: The non-null behaviour of the bingham test. Annales de I’I.H.P. Probabilit´ es et statistiques, 58:567–602, 2022
2022
-
[16]
D. A. Darling. The influence of the maximum term in the addition of independent random variables. Transactions of the American Mathematical Society , 73(1):95–107, 1952
1952
-
[17]
I. L. Dryden. Statistical analysis on high-dimensional spheres and shape spaces. Ann. Stat., 33:1643–1665, 2005
2005
-
[18]
L. Feng, T. Jiang, B. Liu, and W. Xiong. Max-sum tests for cross-sectional independence of high-demensional panel data. Ann. Stat. 50(2), 1124-1143. , 2022
2022
-
[19]
Fern´ andez-de Marcos and E
A. Fern´ andez-de Marcos and E. Garc ´ ıa-Portugu´ es. On new omnibus tests of uniformity on the hypersphere. Test, 32(4):1508–1529, 2023
2023
-
[20]
N. I. Fisher, T. Lewis, and B. J. Embleton. Statistical Analysis of Spherical Data . Cambridge Univ. Press press, 1987
1987
-
[21]
Garc ´ ıa-Portugu´ es, P
E. Garc ´ ıa-Portugu´ es, P. Navarro-Esteban, and J. Cuesta-Albertos. A cram´ er–von mises test of uniformity on the hypersphere. In Statistical Learning and Modeling in Data Analysis: Methods and Applications 12 , pages 107–116. Springer, 2021
2021
-
[22]
Garc ´ ıa-Portugu´ es, P
E. Garc ´ ıa-Portugu´ es, P. Navarro-Esteban, and J. Cuesta-Albertos. On a projection-based class of uniformity tests on the hypersphere. Bernoulli, 29(1):181–204, 2023
2023
-
[23]
Hall and C
P. Hall and C. C. Heyde. Martingale Limit Theory and Its Application . Academic Press, New York., 1980
1980
-
[24]
Heiny and T
J. Heiny and T. Mikosch. Eigenvalues and eigenvectors of heavy-tailed sample covariance matrices with general growth rates: the iid case. Stochastic Processes and their Applications, 127(7):2179–2207, 2017
2017
-
[25]
Heiny and T
J. Heiny and T. Mikosch. The eigenstructure of the sample covariance matrices of high- dimensional stochastic volatility models with heavy tails. Bernoulli, 25(4B):3590–3622, 2019
2019
-
[26]
Heiny and T
J. Heiny and T. Mikosch. Large sample autocovariance matrices of linear processes with heavy tails. Stochastic Processes and their Applications , 141:344–375, 2021
2021
-
[27]
Heiny and J
J. Heiny and J. Yao. Limiting distributions for eigenvalues of sample correlation matrices from heavy-tailed populations. Ann. Stat., 50(6):3249–3280, 2022
2022
-
[28]
T. Jiang. A variance formula related to quantum conductance. Physics Letters A 373, 2117- 2121., 2009
2009
-
[29]
Jiang and T
T. Jiang and T. Pham. Asymptotic distributions of largest pearson correlation coefficients under dependent structures. arXiv preprint arXiv:2304.13102 , 2023
2023 arXiv
-
[30]
Juan and F
J. Juan and F. J. Prieto. Using angles to identify concentrated multivariate outliers. Tech- nometrics, 43:311–322, 2001
2001
-
[31]
Kulik and P
R. Kulik and P. Soulier. Heavy-Tailed Time Series. Springer New York, 2020. 39
2020
-
[32]
Ley and T
C. Ley and T. Verdebout. Modern Directional Statistics. CRC Press, Boca Raton., 2017
2017
-
[33]
K. V. Mardia and P. E. Jupp. Directional Statistics. John Wiley & Sons, Chichester., 2000
2000
-
[34]
Paindaveine and T
D. Paindaveine and T. Verdebout. On high-dimensional sign tests. Bernoulli, 22:1745–1769, 2016
2016
-
[35]
Paindaveine and T
D. Paindaveine and T. Verdebout. Detecting the direction of a signal on high-dimensional spheres: non-null and le cam optimality results. Probability Theory and Related Fields 176, 1165-1216., 2020
2020
-
[36]
Pewsey and E
A. Pewsey and E. Garc ´ ıa-Portugu´ es. Recent advances in directional statistics.Test, 30(1):1– 58, 2021
2021
-
[37]
E. G. Portugu´ es and T. Verdebout. An overview of uniformity tests on the hypersphere. arXiv:1804.00286., 2018
2018 arXiv
-
[38]
Self-normalized large deviations
Qi-Man Shao. Self-normalized large deviations. The Annals of Probability , 25(1):285–328, 1997
1997
-
[39]
Tang and B
Y. Tang and B. Li. A nonparametric test for elliptical distribution based on kernel embedding of probabilities. arXiv preprint arXiv:2306.10594 , 2023
2023 arXiv
-
[40]
L. Wang, B. Peng, and R. Li. A high-dimensional nonparametric multivariate test for mean vector. Journal of the American Statistical Association , 110:1658–1669, 2015
2015
-
[41]
X. Yu, D. Li, and L. Xue. Fisher’s combined probability test for high-dimensional covariance matrices. Journal of the American Statistical Association , 119:511–524, 2022
2022
-
[42]
X. Yu, D. Li, L. Xue, and R. Li. Power-enhanced simultaneous test of high-dimensional mean vectors and covariance matrices with application to gene-set testing. Journal of the American Statistical Association, 118:2548–2561, 2023
2023
-
[43]
X. Yu, J. Yao, and L. Xue. Power enhancement for testing multi-factor asset pricing models via fisher’s method. Journal of Econometrics , 239:105458, 2024
2024
-
[44]
C. Zou, L. Peng, L. Feng, and Z. Wang. Multivariate-sign-based high-dimensional tests for sphericity. Biometrika, 101:229–236, 2014. 40
2014
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.