REVIEW 3 major objections 4 minor 13 references
Exact detection threshold of the packing test
T0 review · 3 major / 4 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read The paper derives the exact detection thresholds and non-null limiting distributions of the packing test for spherical uniformity under FvML and Watson alternatives.
desk verdict Sharp non-null asymptotics for the packing test, with the Watson boundary regime sketched rather than proved. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying mechanism is Poisson (Chen--Stein) approximation for the exceedance indicators $I_{ij} = \mathbf{1}\{|X_i^\top X_j|>t_n\}$, viewed through the tangent--normal decomposition of each observation. The mean of the approximating Poisson variable comes from sharp tail asymptotics of a single inner product, with a Laplace-method saddle-point analysis of the rate function; the variance is controlled by the dependency-graph term $b_2$, whose negligibility requires truncating to small neighborhoods of the saddle points and using uniform large-deviation bounds. The named objects that carry the argument are the rate-function minimizers $A(\rho)=\sqrt{2(1+\rho^2)}$, $a(\rho)=2(\sqrt{2(1+\rho^2)}-\rho)$, $K(\rho)$, and the critical kernel $K_{\rm cr}(\Delta)$; the change in saddle-point structure at $\rho=1$ is what produces the phase transition.
What would settle it
Simulate the Watson model at the critical scale $\Delta_n/\sqrt{p\log n}\to 1$ with large $p$ and $n$, and estimate $n^3\mathbb{P}(|X_1^\top X_2|>t_n,\,|X_1^\top X_3|>t_n)$: if this quantity fails to tend to zero, the Chen--Stein correlation bound in Sections 8.2--9.2 is false and Theorem 4's limiting coefficient is not $-1/2$.
Extended reading notes
Core claim
The central claim is that the packing test statistic $P_n = p\max_{i<j}(X_i^\top X_j)^2 - 4\log n + \log\log n$ has an exact, fully quantified non-null behaviour. Under FvML alternatives with $\kappa_n = \tau p^{3/4}/(\log n)^{1/4}$, $P_n$ converges to a shifted Gumbel law with shift $2\log\cosh(2\tau^2)$, giving the detection threshold $\Theta(p^{3/4}/(\log n)^{1/4})$. Under Watson alternatives with $\Delta_n = (p-2\kappa_n)/2$ and $\rho_n = \Delta_n/\sqrt{p\log n}\to\rho$, the limiting law of $P_n$ is Gumbel in all regimes, but the scaling of the largest squared inner product changes: for $\rho>1$ the null scaling persists and the threshold is $p-2\kappa = \Theta(\sqrt{p\log n})$; for $\rho=1$ the coefficient of $\log\log n/p$ jumps from $-1$ to $-1/2$; for $0<\rho<1$ both the leading and second-order coefficients depend on $\rho$. These results rigorously confirm the empirical observation that the packing test is strictly suboptimal in sample size for both models.
Load-bearing premise
The load-bearing premise is that the correlation term $b_2$, specifically the estimate $\mathbb{P}(I_{12}=1,I_{13}=1)=o(n^{-3})$ in the Watson analysis, is negligible after truncation to saddle-point neighborhoods; without that estimate the Poisson approximation and the four limiting laws do not go through.
Editorial extensions
If this is right
- At the FvML minimax scale $\kappa\asymp p^{3/4}/\sqrt n$, the packing test has asymptotic power equal to its size, so it cannot detect at the optimal rate.
- At its own thresholds, the packing test has explicit asymptotic power: $1-(1-\alpha)^{\cosh(2\tau^2)}$ under FvML and the piecewise formula (12) under Watson.
- Under Watson alternatives with $\rho\le 1$, the packing test has asymptotic power one; for $\rho>1$, its power is strictly between size and one.
- The largest squared inner product scales like $4\log n/p$ whenever $\rho\ge 1$, with the $\log\log n$ coefficient jumping from $-1$ to $-1/2$ at $\rho=1$, and acquires $\rho$-dependent coefficients for $\rho<1$.
Reading between the lines
- A consequence the paper leaves implicit is that at fixed concentration the FvML threshold forces an exponentially large sample: $\log n \sim p^3/\kappa^4$ is necessary for detection.
- The critical scale restriction $\sqrt{\log n}\,(\rho_n-1)\to\Delta$ suggests an uncovered family of intermediate regimes with $\rho_n-1$ of order $(\log n)^{-\gamma}$ for $\gamma\ne 1/2$, where the second-order coefficient may interpolate between $-1$ and $-1/2$.
- A direct way to test the phase transition in simulations is to fit the slope of $pM_n^2$ versus $\log\log n$ at $\rho_n\approx 1$; Theorem 4 predicts the slope $-1/2$, whereas both adjacent regimes give $-1$.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies the packing test for uniformity on the high-dimensional sphere under Fisher--von Mises--Langevin (FvML) and Watson alternatives. Using Poisson approximation, the author derives non-null limiting distributions for the largest squared inner product. Theorem 1 gives a shifted Gumbel limit under FvML alternatives with concentration κ=τ p^{3/4}/(log n)^{1/4}. Theorems 2--4 give Gumbel limits under Watson alternatives in three regimes: local (ρ>1), non-local (ρ∈(0,1)), and critical (ρ=1), where ρ is the limiting value of Δ_n/√(p log n) with Δ_n=(p−2κ_n)/2. The paper concludes that the FvML detection threshold is Θ(p^{3/4}/(log n)^{1/4}) and the Watson threshold is p−2κ=Θ(√(p log n)), with a discontinuous second-order phase transition at ρ=1. Proofs are built on tangent--normal decompositions, Chen--Stein Poisson approximation, Laplace's method, and large-deviation estimates.
Significance. If the boundary gaps are closed, the paper would provide the first exact non-null extreme-value analysis of the packing test under two standard high-dimensional directional models. The results would rigorously confirm the empirical suboptimality of the packing test and reveal a genuinely novel second-order phase transition in the Watson model. The mathematical work is substantial: the tail asymptotics in Propositions 3, 5, 6, and 7 are derived with explicit error rates; the Chen--Stein b2 term is addressed via truncation and saddle-point analysis; and the limiting constants are explicit and parameter-free. The paper is not circular: prior work is cited only for context and for comparison rates, not as an input to the proofs. However, as detailed below, the exact-threshold claims in the abstract and Section 5 go beyond what the stated theorems and the sketched Appendix B actually prove.
major comments (3)
- [§4.1, Eq. (12)] The exact-threshold claim 'p−2κ=Θ(√(p log n))' is not established by the theorems as stated. Theorem 2 covers only ρ_n→ρ∈(1,∞), Theorem 3 only ρ_n→ρ∈(0,1), and Theorem 4 only √(log n)(ρ_n−1)→Δ∈R. The inference from Theorem 2 that the limiting power tends to α as ρ→∞ is a statement about fixed-ρ limits, not about sequences ρ_n→∞; likewise, Theorem 3 gives no control for ρ_n→0. Since the Θ notation requires both Δ_n/√(p log n)→∞ (no detection) and Δ_n/√(p log n)→0 (detection with probability tending to one), the stated threshold is stronger than the proved convergence results. Please either add proofs for these boundary regimes or rephrase the abstract and Section 5 to state the threshold only as a consequence of Theorems 2--4 for the regimes they cover.
- [Appendix B] The proof of power one for ρ_n→1 at arbitrary rates is not complete. The assertion 'Step 2 of the proof of Theorem 4 still yields n^3 P(I12=I13=1)→0' is made without proof, but Step 2 in Section 10.2 relies on the critical-scale change of variables y=L^{1/4}(α1+α2)/√2, z=√(L/2)(α1−α2) and on the limit √(L)(ρ_n−1)→Δ. When √(L)(ρ_n−1)→±∞ the saddle structure and the associated rate (62) are different, so the b2 bound cannot be taken as a black box. Because Eq. (12) and the threshold claim use power one at ρ=1, this missing correlation estimate is load-bearing; if the estimate fails, the Poisson approximation at the boundary would break down and the phase-transition conclusion would not follow.
- [§4.2, Theorem 3] The power formula (12) is stated for all ρ∈[0,∞), but the proof of the ρ∈[0,1) part relies on Theorem 3 with ρ∈(0,1) and on Appendix B for ρ=1; the endpoint ρ=0 is not covered by either. If the intended definition of detection threshold does not require the ρ_n→0 case, please state this explicitly; otherwise supply a proof for Δ_n=o(√(p log n)). The same issue affects the interpretation of the non-local regime as 'asymptotically powerful throughout this regime', since ρ_n→0 is part of that regime by the notation Δ_n/√(p log n)→ρ∈[0,1).
minor comments (4)
- [Appendix A, Lemma 6] The proof omits the minimization ('straightforward minimization of the quadratic function'). Since Lemma 6 supplies the constant (3+c)/2 used in the b2 estimate of Theorem 4, please include the explicit computation.
- [Abstract and §5] The abstract and Section 5 state the Watson threshold without the technical conditions p/(log n)^2→∞ (Theorems 1--2) and p/(log n)^3→∞ (Theorems 3--4); these qualifications should appear wherever the threshold is asserted.
- [§2, Notation] The symbol L is defined both as log n and as a generic sequence diverging to infinity; this dual use is confusing in Propositions 5--7 and should be disambiguated.
- [After Theorem 1 and Eq. (12)] In the displayed power formulas, the exponent on (1−α) is not visible in the manuscript rendering; please verify that β_FvML=1−(1−α)^{cosh(2τ^2)} and β_Wat=1−(1−α)^{(1−ρ^{-2})^{-1/2}} appear correctly in the final version.
Circularity Check
No significant circularity: the detection thresholds are derived from the model distributions by Poisson approximation, and the self-citations are contextual or comparative, not proof inputs.
full rationale
The paper's central claims (Theorems 1-4) are obtained by direct asymptotic analysis of the inner-product tail probabilities under the stated FvML and Watson densities: Propositions 3 and 5 compute the one-pair exceedance probabilities, Propositions 6 and 7 apply Laplace's method and moderate-deviation bounds, and Lemma 5 (Arratia et al.) supplies the Poisson approximation. The threshold statements follow from the computed means and b2-correlation estimates, not from a fitted parameter or from a prior result assumed as input. The citations to Jiang and Pham [2025a,b] are used only for context and for the minimax comparison rates in Sections 2 and 4.1; the theorems themselves do not invoke those papers. There is no redefinition of the object being predicted, no ansatz smuggled in by citation, and no uniqueness theorem imported from the author's prior work. The manuscript does flag a genuine gap in Appendix B: the power-one boundary case rho_n -> 1 is finished by asserting that 'Step 2 of the proof of Theorem 4 still yields n^3 P(I12=I13=1) -> 0' without proof. That is a completeness/rigor gap, and the paper also warns that the critical regime requires technical conditions on p/(log n)^3 and that boundary uniformity is not fully proved. These are limitations about rigor, not circularity: an unverified estimate is different from an estimate that is an input repackaged as an output. Overall, the self-citations are minor and non-load-bearing, giving score 2.
Assumptions & free parameters
assumptions (6)
- standard math Poisson approximation / Chen-Stein bound (Arratia et al. 1989), used as Lemma 5.
- standard math Brascamp-Lieb inequality (Nguyen 2014), used in Proposition 1 for variance and sub-Gaussian bounds.
- standard math Self-normalized moderate-deviation result (Theorem 2.3 in Jing et al. 2003), used in Propositions 6 and 7.
- domain assumption FvML model (5) and Watson model (6) with independent observations.
- domain assumption High-dimensional scaling conditions p/(log n)^2 -> infinity (Theorems 1-2) and p/(log n)^3 -> infinity (Theorems 3-4).
- domain assumption Tangent-normal decomposition representations (15) and (39) with independence between the cosine component T and the uniform remainder V.
Cite this review
Pith. "Pith review of Exact detection threshold of the packing test." pith.science (2026). https://pith.science/paper/EPVIHA3C
@misc{pith2026260800445,
author = {Pith},
title = {Pith review of: Exact detection threshold of the packing test},
year = {2026},
howpublished = {\url{https://pith.science/paper/EPVIHA3C}},
note = {Machine review of arXiv:2608.00445}
}
abstract
Using Poisson approximation techniques, we derive the detection threshold of the packing test in \cite{Jiang13} when testing spherical uniformity under high-dimensional Fisher--von Mises--Langevin (FvML) and Watson alternatives. Our result rigorously confirms the empirical observation that the packing test is strictly suboptimal for testing uniformity in these two popular models. In the high-dimensional FvML model, its detection threshold is precisely \(\kappa=\Theta\lb p^{3/4}/(\log n)^{1/4}\rb\). In the high-dimensional Watson model, its detection threshold is \(p-2\kappa=\Theta(\sqrt{p\log n})\), or equivalently \(\kappa=p/2-\Theta(\sqrt{p\log n})\). The non-null limiting distributions of the packing test under these two models are derived. We show that the limiting scalings of the largest squared inner product undergo a discontinuous phase transition in the Watson model, whereas no analogous phenomenon occurs in the FvML model.
Reference graph
Works this paper leans on
-
[1]
and x7−→c Wat p,κ exp κ(x⊤µ)2 ,x∈S p−1,(6) respectively. Hereκ∈R + is the concentration parameter,µis the unknown location (nuisance) parameter, andc FvML p,κ ,c Wat p,κ are normalizing constants. The optimal detection rates in these two models have been studied in Cutting et al. [2017, 2022]; Jiang and Pham [2025b]. It was shown in Cutting et al
work page 2017
-
[2]
This scaling is also the same as under the null (uniformity)
Although both limiting distributions are of Gumbel type, their scalings are markedly different: Theorem 2 implies that M2 n = 4 logn p − log logn p +o P log logn p ! In particular, the first- and second-order approximations ofM 2 n are independent of the value ofρas long asρ >1. This scaling is also the same as under the null (uniformity). In contrast, Th...
work page 2017
-
[6]
:= ρn π · L (an−α 1α2) √ 2πL ×exp ( −L " ρn(α2 1 +α 2 2)+ α4 1 +α 4 2 4 + 1 2(an−α 1α2)2 #) ×L (an−α1α2)/(2An)·exp ( −an−α 1α2 2an x ) . Indeed, conditionally on eT1, eT2, the event n X⊤ 1 X2 >t◦ n o is equivalent to √ d·Z> √ d· t◦ n− eT1eT2 q 1− eT 2 1 q 1− eT 2 2 . A direct computation gives p s◦ n =a n √ L− an An logL−x 2an √ L +o(L−1/2). Moreover, uni...
work page 2003
-
[7]
u2 2 +v 2 + u4 8 + v4 4 # + 1 2 (2−uv−ε )2 + ≥(1−3ε)
Define L:=logn,I ◦◦ i j :=1{|S i j|>t◦◦n},W ◦◦ := X 1≤i<j≤n I◦◦ i j. Then ( pM 2 n≤4 logn− 1 2 log logn+x ) ={M n≤t◦◦ n}={W ◦◦ =0}. LetI n :={(i,j) : 1≤i<j≤n}.By Lemma 5, Pκn,e1(W◦◦ =0)−exp{−λ ◦◦ n} ≤b◦◦ 1n +b◦◦ 2n, where λ◦◦ n :=E κn,e1W◦◦ = n 2 ! Pκn,e1(I◦◦ 12 =1), and b◦◦ 1n≲n 3Pκn,e1(I◦◦ 12 =1) 2,b ◦◦ 2n≲n 3Pκn,e1 I◦◦ 12 =1,I ◦◦ 13 =1 . As in the proo...
work page 2003
-
[8]
for a proof. Lemma 5(Arratia et al. [1989]).Let{I α}α∈I be Bernoulli variables with a dependency graph: I α is independent of{I β :β<B α}. Let W= P α∈I Iα andλ=EW. Define b1 := X α∈I X β∈Bα EIα EIβ,b 2 := X α∈I X β∈Bα β,α E(IαIβ). Then dTV (L(W),Poisson(λ) )≤min(1,λ −1)(b1 +b 2). Lemma 6.There exists a universal constant c>0such that inf ε∈[0, 1 10],q≥0 (...
work page 1989
-
[1989]
[2026]; Jammalamadaka and Janson
(see also Feng et al. [2026]; Jammalamadaka and Janson
work page 2026
-
[2003]
Li Jing, Pascal Vincent, Yann LeCun, and Yuandong Tian. Understanding dimensional collapse in contrastive self-supervised learning.arXiv preprint arXiv:2110.09348,
-
[2014]
This proves the second claim in (17)
and the references therein], we obtain Var(T)≤1/(p−3). This proves the second claim in (17). The first claim in (17) follows, under the standing assumptionκ=o(p), from the proof of Lemma 4 in Paindaveine and Verdebout [2020]. We now prove the sub-Gaussian bound (18) using the Herbst argument. Forλ∈R, define the tilted lawν λ by dνλ(z) := eλz EeλT dν(z). H...
work page 2020
Show all 13 references
-
[2015]
Asymptotic analysis of high-dimensional uniformity tests under heavy-tailed alternatives.arXiv preprint arXiv:2506.00393, 2025a
Tiefeng Jiang and Tuan Pham. Asymptotic analysis of high-dimensional uniformity tests under heavy-tailed alternatives.arXiv preprint arXiv:2506.00393, 2025a. Tiefeng Jiang and Tuan Pham. Detecting non-uniform patterns on high-dimensional hyperspheres. arXiv preprint arXiv:2506...
-
[2018]
Learning with hyperspherical uniformity
Weiyang Liu, Rongmei Lin, Zhen Liu, Li Xiong, Bernhard Sch ¨olkopf, and Adrian Weller. Learning with hyperspherical uniformity. In24th International Conference on Artificial Intelligence and Statistics (AISTATS 2021), pages 1180–1188. PMLR,
2021
-
[2021]
E. G. Portugu ´es and T. Verdebout. An overview of uniformity tests on the hypersphere. arXiv:1804.00286.,
-
[2022]
High-dimensional sobolev tests on hyperspheres.arXiv preprint arXiv:2501.10898,
Bruno Ebner, Eduardo Garc´ıa-Portugu´es, and Thomas Verdebout. High-dimensional sobolev tests on hyperspheres.arXiv preprint arXiv:2501.10898,
-
[2025]
In particular, any linear combination ofR n andB n in (2) and (3), respectively, that is not proportional toR n fails to achieve the optimal detection rate
further establish a somewhat sur- prising result: within a general class of Sobolev-based tests that includes the Rayleigh and Bingham tests as special cases, the Rayleigh test is the only one that achieves the optimal detection rate under the FvML model. In particular, any li...
2017
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.