Pith. sign in

REVIEW 3 major objections 4 minor 13 references

Exact detection threshold of the packing test

T0 review · 3 major / 4 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read The paper derives the exact detection thresholds and non-null limiting distributions of the packing test for spherical uniformity under FvML and Watson alternatives.

desk verdict Sharp non-null asymptotics for the packing test, with the Watson boundary regime sketched rather than proved. read the letter →

arxiv 2608.00445 v1 pith:EPVIHA3C submitted 2026-08-01 math.ST stat.TH

classification math.STstat.TH MSC 62H1562E2060G7060F10
keywords packingtestsphericaluniformitydetectionthresholdFisher-vonMises-LangevindistributionWatsonPoissonapproximationChen-Steinmethodphasetransition
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper pins down the exact detection thresholds and non-null limiting laws of the packing test for spherical uniformity, in the two standard high-dimensional parametric alternatives. Under Fisher--von Mises--Langevin alternatives, the threshold is $\kappa = \Theta(p^{3/4}/(\log n)^{1/4})$; under Watson alternatives, it is $p-2\kappa = \Theta(\sqrt{p\log n})$. Because these thresholds are strict, the packing test is suboptimal in sample size compared with the known minimax rates in both models. The paper also shows that, in the Watson model, the asymptotic scaling of the largest squared inner product undergoes a discontinuous second-order phase transition at the critical regime, while no analogous phenomenon occurs in the FvML model.

What carries the argument

The carrying mechanism is Poisson (Chen--Stein) approximation for the exceedance indicators $I_{ij} = \mathbf{1}\{|X_i^\top X_j|>t_n\}$, viewed through the tangent--normal decomposition of each observation. The mean of the approximating Poisson variable comes from sharp tail asymptotics of a single inner product, with a Laplace-method saddle-point analysis of the rate function; the variance is controlled by the dependency-graph term $b_2$, whose negligibility requires truncating to small neighborhoods of the saddle points and using uniform large-deviation bounds. The named objects that carry the argument are the rate-function minimizers $A(\rho)=\sqrt{2(1+\rho^2)}$, $a(\rho)=2(\sqrt{2(1+\rho^2)}-\rho)$, $K(\rho)$, and the critical kernel $K_{\rm cr}(\Delta)$; the change in saddle-point structure at $\rho=1$ is what produces the phase transition.

What would settle it

Simulate the Watson model at the critical scale $\Delta_n/\sqrt{p\log n}\to 1$ with large $p$ and $n$, and estimate $n^3\mathbb{P}(|X_1^\top X_2|>t_n,\,|X_1^\top X_3|>t_n)$: if this quantity fails to tend to zero, the Chen--Stein correlation bound in Sections 8.2--9.2 is false and Theorem 4's limiting coefficient is not $-1/2$.

Watch

Extended reading notes

Core claim

The central claim is that the packing test statistic $P_n = p\max_{i<j}(X_i^\top X_j)^2 - 4\log n + \log\log n$ has an exact, fully quantified non-null behaviour. Under FvML alternatives with $\kappa_n = \tau p^{3/4}/(\log n)^{1/4}$, $P_n$ converges to a shifted Gumbel law with shift $2\log\cosh(2\tau^2)$, giving the detection threshold $\Theta(p^{3/4}/(\log n)^{1/4})$. Under Watson alternatives with $\Delta_n = (p-2\kappa_n)/2$ and $\rho_n = \Delta_n/\sqrt{p\log n}\to\rho$, the limiting law of $P_n$ is Gumbel in all regimes, but the scaling of the largest squared inner product changes: for $\rho>1$ the null scaling persists and the threshold is $p-2\kappa = \Theta(\sqrt{p\log n})$; for $\rho=1$ the coefficient of $\log\log n/p$ jumps from $-1$ to $-1/2$; for $0<\rho<1$ both the leading and second-order coefficients depend on $\rho$. These results rigorously confirm the empirical observation that the packing test is strictly suboptimal in sample size for both models.

Load-bearing premise

The load-bearing premise is that the correlation term $b_2$, specifically the estimate $\mathbb{P}(I_{12}=1,I_{13}=1)=o(n^{-3})$ in the Watson analysis, is negligible after truncation to saddle-point neighborhoods; without that estimate the Poisson approximation and the four limiting laws do not go through.

Editorial extensions

If this is right

  • At the FvML minimax scale $\kappa\asymp p^{3/4}/\sqrt n$, the packing test has asymptotic power equal to its size, so it cannot detect at the optimal rate.
  • At its own thresholds, the packing test has explicit asymptotic power: $1-(1-\alpha)^{\cosh(2\tau^2)}$ under FvML and the piecewise formula (12) under Watson.
  • Under Watson alternatives with $\rho\le 1$, the packing test has asymptotic power one; for $\rho>1$, its power is strictly between size and one.
  • The largest squared inner product scales like $4\log n/p$ whenever $\rho\ge 1$, with the $\log\log n$ coefficient jumping from $-1$ to $-1/2$ at $\rho=1$, and acquires $\rho$-dependent coefficients for $\rho<1$.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A consequence the paper leaves implicit is that at fixed concentration the FvML threshold forces an exponentially large sample: $\log n \sim p^3/\kappa^4$ is necessary for detection.
  • The critical scale restriction $\sqrt{\log n}\,(\rho_n-1)\to\Delta$ suggests an uncovered family of intermediate regimes with $\rho_n-1$ of order $(\log n)^{-\gamma}$ for $\gamma\ne 1/2$, where the second-order coefficient may interpolate between $-1$ and $-1/2$.
  • A direct way to test the phase transition in simulations is to fit the slope of $pM_n^2$ versus $\log\log n$ at $\rho_n\approx 1$; Theorem 4 predicts the slope $-1/2$, whereas both adjacent regimes give $-1$.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper studies the packing test for uniformity on the high-dimensional sphere under Fisher--von Mises--Langevin (FvML) and Watson alternatives. Using Poisson approximation, the author derives non-null limiting distributions for the largest squared inner product. Theorem 1 gives a shifted Gumbel limit under FvML alternatives with concentration κ=τ p^{3/4}/(log n)^{1/4}. Theorems 2--4 give Gumbel limits under Watson alternatives in three regimes: local (ρ>1), non-local (ρ∈(0,1)), and critical (ρ=1), where ρ is the limiting value of Δ_n/√(p log n) with Δ_n=(p−2κ_n)/2. The paper concludes that the FvML detection threshold is Θ(p^{3/4}/(log n)^{1/4}) and the Watson threshold is p−2κ=Θ(√(p log n)), with a discontinuous second-order phase transition at ρ=1. Proofs are built on tangent--normal decompositions, Chen--Stein Poisson approximation, Laplace's method, and large-deviation estimates.

Significance. If the boundary gaps are closed, the paper would provide the first exact non-null extreme-value analysis of the packing test under two standard high-dimensional directional models. The results would rigorously confirm the empirical suboptimality of the packing test and reveal a genuinely novel second-order phase transition in the Watson model. The mathematical work is substantial: the tail asymptotics in Propositions 3, 5, 6, and 7 are derived with explicit error rates; the Chen--Stein b2 term is addressed via truncation and saddle-point analysis; and the limiting constants are explicit and parameter-free. The paper is not circular: prior work is cited only for context and for comparison rates, not as an input to the proofs. However, as detailed below, the exact-threshold claims in the abstract and Section 5 go beyond what the stated theorems and the sketched Appendix B actually prove.

major comments (3)
  1. [§4.1, Eq. (12)] The exact-threshold claim 'p−2κ=Θ(√(p log n))' is not established by the theorems as stated. Theorem 2 covers only ρ_n→ρ∈(1,∞), Theorem 3 only ρ_n→ρ∈(0,1), and Theorem 4 only √(log n)(ρ_n−1)→Δ∈R. The inference from Theorem 2 that the limiting power tends to α as ρ→∞ is a statement about fixed-ρ limits, not about sequences ρ_n→∞; likewise, Theorem 3 gives no control for ρ_n→0. Since the Θ notation requires both Δ_n/√(p log n)→∞ (no detection) and Δ_n/√(p log n)→0 (detection with probability tending to one), the stated threshold is stronger than the proved convergence results. Please either add proofs for these boundary regimes or rephrase the abstract and Section 5 to state the threshold only as a consequence of Theorems 2--4 for the regimes they cover.
  2. [Appendix B] The proof of power one for ρ_n→1 at arbitrary rates is not complete. The assertion 'Step 2 of the proof of Theorem 4 still yields n^3 P(I12=I13=1)→0' is made without proof, but Step 2 in Section 10.2 relies on the critical-scale change of variables y=L^{1/4}(α1+α2)/√2, z=√(L/2)(α1−α2) and on the limit √(L)(ρ_n−1)→Δ. When √(L)(ρ_n−1)→±∞ the saddle structure and the associated rate (62) are different, so the b2 bound cannot be taken as a black box. Because Eq. (12) and the threshold claim use power one at ρ=1, this missing correlation estimate is load-bearing; if the estimate fails, the Poisson approximation at the boundary would break down and the phase-transition conclusion would not follow.
  3. [§4.2, Theorem 3] The power formula (12) is stated for all ρ∈[0,∞), but the proof of the ρ∈[0,1) part relies on Theorem 3 with ρ∈(0,1) and on Appendix B for ρ=1; the endpoint ρ=0 is not covered by either. If the intended definition of detection threshold does not require the ρ_n→0 case, please state this explicitly; otherwise supply a proof for Δ_n=o(√(p log n)). The same issue affects the interpretation of the non-local regime as 'asymptotically powerful throughout this regime', since ρ_n→0 is part of that regime by the notation Δ_n/√(p log n)→ρ∈[0,1).
minor comments (4)
  1. [Appendix A, Lemma 6] The proof omits the minimization ('straightforward minimization of the quadratic function'). Since Lemma 6 supplies the constant (3+c)/2 used in the b2 estimate of Theorem 4, please include the explicit computation.
  2. [Abstract and §5] The abstract and Section 5 state the Watson threshold without the technical conditions p/(log n)^2→∞ (Theorems 1--2) and p/(log n)^3→∞ (Theorems 3--4); these qualifications should appear wherever the threshold is asserted.
  3. [§2, Notation] The symbol L is defined both as log n and as a generic sequence diverging to infinity; this dual use is confusing in Propositions 5--7 and should be disambiguated.
  4. [After Theorem 1 and Eq. (12)] In the displayed power formulas, the exponent on (1−α) is not visible in the manuscript rendering; please verify that β_FvML=1−(1−α)^{cosh(2τ^2)} and β_Wat=1−(1−α)^{(1−ρ^{-2})^{-1/2}} appear correctly in the final version.

Circularity Check

0 steps flagged · score 2.0 of 10

No significant circularity: the detection thresholds are derived from the model distributions by Poisson approximation, and the self-citations are contextual or comparative, not proof inputs.

full rationale

The paper's central claims (Theorems 1-4) are obtained by direct asymptotic analysis of the inner-product tail probabilities under the stated FvML and Watson densities: Propositions 3 and 5 compute the one-pair exceedance probabilities, Propositions 6 and 7 apply Laplace's method and moderate-deviation bounds, and Lemma 5 (Arratia et al.) supplies the Poisson approximation. The threshold statements follow from the computed means and b2-correlation estimates, not from a fitted parameter or from a prior result assumed as input. The citations to Jiang and Pham [2025a,b] are used only for context and for the minimax comparison rates in Sections 2 and 4.1; the theorems themselves do not invoke those papers. There is no redefinition of the object being predicted, no ansatz smuggled in by citation, and no uniqueness theorem imported from the author's prior work. The manuscript does flag a genuine gap in Appendix B: the power-one boundary case rho_n -> 1 is finished by asserting that 'Step 2 of the proof of Theorem 4 still yields n^3 P(I12=I13=1) -> 0' without proof. That is a completeness/rigor gap, and the paper also warns that the critical regime requires technical conditions on p/(log n)^3 and that boundary uniformity is not fully proved. These are limitations about rigor, not circularity: an unverified estimate is different from an estimate that is an input repackaged as an output. Overall, the self-citations are minor and non-load-bearing, giving score 2.

Assumptions & free parameters 0 free parameters · 6 assumptions · 0 invented entities

The paper is mathematically self-contained given standard probability tools. No free parameters are fitted to data; the limiting parameters tau, rho, and Delta in the theorems parametrize the local alternatives and are not chosen to make the derivation work. No new entities are postulated.

assumptions (6)
  • standard math Poisson approximation / Chen-Stein bound (Arratia et al. 1989), used as Lemma 5.
    The core tool for asymptotic independence of exceedance indicators; cited as an established theorem.
  • standard math Brascamp-Lieb inequality (Nguyen 2014), used in Proposition 1 for variance and sub-Gaussian bounds.
    Standard concentration tool invoked with reference.
  • standard math Self-normalized moderate-deviation result (Theorem 2.3 in Jing et al. 2003), used in Propositions 6 and 7.
    Provides the tail asymptotic for the uniform component of the inner product.
  • domain assumption FvML model (5) and Watson model (6) with independent observations.
    The paper's results are stated for these two parametric alternatives; the models are standard in directional statistics.
  • domain assumption High-dimensional scaling conditions p/(log n)^2 -> infinity (Theorems 1-2) and p/(log n)^3 -> infinity (Theorems 3-4).
    Needed for the Laplace method and uniform tail bounds; the paper honestly notes that p/(log n)^3 may be improvable.
  • domain assumption Tangent-normal decomposition representations (15) and (39) with independence between the cosine component T and the uniform remainder V.
    Standard construction for rotationally symmetric distributions on the sphere; used throughout the proofs.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Exact detection threshold of the packing test." pith.science (2026). https://pith.science/paper/EPVIHA3C

@misc{pith2026260800445,
  author       = {Pith},
  title        = {Pith review of: Exact detection threshold of the packing test},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/EPVIHA3C}},
  note         = {Machine review of arXiv:2608.00445}
}
abstract

Using Poisson approximation techniques, we derive the detection threshold of the packing test in \cite{Jiang13} when testing spherical uniformity under high-dimensional Fisher--von Mises--Langevin (FvML) and Watson alternatives. Our result rigorously confirms the empirical observation that the packing test is strictly suboptimal for testing uniformity in these two popular models. In the high-dimensional FvML model, its detection threshold is precisely \(\kappa=\Theta\lb p^{3/4}/(\log n)^{1/4}\rb\). In the high-dimensional Watson model, its detection threshold is \(p-2\kappa=\Theta(\sqrt{p\log n})\), or equivalently \(\kappa=p/2-\Theta(\sqrt{p\log n})\). The non-null limiting distributions of the packing test under these two models are derived. We show that the limiting scalings of the largest squared inner product undergo a discontinuous phase transition in the Watson model, whereas no analogous phenomenon occurs in the FvML model.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

13 extracted references · 9 canonical work pages

  1. [1]

    Hereκ∈R + is the concentration parameter,µis the unknown location (nuisance) parameter, andc FvML p,κ ,c Wat p,κ are normalizing constants

    and x7−→c Wat p,κ exp κ(x⊤µ)2 ,x∈S p−1,(6) respectively. Hereκ∈R + is the concentration parameter,µis the unknown location (nuisance) parameter, andc FvML p,κ ,c Wat p,κ are normalizing constants. The optimal detection rates in these two models have been studied in Cutting et al. [2017, 2022]; Jiang and Pham [2025b]. It was shown in Cutting et al

  2. [2]

    This scaling is also the same as under the null (uniformity)

    Although both limiting distributions are of Gumbel type, their scalings are markedly different: Theorem 2 implies that M2 n = 4 logn p − log logn p +o P log logn p ! In particular, the first- and second-order approximations ofM 2 n are independent of the value ofρas long asρ >1. This scaling is also the same as under the null (uniformity). In contrast, Th...

  3. [6]

    Indeed, conditionally on eT1, eT2, the event n X⊤ 1 X2 >t◦ n o is equivalent to √ d·Z> √ d· t◦ n− eT1eT2 q 1− eT 2 1 q 1− eT 2 2

    := ρn π · L (an−α 1α2) √ 2πL ×exp ( −L " ρn(α2 1 +α 2 2)+ α4 1 +α 4 2 4 + 1 2(an−α 1α2)2 #) ×L (an−α1α2)/(2An)·exp ( −an−α 1α2 2an x ) . Indeed, conditionally on eT1, eT2, the event n X⊤ 1 X2 >t◦ n o is equivalent to √ d·Z> √ d· t◦ n− eT1eT2 q 1− eT 2 1 q 1− eT 2 2 . A direct computation gives p s◦ n =a n √ L− an An logL−x 2an √ L +o(L−1/2). Moreover, uni...

  4. [7]

    u2 2 +v 2 + u4 8 + v4 4 # + 1 2 (2−uv−ε )2 + ≥(1−3ε)

    Define L:=logn,I ◦◦ i j :=1{|S i j|>t◦◦n},W ◦◦ := X 1≤i<j≤n I◦◦ i j. Then ( pM 2 n≤4 logn− 1 2 log logn+x ) ={M n≤t◦◦ n}={W ◦◦ =0}. LetI n :={(i,j) : 1≤i<j≤n}.By Lemma 5, Pκn,e1(W◦◦ =0)−exp{−λ ◦◦ n} ≤b◦◦ 1n +b◦◦ 2n, where λ◦◦ n :=E κn,e1W◦◦ = n 2 ! Pκn,e1(I◦◦ 12 =1), and b◦◦ 1n≲n 3Pκn,e1(I◦◦ 12 =1) 2,b ◦◦ 2n≲n 3Pκn,e1 I◦◦ 12 =1,I ◦◦ 13 =1 . As in the proo...

  5. [8]

    √ 2q+ q2 2 √ 2 # + 1 2 (2−q−ε )2 + ) > 3+c 2 . Proof of Lemma 6.Fixε∈[0,1/10]. Consider the caseq≥2−ε. Then inf q≥0 ( (1−2ε )

    for a proof. Lemma 5(Arratia et al. [1989]).Let{I α}α∈I be Bernoulli variables with a dependency graph: I α is independent of{I β :β<B α}. Let W= P α∈I Iα andλ=EW. Define b1 := X α∈I X β∈Bα EIα EIβ,b 2 := X α∈I X β∈Bα β,α E(IαIβ). Then dTV (L(W),Poisson(λ) )≤min(1,λ −1)(b1 +b 2). Lemma 6.There exists a universal constant c>0such that inf ε∈[0, 1 10],q≥0 (...

  6. [1989]

    [2026]; Jammalamadaka and Janson

    (see also Feng et al. [2026]; Jammalamadaka and Janson

  7. [2003]

    Understanding dimensional collapse in contrastive self-supervised learning.arXiv preprint arXiv:2110.09348,

    Li Jing, Pascal Vincent, Yann LeCun, and Yuandong Tian. Understanding dimensional collapse in contrastive self-supervised learning.arXiv preprint arXiv:2110.09348,

  8. [2014]

    This proves the second claim in (17)

    and the references therein], we obtain Var(T)≤1/(p−3). This proves the second claim in (17). The first claim in (17) follows, under the standing assumptionκ=o(p), from the proof of Lemma 4 in Paindaveine and Verdebout [2020]. We now prove the sub-Gaussian bound (18) using the Herbst argument. Forλ∈R, define the tilted lawν λ by dνλ(z) := eλz EeλT dν(z). H...

Show all 13 references
  1. [2015]

    Asymptotic analysis of high-dimensional uniformity tests under heavy-tailed alternatives.arXiv preprint arXiv:2506.00393, 2025a

    Tiefeng Jiang and Tuan Pham. Asymptotic analysis of high-dimensional uniformity tests under heavy-tailed alternatives.arXiv preprint arXiv:2506.00393, 2025a. Tiefeng Jiang and Tuan Pham. Detecting non-uniform patterns on high-dimensional hyperspheres. arXiv preprint arXiv:2506...

  2. [2018]

    Learning with hyperspherical uniformity

    Weiyang Liu, Rongmei Lin, Zhen Liu, Li Xiong, Bernhard Sch ¨olkopf, and Adrian Weller. Learning with hyperspherical uniformity. In24th International Conference on Artificial Intelligence and Statistics (AISTATS 2021), pages 1180–1188. PMLR,

  3. [2021]

    E. G. Portugu ´es and T. Verdebout. An overview of uniformity tests on the hypersphere. arXiv:1804.00286.,

  4. [2022]

    High-dimensional sobolev tests on hyperspheres.arXiv preprint arXiv:2501.10898,

    Bruno Ebner, Eduardo Garc´ıa-Portugu´es, and Thomas Verdebout. High-dimensional sobolev tests on hyperspheres.arXiv preprint arXiv:2501.10898,

  5. [2025]

    In particular, any linear combination ofR n andB n in (2) and (3), respectively, that is not proportional toR n fails to achieve the optimal detection rate

    further establish a somewhat sur- prising result: within a general class of Sobolev-based tests that includes the Rayleigh and Bingham tests as special cases, the Rayleigh test is the only one that achieves the optimal detection rate under the FvML model. In particular, any li...

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.