REVIEW 3 major objections 6 minor 22 references
First-order Edgeworth expansions for linear rank statistics proved via Stein's method
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · glm-5.2
2026-07-07 19:46 UTC pith:H43QDT6U
load-bearing objection Consolidated republication of the 1989 AOS paper with a corrected distribution definition; first-order Edgeworth expansion proofs are sound but depend on a condition only verified for product-form matrices. the 3 major comments →
Edgeworth Expansions for Linear Rank Statistics -- Consolidated Version
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central mechanism is a combinatorial coupling that creates a chain of four permutations π_1 → π_2 → π_3 → π_4, each obtained from the previous by composing with a fixed permutation determined by a random index vector I. The successive increments ΔT_1, ΔT_2, ΔT_3 decompose T_A into pieces with controlled dependence: (I_1, J_1) is independent of π_3, (I_1, I_2, J_1, J_2) is independent of π_2, and (I_1,...,I_4, J_1,...,J_4) is independent of π_1. This layered independence suffices to turn Stein's identity E(q(T_A)) - Φ(q) = E(f'(T_A)) - E(T_A f(T_A)) into a Taylor-expansion argument where each remainder term involves only lower-order dependence, yielding the basic equation (3.20) from Thee
What carries the argument
The basic equation (3.20): E(q(T_A)) - Φ(q) + (1/2)E(T_A³)·E(T_A f'(T_A)) = R(q), where R(q) is a remainder expressible in terms of second derivatives of the Stein solution f and the increments ΔT_k. The cubic moment E(T_A³)/n captures the first-order Edgeworth correction λ_{1,A}, and bounding R(q) by D_A² gives the O(n^{-1}) rate.
Load-bearing premise
The main theorem for distribution functions (Theorem 2.4) requires condition (2.5): that the second differences of the distribution functions of all 8-row-and-column-deleted submatrices of the centered matrix  are bounded by C₁(D_A² + y²). This condition is natural and nearly necessary, but the paper only verifies it for the special case a_{ij} = e_i d_j (via van Zwet's conditions); for general matrices A it remains an abstract assumption whose practical verifiability is unc
What would settle it
If a matrix A with σ_A > 0 could be constructed where condition (2.5) fails — i.e., where second differences of F_B grow faster than D_A² + y² for some submatrix B ∈ N(8, Â) — and the Edgeworth expansion error ||F_A - e_{1,A}|| were simultaneously observed to exceed K·D_A² for all constants K, this would show the condition is not merely sufficient but necessary.
If this is right
- The method provides a template for Edgeworth expansions in other combinatorial probability settings where full independence is absent but partial independence can be manufactured through coupling constructions.
- The second-order expansions (Theorem 2.7), while stated without proof here, suggest that the same combinatorial framework extends to higher-order corrections if one extends the permutation chain to length 5 and uses 16-element index vectors.
- The connection to van Zwet's conditions shows that the abstract second-difference condition (2.5) is satisfiable under standard moment and anti-concentration assumptions on score functions, making the expansion applicable to common rank-test statistics.
- The smooth-function result (Theorem 2.1) could be used to refine Berry-Esseen-type bounds for expectations of smooth functions of rank statistics, which arise in power calculations for rank tests.
Where Pith is reading between the lines
- If condition (2.5) could indeed be verified for the full matrix  alone rather than for all submatrices B ∈ N(8, Â) — as the author conjectures in Remark 2.11(b) — the result would become directly applicable to arbitrary matrices A without case-by-case verification, substantially broadening its practical reach.
- The (log n)² factor in Theorem 2.12(a) arises from the characteristic-function bound in the intermediate frequency range b₃ log n ≤ |t| ≤ b₄ n^{3/2}; if sharper characteristic-function estimates were available (e.g., through non-trivial cancellation arguments), this logarithmic penalty might be removable, matching the optimal O(n^{-1}) rate.
- The combinatorial coupling construction with 8-element index vectors for first order and 16-element vectors for second order suggests a general pattern: an m-th order Edgeworth expansion would require 2^{m+2}-element index vectors and a chain of m+2 permutations, with the complexity growing exponentially in the order.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript establishes a first-order Edgeworth expansion for the distribution function of general linear rank statistics T_A under the null hypothesis, with remainder O(D_A^2) = O(δ_A/n), which is typically O(n^{-1}). The proof combines Stein's method with an extension of Bolthausen's combinatorial method, yielding a basic identity (Lemma 3.20) from which both a smooth-function expansion (Theorem 2.1) and a distribution-function expansion (Theorem 2.4) are derived. The key condition (2.5) requires bounded second differences of F_B for all submatrices B in N(8, Â). For the product-form case a_{ij} = e_i d_j, the author verifies (2.5) via van Zwet's (1982) characteristic function estimates (Theorem 2.12(a)). Second-order results (Theorems 2.7, 2.12(b)) are stated but not proved, with proofs deferred to Schneller (2025). This is a consolidated version of the 1989 Annals of Statistics paper incorporating an erratum correcting the distribution of the random vector I in Section 3.
Significance. The paper provides a self-contained derivation of the first-order Edgeworth expansion for general linear rank statistics, a problem of longstanding interest in nonparametric statistics. The constants K_1,...,K_8 are absolute and not fitted to data, and the expansion coefficients λ_{1,A}, λ_{2,A} are defined directly from matrix entries, so the result is parameter-free in the sense relevant to the axioms. The connection to van Zwet's conditions (Theorem 2.12) provides a concrete, verifiable pathway for the important two-sample and general product-form cases. The transparency about the scope limitation of condition (2.5) — including the explicit conjecture in Remark 2.11(b) that the condition on submatrices B ∈ N(8, Â) should suffice for  alone — is commendable. The erratum correction to the distribution of I (Section 3, equations (3.7)–(3.10)) is properly incorporated.
major comments (3)
- [Theorem 2.4, condition (2.5); proof of Lemma 5.3, A_3 term] The load-bearing assumption (2.5) requires the second-difference bound for all B ∈ N(8, Â), and this requirement arises structurally in the proof of Lemma 5.3: when bounding A_3, the conditioning on (I, J) = (i, j) produces a submatrix B obtained by deleting 8 rows and 8 columns, and the bound |Δ²_D F_B| ≤ 2C₁D² invokes (2.5) for this B. The author is transparent about this in Remark 2.11(b). However, the practical verifiability of (2.5) for general matrices A remains unclear: it is only verified for the product-form case a_{ij} = e_i d_j in Section 6 (via van Zwet's conditions). For a general matrix A satisfying σ_A > 0, no method is provided to check (2.5). This is not an internal inconsistency — the theorem is correct conditional on (2.5) — but it limits the practical scope of Theorem 2.4. The author should add a brief discussion (perhaps 2–3 sentences in Section 2) clarifying which (
- [Proposition 5.7] Proposition 5.7 is a key ingredient in the proof of Theorem 2.4 (used in (5.4) to bound A_2 in Lemma 5.3), but its proof is given only as a 'Sketch of Proof' that delegates substantial steps to Bolthausen (1984) and to Schneller (2025), Chapter 3, Section 6 (see Remark 5.9). The estimate (5.8) in particular is asserted without full justification. Since Proposition 5.7 is load-bearing for the main theorem, the sketch should be expanded to at least outline the proof of (5.8), or the reader should be given a precise pointer to where in Schneller (2025) the complete argument appears.
- [Section 4, truncation argument] The truncation to |a_{ij}| ≤ 1 (used in both Sections 4 and 5) is delegated to Schneller (2025), Chapter 3, Section 4, with only a reference to Bolthausen (1984), pages 382–383 for the basic ideas. Since the truncation underlies both main theorems, a brief summary of the truncation procedure and its effect on the constants would improve self-containedness. This is a presentation gap rather than a correctness concern.
minor comments (6)
- [Section 4, first line] 'We devide' should be 'We divide'.
- [Section 5, Lemma 5.3 proof] The notation 'c.C₁ + c.' in the parenthetical reference to (4.2), (4.3), (4.4) is unclear; presumably this means 'c·C₁ + c·' with specific constants. Clarify.
- [Lemma 6.7 proof] In the proof of Lemma 6.7, the bound on I₃ uses the substitution v = tθ and the inequality θ ≥ 1/(4n). The chain of inequalities leading to I₃ → 0 could be stated more explicitly, as the role of the moment condition (6.2) on Û is somewhat compressed.
- [Section 3, Table 2] Table 2 uses the notation 'Same as π₂' and 'Same as π₃' in a way that could be confusing on first reading; a brief clarifying sentence about what 'Same' means (that the σ-field generated is the same) would help.
- [References] The reference to Schneller (2025) as arXiv:2511.12187 should be checked for consistency with the arXiv identifier format used elsewhere.
- [Section 2, equation (2.16)] The second inequality in (2.16) involves the exponents ((4/k)−1)₊ and ((4/s)−1)₊. It would help the reader to briefly recall (here or in Remark 2.18) under what conditions on k and s these exponents vanish, so that the rate is purely O((log n)² n⁻¹).
Simulated Author's Rebuttal
We thank the referee for a careful and constructive report. The three major comments all concern presentation gaps rather than correctness issues: (1) the practical verifiability of condition (2.5) for general matrices, (2) the sketch-only proof of Proposition 5.7, and (3) the truncation argument delegated to external sources. We agree with all three and will revise accordingly.
read point-by-point responses
-
Referee: The load-bearing assumption (2.5) requires the second-difference bound for all B in N(8, Â), and this requirement arises structurally in the proof of Lemma 5.3... For a general matrix A satisfying σ_A > 0, no method is provided to check (2.5)... The author should add a brief discussion (perhaps 2–3 sentences in Section 2) clarifying which classes of matrices beyond the product-form case satisfy (2.5).
Authors: The referee is correct that condition (2.5) is only verified for the product-form case a_{ij} = e_i d_j in Section 6, and that no general verification method is provided for arbitrary matrices. This is a genuine limitation of the practical scope of Theorem 2.4, and we agree it should be made more transparent in the manuscript itself rather than leaving the reader to infer it. We will add a brief discussion in Section 2 (after Remark 2.11) clarifying that (2.5) is currently verified only for product-form matrices via van Zwet's conditions (Theorem 2.12(a)), and that for general matrices satisfying σ_A > 0, the condition remains a structural assumption whose verifiability is an open problem. We will also note that Remark 2.11(b) already conjectures that the condition on submatrices B ∈ N(8, Â) should suffice for  alone, which would, if true, reduce the verification burden. This addition will not change any results or proofs. revision: yes
-
Referee: Proposition 5.7 is a key ingredient in the proof of Theorem 2.4 (used in (5.4) to bound A_2 in Lemma 5.3), but its proof is given only as a 'Sketch of Proof' that delegates substantial steps to Bolthausen (1984) and to Schneller (2025), Chapter 3, Section 6 (see Remark 5.9). The estimate (5.8) in particular is asserted without full justification. Since Proposition 5.7 is load-bearing for the main theorem, the sketch should be expanded to at least outline the proof of (5.8), or the reader should be given a precise pointer to where in Schneller (2025) the complete argument appears.
Authors: We agree that Proposition 5.7 is load-bearing and that the current sketch, particularly regarding estimate (5.8), is too compressed. The estimate (5.8) is in fact established within the sketch itself: the bound is decomposed into B_1 and B_2, and each is bounded explicitly — B_1 via the truncation set Γ and B_2 via the conditional expectation bound E(|T_A| | π(i)=j) ≤ c_8, which follows from the same arguments as in the proof of Lemma 4.1 (cf. Bolthausen 1984, p. 385, lines 3–14). However, we concede that the logical flow is not clearly delineated, and the connection between (5.8) and the subsequent B_1, B_2 decomposition could be made more explicit. We will expand the sketch to clearly separate the derivation of (5.8) from its application, and to indicate more precisely which steps follow Bolthausen (1984) and which are original. We will also add a precise pointer to Schneller (2025), Chapter 3, Section 6, where a complete (though slightly different) proof for the case |â_{ij}| ≤ 1 is given. revision: yes
-
Referee: The truncation to |a_{ij}| ≤ 1 (used in both Sections 4 and 5) is delegated to Schneller (2025), Chapter 3, Section 4, with only a reference to Bolthausen (1984), pages 382–383 for the basic ideas. Since the truncation underlies both main theorems, a brief summary of the truncation procedure and its effect on the constants would improve self-containedness.
Authors: The referee is right that the truncation is a foundational step for both main theorems and that the current treatment — a single sentence referencing Schneller (2025) and Bolthausen (1984) — is insufficient for self-containedness. We will add a brief summary paragraph at the beginning of Section 4 (before Lemma 4.1) outlining the truncation procedure: the key idea is to replace large entries â_{ij} by truncated versions a'_{ij} = d(â'_{ij}) where â'_{ij} = â_{ij} 1{|â_{ij}| ≤ 1/2}, and to control the resulting approximation error. We will note that this may require reducing the constant ε₀ in the condition β_A ≤ ε₀ n, and that the effect on the constants K_1, ..., K_5 in Theorems 2.1 and 2.4 is absorbed into the absolute nature of these constants. The full details will remain referenced to Schneller (2025), Chapter 3, Section 4, but the summary will allow the reader to understand the logical role of the truncation without consulting the external source. revision: yes
Circularity Check
No significant circularity: the derivation is self-contained with absolute constants and externally verified conditions; self-citations are for deferred proofs, not load-bearing definitions.
full rationale
The paper's central result (Theorem 2.4) is derived from first principles using Stein's method and the combinatorial method of Bolthausen (1984), an external source. The expansion coefficients λ_{1,A}, λ_{2,A} are defined directly from the matrix entries (Section 2), not fitted to data. The constants K_1,...,K_8 are absolute positive constants, not parameters tuned to outputs. Condition (2.5) is a structural assumption on second differences of F_B, not a definition of the conclusion. The verification of (2.5) for the product-form case (Theorem 2.12(a), Section 6) uses van Zwet (1982), an external result, to establish the characteristic function estimate (Lemma 6.5), which is then converted to (2.5) via Lemma 6.3 — a genuine mathematical derivation, not a renaming. Self-citations to Schneller (1987/2025) appear for: (i) the truncation argument (Section 4, referencing Chapter 3, Section 4), (ii) complete proofs of the second-order result (Section 7), and (iii) Proposition 5.7's complete proof (Remark 5.9). These are deferred proofs of lemmas whose statements are given in the paper, not self-citations that define the central result. The proof of Lemma 5.3 invokes (2.5) for submatrices B ∈ N(8, Â) because the combinatorial construction produces such submatrices when conditioning on (I, J) — this is a genuine structural requirement of the proof method, not circular reasoning. The author is transparent in Remark 2.11(b) that (2.5) for  alone might suffice but a proof 'eludes me.' This is a scope limitation, not circularity. No step in the derivation chain reduces to its inputs by construction.
Axiom & Free-Parameter Ledger
axioms (5)
- standard math Stein's equation (3.2): f'(x) - x f(x) = q(x) - Φ(q) characterizes the standard normal distribution, and the solution (3.1) satisfies the stated bounds (4.5) from Erickson (1974).
- standard math Bolthausen's Berry-Esseen bound (1.1): sup_z |F_A(z) - Φ(z)| ≤ K β_A/n for some absolute constant K > 0.
- standard math van Zwet's characteristic function estimate: under conditions (2.13)–(2.15), the characteristic function of T_A satisfies |F̂_B(t)| ≤ b_1 n^{-b_2 log n} for b_3 log n ≤ |t| ≤ b_4 n^{3/2}.
- domain assumption The truncation |a_{ij}| ≤ 1 can be established without loss of generality for the purposes of the proofs.
- domain assumption Condition (2.5) is 'almost necessary' for Edgeworth expansions, analogous to Bickel and Robinson (1982) conditions in the i.i.d. case.
read the original abstract
An Edgeworth expansion of first order is established for general linear rank statistics under the null hypothesis with a remainder term that is usually of order $n^{-1}$. Furthermore, corresponding results for the second order are formulated, but not proved here. The proof for the first order is based on Stein's method and on an extension of the combinatorial method of Bolthausen. It is also shown that conditions of van Zwet imply up to a small factor our conditions for the validity of Edgeworth expansions. Moreover, our proof for the first order also provides us with a result about Edgeworth expansions for smooth functions.
Reference graph
Works this paper leans on
-
[1]
Barbour, A. D.(1986). Asymptotic expansions based on smooth functions in the central limit theorem.Probab. Theory Related Fields72289–303. DOI: 10.1007/BF00699108
-
[2]
Bickel, P. J. and Robinson, J.(1982). Edgeworth expansions and smoothness.Ann. Probab.10500–503. DOI: 10.1214/aop/1176993873
-
[3]
Bickel, P. J. and van Zwet, W. R.(1978). Asymptotic expansions for the power of distribution-free tests in the two-sample problem.Ann. Statist.6937–1004. DOI: 10.1214/aos/1176344305
-
[4]
An estimate of the remainder in a combinatorial central limit theorem.Z
Bolthausen, E.(1984). An estimate of the remainder in a combinatorial central limit theorem.Z. Wahrsch. verw. Gebiete66379–386. DOI: 10.1007/BF00533704
-
[5]
Does, R. J. M. M.(1982). Berry-Esseen theorems for simple linear rank statistics under the null-hypothesis.Ann. Probab.10982–991. DOI: 10.1214/aop/1176993719
-
[6]
Does, R. J. M. M.(1983). An Edgeworth expansion for simple linear rank statistics under the null-hypothesis.Ann. Statist.11607–624. DOI: 10.1214/aos/1176346166. 24Walter Schneller
-
[7]
The Annals of Probability , author =
Erickson, R. V.(1974).L 1 bounds for asymptotic normality ofm-dependent sums using Stein’s technique.Ann. Probab.2522–529. DOI: 10.1214/aop/1176996670
-
[8]
Feller, W.(1971).An Introduction to Probability Theory and Its Application2, 2nd ed
work page 1971
- [9]
-
[10]
DOI: 10.1016/B978-0-12-642350-1.X5017-6
Press, San Diego and London. DOI: 10.1016/B978-0-12-642350-1.X5017-6
-
[11]
Ho, S. T. and Chen, L. H. Y.(1978). AnL p bound for the remainder in a combinatorial central limit theorem.Ann. Probab.6231–249. DOI: 10.1214/aop/1176995570
-
[12]
Hoeffding, W.(1951). A combinatorial central limit theorem.Ann. Math. Statist.22 558–566. DOI: 10.1214/aoms/1177729545
-
[13]
On the Hoeffding’s combinatorial central limit theorem.Ann
Motoo, M.(1957). On the Hoeffding’s combinatorial central limit theorem.Ann. Inst. Statist. Math.8145–154. DOI: 10.1007/BF02863580
-
[14]
Schwarz,Estimating the dimension of a model, The annals of statistics, 6 (1978), pp
Robinson, J.(1978). An asymptotic expansion for samples from a finite population.Ann. Statist.61005–1011. DOI: 10.1214/aos/1176344306
-
[15]
Schneller, W.(1987).Edgeworth-Entwicklungen f¨ ur lineare Rangstatistiken. Ph.D. thesis, Technische Universit¨ at Berlin
work page 1987
-
[16]
A short proof of Motoo’s combinatorial central limit theorem using Stein’s method.Probab
Schneller, W.(1988). A short proof of Motoo’s combinatorial central limit theorem using Stein’s method.Probab. Theory Related Fields78249–252. DOI: 10.1007/BF00322021
-
[17]
Edgeworth expansions for linear rank statistics.Ann
Schneller, W.(1989). Edgeworth expansions for linear rank statistics.Ann. Statist.17 1103–1123. DOI: 10.1214/aos/1176347258
-
[18]
Schneller, W.(2025).Edgeworth Expansions for Linear Rank Statistics Using Stein’s Method[English translation of an updated version of the author’s Ph.D. thesis from 1987]. Available at arXiv:2511.12187
-
[19]
Erratum: Edgeworth expansions for linear rank statistics.Ann
Schneller, W.(2026). Erratum: Edgeworth expansions for linear rank statistics.Ann. Statist.541649–1653. DOI: 10.1214/25-AOS2611
-
[20]
Schneller, W.(2026a). Supplement A to ”Erratum: Edgeworth expansions for linear rank statistics.” DOI: 10.1214/25-AOS2611SUPPA
-
[21]
Schneller, W.(2026b). Supplement B to ”Erratum: Edgeworth expansions for linear rank statistics.” DOI: 10.1214/25-AOS2611SUPPB
-
[22]
Stein, C.(1972). A bound for the error in the normal approximation to the distribution of a sum of dependent random variables.Proc. Sixth Berkeley Symp. Math. Statist. Probab.2 583–602. Univ. California Press URL: https://projecteuclid.org/ebooks/berkeley- symposium-on-mathematical-statistics-and-probability/Proceedings-of-the-Sixth- Berkeley-Symposium-on...
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.