Pith. sign in

REVIEW 3 major objections 5 minor 29 references

Multi-sample rank tests for location against Lehmann-type alternatives

T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Under the new Lehmann-type alternative in (2), every rank test statistic is distribution-free.

desk verdict Sound distribution-free theorem for a new k-sample Lehmann alternative, but the practical recommendations rest on an unsupported chi-square approximation and contradictory power claims. read the letter →

arxiv 2506.01914 v1 pith:V2CFDEFP submitted 2025-06-02 math.ST stat.TH

classification math.STstat.TH MSC 62G1062G2062G3062E15
keywords multi-sampletestingproblemorderedalternativeLehmann-typealternativesranktestsprecedencestatisticsexceedanceJonckheere-Terpstratestdistribution-free
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper proposes a new way to model the ordered alternative in the k-sample homogeneity testing problem and uses it to build rank tests whose properties do not depend on the unknown baseline distribution. The alternative models the k populations as the k order statistics of a single sample from a common continuous distribution F, so the i-th population has distribution $F^*_i(x) = P(Z_{[i]} \le x)$. The authors prove that under this Lehmann-type alternative the probability of any rank order, and hence the distribution of any rank statistic, is free of F. They introduce two families of multi-sample tests based on precedence and exceedance statistics, $M_\rho$ and $V_\rho$, extending two-sample procedures, and compare their power with the Jonckheere-Terpstra test by simulation. If the order-statistic model is a reasonable description of stochastic ordering, the practical gain is that critical values and power can be tabulated without specifying the underlying distribution.

What carries the argument

The load-bearing object is the order-statistic Lehmann alternative in (2), written as $F^*_i(x) = P(Z_{[i]} \le x)$ for order statistics $Z_{[1]} \le \dots \le Z_{[k]}$ of k iid variables with distribution F. Because $F^*_i$ is a polynomial in $F(x)$ and $1 - F(x)$, the rank-order probability in Theorem 1 is a fixed integral over $0 \le y \le 1$ after substituting $y = F(x)$, so the unknown F cancels. For the power-function subclass $F_i(x) = [F(x)]^{\eta_i}$, Theorem 3 gives the explicit rank-order probability in (8), a product of $\eta$-powers divided by successive partial sums of $\eta$-values, which makes the distribution-free property transparent. The test statistics $M_\rho$ and $V_\rho$ are built from trimmed extreme ranks: $M_\rho$ takes the largest deviation of each sample's upper and lower ranks from the ideal contiguous ranking, while $V_\rho$ sums those deviations; rejecting for small values of either statistic implements the ordered alternative.

What would settle it

Take k=3 samples of size 1, so the observation from population i is the i-th order statistic of three iid draws from F. Simulate the rank-order probability for a non-identity permutation under F = Uniform(0,1) and under F = Beta(0.5,0.5). If the two simulated probabilities differ beyond Monte Carlo error, the distribution-free claim in Theorem 1 fails; if they agree and match the integral in (4), the claim is confirmed.

Watch

Extended reading notes

Core claim

The central claim is that the new Lehmann-type alternative defined by order statistics, $H_L$ in equation (2), makes every rank test statistic distribution-free in the k-sample ordered-alternative problem. Concretely, if the i-th sample has distribution $F^*_i(x) = P(Z_{[i]} \le x)$, where $Z_{[1]} \le \dots \le Z_{[k]}$ are order statistics of k independent observations from a common continuous F, then the probability of any rank permutation r is given by an integral over the unit cube that no longer contains F; the change of variables $y = F(x)$ removes the unknown baseline. The paper further proves an explicit product formula for the more general power-function alternative $F_i(x) = [F(x)]^{\eta_i}$, extending Savage's two-sample result, and presents two families of test statistics, $M_\rho$ and $V_\rho$, based on maximum deviations and on sums of precedence/exceedance gaps, with Monte Carlo critical values and power comparisons against the Jonckheere-Terpstra test.

Load-bearing premise

The load-bearing premise is that the real-world ordered alternative can be represented as order statistics from one common baseline distribution; if the actual populations are ordered in some other way, such as by location shifts, the distribution-free property proved for this model does not transfer to that setting.

Editorial extensions

If this is right

  • Critical values and null distributions of $M_\rho$ and $V_\rho$ can be tabulated once and reused for every continuous baseline F, since the distribution of any rank statistic under $H_L$ does not depend on F.
  • Power comparisons against the Jonckheere-Terpstra test are meaningful without specifying F; the simulations show JT has the highest power among the compared tests, with $V_\rho$ close behind and $M_\rho$ lower.
  • The trimming parameter $\rho$ gives a family of tests whose robustness to outliers can be tuned; setting $\rho > 0$ discards extreme ranks within each sample.
  • The explicit formula (8) for the power-function alternative allows exact or near-exact computation of rank-order probabilities for $F_i(x) = [F(x)]^{\eta_i}$, extending Savage's two-sample formula to k samples.
  • The $\chi^2$ approximation to the null distribution of $V_\rho$ matches the exact tail reasonably, so the proposed tests can be applied with approximate critical regions in practice.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the order-statistic model is a plausible description of dose-response or treatment-level experiments, the distribution-free property means that a single set of tables can serve for all baselines; the paper does not argue the model's empirical realism.
  • The same construction could be adapted to two-sided or unordered alternatives by replacing the extremal rank set E with other target permutations, a direction the paper does not pursue.
  • The connection to step-stress lifetime models noted in the closing remarks suggests the tests could be applied to accelerated life testing where stress levels induce ordered order-statistic behavior; this is an extension of the paper's remarks, not one of its claims.
  • The explicit rank probabilities in (8) could be combined with the Neyman-Pearson lemma to derive locally most powerful rank tests for the power-function alternative, following the logic the paper attributes to Lehmann; the paper stops short of doing this.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes a new Lehmann-type alternative for the k-sample ordered-alternative problem, defined through the order statistics of k i.i.d. variables from a common continuous distribution F, so that the i-th population CDF is the CDF of the i-th order statistic. The central theoretical claim is that under this model the distribution of any rank statistic is free of the baseline F; Theorem 1 expresses the probability of a rank order as an F-free integral, and Theorem 3 specializes Savage's formula to the power alternative Fi(x)=F(x)^\eta_i. Two families of rank tests, M_rho and V_rho, are defined as extensions of two-sample precedence/exceedance statistics, null critical values are tabulated by simulation for k=3,4,5 and equal sample sizes, a chi-square approximation for V_rho is suggested, and power comparisons against the Jonckheere-Terpstra test are reported under two Lehmann-type models.

Significance. If the distribution-free claim is correct, it is a clean and useful extension of the two-sample Lehmann result: it gives practitioners a nonparametric ordered-alternative model with explicit rank probabilities and distribution-free tests under the alternative. Theorem 1 is mathematically sound: the integral in Eq. (4) is F-free by the probability integral transform, and the proposed M_rho and V_rho statistics are functions of the rank vector only, so their distributions under HL do not depend on the unknown baseline. The paper also credits the earlier two-sample statistics on which the new tests are built. However, the applied part of the paper—the chi-square approximation and the power conclusions—contains unsupported and internally inconsistent claims that must be resolved before the practical recommendations can be accepted.

major comments (3)
  1. [Section 5, Table 4] The text states that for the V_rho test 'we can use the critical regions derived by the chi-square approximation instead of the exact regions,' but no derivation, degrees of freedom, or goodness-of-fit justification is provided. Table 4 itself shows a substantial discrepancy: for k=3, n=10, rho=0, the exact critical value 46 has tail probability 0.0347, while the chi-square critical value 46.6 has tail probability 0.05, making the approximate critical region about 44% more liberal in tail probability. This is a load-bearing practical claim and needs either a proper derivation or an explicit warning that the approximation is only rough and should not replace the exact tables.
  2. [Section 6.1, after Figures 3-5] The power statements are internally contradictory. The text says 'the power of V_rho-test for small rho (0 <= rho <= 0.25) approaches its level for larger sample sizes' and then, two sentences later, 'The power of V_rho-tests is very high, close to the power of the Jonckheere-Terpstra test.' If power approaches the nominal level, it is not very high except in the trivial sense of being near 0.05. The same contradiction appears in Section 6.2. The authors must decide which statement is correct and report the actual power values or figures that support it.
  3. [Section 6, power study] The power comparison is not reproducible and lacks precision. Ten thousand simulation runs are mentioned, but no Monte Carlo standard errors or confidence intervals are reported, so the reader cannot judge whether differences between tests are meaningful, especially when power is near the nominal level. If the claim is that V_rho power is close to the Jonckheere-Terpstra test, a numerical table with standard errors is needed. If, instead, the tests' power approaches the level as sample size grows, this would suggest inconsistency and would undermine the practical recommendation; this must be stated explicitly and reconciled with the 'very high' claim.
minor comments (5)
  1. [Section 1, Eq. (2)] The notation HL is used both for the new order-statistic alternative and for Savage's power alternative HL_eta; please use distinct names (e.g., HL^ord and HLk,eta) to avoid ambiguity.
  2. [Section 4.2, after Eq. (13)] In the notation line 'Xi,1,...,Xi,ni ~ Fj, i=1,...,k', the subscript j in Fj appears to be a typo for i; please correct it.
  3. [Tables 1-3] The caption says the critical values are 'at near 5% level of significance', and the tail probabilities in parentheses are not defined clearly; please state explicitly that they are left-tail probabilities for M_rho and V_rho and right-tail probabilities for JT, before the randomization adjustment in Eq. (15).
  4. [Sections 6.1 and 6.2] The two alternative models, the order-statistic HL in (2) and the power alternative HLk,eta with eta_i=i in (7), are both called Lehmann-type; please use separate names so that the reader can tell which simulation is being discussed.
  5. [Section 7] The hazard-rate remark about G(x)=F(x)^eta and G*(x)=1-(1-F(x))^eta is interesting but not connected to the proposed tests; either draw a connection or remove it as tangential.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: the distribution-free claim under HL is derived in Theorem 1 by the probability integral transform, the M_rho and V_rho families are defined explicitly in the paper, self-citations are motivational only, and the power study is benchmarked externally against Jonckheere-Terpstra.

full rationale

The paper's central claim, that every rank statistic is distribution-free under the Lehmann-type alternative HL in (2), is a genuine derivation rather than a restatement of inputs. Theorem 1 computes P(R = r | HL) as the integral in (4), obtained from (5) by the change of variables x_{i,j} = F^{-1}(y_{i,j}); the integrand and the integration region 0 ≤ y_{a_1(r),b_1(r)} ≤ ... ≤ y_{a_n(r),b_n(r)} ≤ 1 contain no F, so the rank probabilities are F-free. The alternative HL is defined through order statistics of k iid observations with CDF F (equation (2)), not through the test statistics, so the distribution-free property is a derived consequence, not an input assumption. The proposed statistics M_rho and V_rho are fully specified in the paper in equations (10) and (13) as explicit functions of the rank vector; their null critical values are obtained in this paper by Monte Carlo simulation (Section 5), not imported from the cited two-sample work of Stoimenova and Balakrishnan, so those self-citations are motivational rather than load-bearing. The power study compares against the external Jonckheere-Terpstra benchmark and simulates from the model at Uniform(0,1), which is an application of Theorem 1 rather than a circular step. Two non-circularity concerns are noted for the correctness pass: the Section 5 chi-square approximation is asserted without derivation and Table 4 itself shows noticeable discrepancies (e.g., exact tail 0.0347 at c.v. 46 versus chi-square tail 0.05 at c.v. 46.6 for n = 10, rho = 0), and Section 6.1 contains apparently contradictory statements about the power of V_rho for small rho. These affect reliability, not circularity. Finding: no significant circularity.

Assumptions & free parameters 2 free parameters · 4 assumptions · 0 invented entities

The central theoretical claim rests only on standard probability and the cited Savage theorem. The main tuning choices are rho and the chi-square approximation fit, neither of which is estimated from data in a way that would induce circularity.

free parameters (2)
  • rho = 0 to 0.25 in steps of 0.05
    Trimming proportion used in the test statistics M_rho and V_rho; chosen by the authors in the simulation study, not estimated from data.
  • chi-square approximation parameters = not specified, mean-matched to exact null distribution
    The chi-square approximation of the V_rho null distribution is constructed to have the same expectation, requiring fitted scale or degrees-of-freedom that are not reported.
assumptions (4)
  • domain assumption The k samples are mutually independent and each distribution is absolutely continuous.
    Stated in Section 1 and used throughout for rank-based inference and for the no-ties condition.
  • standard math Savage's theorem (Theorem 2 in the paper) giving the rank probability under power alternatives.
    Used without proof to derive Theorem 3, the k-sample power alternative rank probability.
  • standard math The inverse CDF F^{-1} is non-decreasing and continuous, so the change of variables in Theorem 1 is valid.
    Invoked in the proof of Theorem 1.
  • domain assumption Monte Carlo simulation with 10,000 replications approximates the null and alternative distributions well enough for critical values and power.
    The paper does not provide standard errors or theoretical guarantees, yet uses these values for conclusions.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Multi-sample rank tests for location against Lehmann-type alternatives." pith.science (2026). https://pith.science/paper/V2CFDEFP

@misc{pith2026250601914,
  author       = {Pith},
  title        = {Pith review of: Multi-sample rank tests for location against Lehmann-type alternatives},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/V2CFDEFP}},
  note         = {Machine review of arXiv:2506.01914}
}
abstract

This paper deals with testing the equality of $k$ ($k\ge 2$) distribution functions against possible stochastic ordering among them. Two classes of rank tests are proposed for this testing problem. The statistics of the tests under study are based on precedence and exceedance statistics and are natural extension of corresponding statistics for the two-sample testing problem. Furthermore, as an extension of the Lehmann alternative for the two-sample location problem, we propose a new subclass of the general alternative for the stochastic order of multiple samples. We show that under the new Lehmann-type alternative any rank test statistics is distribution free. The power functions of the two new families of rank tests are compared to the power performance of the Jonckheere-Terpstra rank test.

Figures

Figures reproduced from arXiv: 2506.01914 by the authors.

Figure 1
Figure 1. Null distributions of Vρ for 3 samples with equal sample sizes n3 = 20, and ρ = 0(0.05)0.25 with χ 2 -approximation [PITH_FULL_IMAGE:figures/full_fig_p015_1.png] view at source ↗
Figure 2
Figure 2. Boxplots of the data of Example We compute the statistics Mρ and Vρ for ρ = 0(0.05)0.25 [PITH_FULL_IMAGE:figures/full_fig_p016_2.png] view at source ↗
Figure 3
Figure 3. Power of Vρ and Mρ tests vs. Lehmann alternative for 3 samples 6.1. Lehmann-type alternative with order statistics Now, we will demonstrate the use of Monte Carlo simulation method for the com￾putation of the power of the exceedance tests Mρ tests and Vρ-tests against the Lehmann￾type alternatives in (2). For this purpose, we have to generate sets of data from the distributions of the order statistics in (2) and com… view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: Power of Vρ and Mρ tests vs. Lehmann alternative for 4 samples [PITH_FULL_IMAGE:figures/full_fig_p019_4.png]
Figure 5
Figure 5. Figure 5: Power of Vρ and Mρ tests vs. Lehmann alternativ for 5 samples 19 [PITH_FULL_IMAGE:figures/full_fig_p019_5.png]
Figure 6
Figure 6. Figure 6: Power of Vρ and Mρ tests for 3 samples vs. Lehmann alternative Fni (x) = F(x) i [PITH_FULL_IMAGE:figures/full_fig_p021_6.png]
Figure 7
Figure 7. Figure 7: Power of Vρ and Mρ tests for 4 samples vs. Lehmann alternative Fni (x) = F(x) i 21 [PITH_FULL_IMAGE:figures/full_fig_p021_7.png]
Figure 8
Figure 8. Figure 8: Power of Vρ and Mρ tests for 5 samples vs. Lehmann alternative Fni (x) = F(x) i random variable with pdf f(t) and CDF F(t), then the hazard and reversed hazard rate are defined respectively by hF (x) = f(x) 1 − F(x) and rF (x) = f(x) F(x) ; x > 0. If G(x) and G∗ (x) ar…

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

29 extracted references · 29 canonical work pages

  1. [1]

    Comparing the performance of nonparamet- ric tests for equality of location against ordered alternatives

    Altunkaynak, B., Gamgam, H., 2021. Comparing the performance of nonparamet- ric tests for equality of location against ordered alternatives. Communications in Statistics-Simulation and Computation 50, 63–84

  2. [2]

    A general theory of hypothesis testing based on rankings

    Alvo, M., Pan, J., 1997. A general theory of hypothesis testing based on rankings. Journal of Statistical Planning and Inference 61, 219–248

  3. [3]

    My musings on a pioneering work of erich lehmann and its rediscoveries on some families of distributions

    Balakrishnan, N., 2021. My musings on a pioneering work of erich lehmann and its rediscoveries on some families of distributions. Communications in Statistics- Theory and Methods 51, 8066–8073

  4. [4]

    A sequential order statistics approach to step-stress testing

    Balakrishnan, N., Kamps, U., Kateri, M., 2012. A sequential order statistics approach to step-stress testing. Annals of the Institute of Statistical Mathematics 64, 303–318

  5. [5]

    Precedence-type tests and applications

    Balakrishnan, N., Ng, H.T., 2006. Precedence-type tests and applications. volume 472. John Wiley & Sons

  6. [6]

    On rank statistics: An approach via metrics on the permutation group

    Critchlow, D.E., 1992. On rank statistics: An approach via metrics on the permutation group. J. Stat. Plann. Inference 32, 325–346

  7. [7]

    Rank Tests for ”Lehmann’s Alternative

    Davies, R.B., 1971. Rank Tests for ”Lehmann’s Alternative. JASA 66, 879–883

  8. [8]

    C-sample tests of homogeneity against or- dered alternatives, in: Optimizing Methods in Statistics

    Govindarajulu, Z., Haller Jr, H.S., 1971. C-sample tests of homogeneity against or- dered alternatives, in: Optimizing Methods in Statistics. Elsevier, p. 479

Show all 29 references
  1. [9]

    A two-sample rank test on location

    Haga, T., 1959. A two-sample rank test on location. Annals of the Institute of statistical mathematics 11, 211–219

  2. [10]

    Power calculations for preclinical studies using a k-sample rank test and the lehmann alternative hypothesis

    Heller, G., 2006. Power calculations for preclinical studies using a k-sample rank test and the lehmann alternative hypothesis. Statistics in medicine 25, 2543–2553

  3. [11]

    “optimum” nonparametric tests, in: The Collected Works of Wassily Hoeffding

    Hoeffding, W., 1951. “optimum” nonparametric tests, in: The Collected Works of Wassily Hoeffding. Springer, pp. 227–236

  4. [12]

    A distribution-free k-sample test against ordered alternatives

    Jonckheere, A., 1954. A distribution-free k-sample test against ordered alternatives. Biometrika 41, 133–145

  5. [13]

    Inference in step-stress models based on failure rates

    Kateri, M., Kamps, U., 2015. Inference in step-stress models based on failure rates. Statistical Papers 56, 639–660. 23

  6. [14]

    Hazard rate modeling of step-stress experiments

    Kateri, M., Kamps, U., 2017. Hazard rate modeling of step-stress experiments. Annual Review of Statistics and Its Application 4, 147–168

  7. [15]

    The distribution of two-sample location exceedance test statistics under Lehmann alternatives

    Katzenbeisser, W., 1985. The distribution of two-sample location exceedance test statistics under Lehmann alternatives. Statistical Papers (Statistische Hefte) 26, 131– 138. K¨ossler, W., 2005. Some c-sample rank tests of homogeneity against ordered alterna- tives based on u-s...

  8. [16]

    The power of rank tests

    Lehmann, E., 1953. The power of rank tests. Ann. Math. Stat. 24, 23–43

  9. [17]

    Testing statistical hypotheses

    Lehmann, E.L., Romano, J.P., 2022. Testing statistical hypotheses. Springer

  10. [18]

    Statistical methods based on ranks

    Lehmann, E.L., et al., 1975. Statistical methods based on ranks. Nonparametrics. San

  11. [19]

    On the choice of precedence tests

    Lin, C., Sukhatme, S., 1992. On the choice of precedence tests. Communications in Statistics-Theory and Methods 21, 2949–2968

  12. [20]

    Powers of two-sample rank tests under Lehmann alternatives

    Lin, C.H., 1990. Powers of two-sample rank tests under Lehmann alternatives. Iowa State University

  13. [21]

    Rank tests of maximal power against lehmann-type alternatives

    Peto, R., 1972. Rank tests of maximal power against lehmann-type alternatives. Biometrika 59, 472–475

  14. [22]

    Contributions to the theory of rank order statistics-the two-sample case

    Savage, I.R., 1956. Contributions to the theory of rank order statistics-the two-sample case. The Annals of Mathematical Statistics 27, 590–615

  15. [23]

    Tables of the distribution of the mann-whitney-wilcoxon u- statistic under lehmann alternatives

    Shorack, R.A., 1967. Tables of the distribution of the mann-whitney-wilcoxon u- statistic under lehmann alternatives. Technometrics 9, 666–677

  16. [24]

    Theory of rank tests

    Sidak, Z., Sen, P.K., Hajek, J., 1999. Theory of rank tests. Elsevier

  17. [25]

    Power of exceedance-type tests under Lehmann alternatives

    Stoimenova, E., 2011. Power of exceedance-type tests under Lehmann alternatives. Comm. Statist. - Th. M. 40 (4), 731–744

  18. [26]

    A class of exceedance-type statistics for the two-sample problem

    Stoimenova, E., Balakrishnan, N., 2011. A class of exceedance-type statistics for the two-sample problem. J. Stat. Plann. Inf. 141 (9), 3244–3255

  19. [27]

    ˇSid´ak-type tests for the two-sample problem based on precedence and exceedance statistics

    Stoimenova, E., Balakrishnan, N., 2017. ˇSid´ak-type tests for the two-sample problem based on precedence and exceedance statistics. Statistics 51, 247–264

  20. [28]

    Powers of two-sample rank tests under the lehmann alternatives

    Sukhatme, S., 1992. Powers of two-sample rank tests under the lehmann alternatives. The American Statistician 46, 212–214. 24

  21. [29]

    The asymptotic normality and consistency of Kendall’s test against trend, when ties are present in one ranking

    Terpstra, T., 1952. The asymptotic normality and consistency of Kendall’s test against trend, when ties are present in one ranking. Indag. Math. 14, 327–333. van der Laan, P., Chakraborti, S., 2001. Precedence tests and Lehmann alternatives. Statist. Papers 42, 301–312. V ock,...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.