Pith. sign in

REVIEW 4 major objections 4 minor 59 references

This paper extends the Kolmogorov-Smirnov, Cramer-von Mises, and Anderson-Darling tests to weighted samples and shows, through a single label-permutation calibration, that all three keep false-positive rates near nominal while preserving th

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 06:42 UTC pith:2F5TJNFJ

load-bearing objection Useful weighted ECDF tests for balance diagnostics, but the 'works for any weights' claim is only validated for exogenous weights—under IPTW the permutation null is not obviously correct. the 4 major comments →

arxiv 2607.21782 v1 pith:2F5TJNFJ submitted 2026-07-23 stat.ME

Weighted Extensions of the Kolmogorov-Smirnov, Cramer-von Mises, and Anderson-Darling Tests for Assessing Covariate Balance

classification stat.ME MSC 62G1062G3062G09
keywords covariate balanceweighted goodness-of-fitKolmogorov-Smirnov testAnderson-Darling testCramer-von Mises testpermutation inferencecase weightspropensity score weighting
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

After matching, weighting, or any other adjustment, researchers need to check whether a covariate's whole distribution is balanced across groups, not just its mean or variance. This paper makes three classical distribution-comparison tests—Kolmogorov-Smirnov, Cramer-von Mises, and Anderson-Darling—usable on weighted samples, and calibrates all three with one label-permutation procedure that treats each case's weight as a fixed attribute. In simulations with heavily varying weights, all three tests rejected at rates close to the nominal 5% under a true null; KS remained the most powerful for a central discrepancy, AD was clearly best for a tail discrepancy, and AD and CVM performed comparably for a diffuse shift. The paper's practical conclusion is that AD is a reasonable general-purpose default for routine covariate-balance checks, with KS reserved for when a central imbalance is specifically suspected.

Core claim

The paper's central claim is that the Kolmogorov-Smirnov, Cramer-von Mises, and Anderson-Darling statistics can be lifted to any positive case-weight setting by replacing ordinary empirical CDFs with weighted empirical CDFs, and in the Anderson-Darling statistic by replacing the sample size N with the Kish effective sample size inside the variance-standardization term. All three statistics are then calibrated with one label-permutation procedure that shuffles group labels over fixed (value, weight) pairs. The paper argues, and supports by simulation, that this yields tests whose false-positive rates stay near nominal under substantial weight variability, that the classical sensitivity orderi

What carries the argument

The apparatus is (1) the weighted empirical CDF, where each observation contributes w_i/sum(w) instead of 1/n; (2) the Kish effective sample size inserted into AD's variance standardization, which prevents the weighted null variance from being understated when weights vary; and (3) the permutation p-value, computed as the proportion of relabeled datasets whose statistic equals or exceeds the observed one. These pieces make the weighted statistics reduce exactly to their classical forms when all weights are equal, and they make the inference independent of how the weights were generated.

Load-bearing premise

The load-bearing premise is that valid inference follows from permuting treatment labels across fixed (value, weight) pairs under the null; this exchangeability can break when weights are themselves estimated from treatment and covariates, because each weight already encodes information about the label being shuffled.

What would settle it

Generate a dataset in which treatment depends on a covariate through a logistic model, construct inverse-propensity weights, and then apply the three permutation tests to a covariate that is truly balanced after weighting; if the tests reject more often than nominal across sample sizes, the claim that calibration is independent of how weights were generated is falsified.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Applied researchers can now run full-distribution balance checks on weighted data, rather than relying only on standardized mean differences and variance ratios.
  • A single permutation calibration works across all three statistics and any positive weighting scheme, so no separate null-distribution formula is needed for each weighting method.
  • The known sensitivity profiles carry over: KS is the most powerful test for a centrally located imbalance, AD is dramatically the best for tail imbalance, and AD and CVM are comparably good (and both better than KS) for a diffuse shift.
  • The paper's recommendation is to use AD as the routine default for balance diagnostics, to prefer KS when a central imbalance is specifically suspected, and to consider reporting AD and KS together to cover both blind spots.
  • At very severe weight variability (effective sample size ratio 0.2), all three tests show a modest false-positive inflation—about 0.064 to 0.067 at nominal 0.05—so calibration is not perfectly invariant to extreme weights.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If weights are estimated from treatment and covariates, as with inverse-propensity weighting, each unit's weight already encodes its treatment propensity; the fixed-weight permutation null may then not match the true null. A direct simulation with treatment-dependent weights would settle whether the 'no assumption about weight origin' claim holds beyond exogenously assigned weights.
  • The diffuse-scenario result suggests AD and CVM coincide near a uniform shift because AD's variance standardization has nothing to emphasize. Varying the discrepancy's functional form or the base distribution's shape could map the boundary between the tests' regimes.
  • Because every formula uses weights only in ratios, the tests are invariant to rescaling weights; this likely makes them applicable to survey weights or post-stratification weights without any extra adjustment.
  • The same weighted-ECDF-plus-permutation framework could be extended immediately to the Kuiper or Wasserstein statistics, which the paper itself names as future directions, adding tests sensitive to cyclical or cost-scaled discrepancies.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The manuscript extends the two-sample Kolmogorov-Smirnov, Cramér–von Mises, and Anderson-Darling tests to case-weighted data for covariate-balance assessment. It defines weighted ECDFs and weighted statistics, and calibrates all three using a single label-permutation procedure that treats each observation's weight as a fixed attribute. The simulation study uses exogenous lognormal weights at sample sizes 1,000–4,000 and reports Type I error near nominal and preserved comparative advantages (KS for central, AD for tail, AD/CVM for diffuse discrepancies). An IPTW application to BMI is presented, and Stata implementations are referenced. The paper recommends AD as a reasonable general-purpose default.

Significance. If the permutation calibration is valid for weights of the kind used in practice, the paper fills a practical gap: applied covariate-balance assessment after IPTW or entropy balancing currently relies mostly on moment-based summaries, and weighted distributional tests are not routinely available. The weighted KS and CVM forms follow naturally from weighted ECDFs, the shared permutation inference is transparent and reduces to the classical unweighted test, and the simulation includes paired power comparisons and a weight-variability sensitivity check. These are genuine strengths. However, the central generality claim — that the inference 'requires no assumption about how the weights were generated' — is not supported by the simulations, which assign weights exogenously and independently of treatment; the applied example uses IPTW weights. The AD_w standardization through Kish effective sample size is asserted rather than derived. These points affect the validity of the headline claims and require revision.

major comments (4)
  1. [Sec. 2.5 / Sec. 2.6.1] The claim that the label-permutation procedure 'requires no assumption about how the weights were generated' is broader than what is shown. Permuting labels among fixed (value, weight) pairs generates a null under which (X, W) are jointly exchangeable across treatment labels. For IPTW or entropy-balancing weights, W is a function of G and X; even when weighted covariate distributions are equal, (X, W) is typically not exchangeable with G. The simulation (Eq. 12) assigns weights independently of y, g, and p, so Tables 2 and 5 validate the procedure only under exogenous weights. The applied example (Sec. 4.1) uses IPTW weights, so the p-values in Table 6 inherit this unvalidated assumption. Please either prove/examine validity for estimated weights, add simulations with estimated propensity or entropy weights, or explicitly restrict the claim.
  2. [Sec. 2.4, Eqs. (7)-(9)] The Anderson-Darling standardization is not derived. Eq. (7) contains an ambiguous factor 'N·2' (and Eq. (9) an analogous 'ne·2') whose origin is unclear; the usual null variance of the ECDF difference involves a term like D(1-D)(1/n1+1/n0), not an unexplained N×2 constant. Since permutation p-values are invariant to a common multiplicative constant, this ambiguity does not invalidate the reported p-values, but it affects the reported statistic and the interpretation of AD_w. More importantly, replacing N by the Kish effective sample size n_e (Eq. 8) is asserted as 'the standard adjustment' without derivation; this choice changes the relative weighting across x and should be justified or supported by simulation under estimated weights.
  3. [Table 5 / Sec. 3.6] At the most severe weight variability examined (r_e = 0.2), all three tests show Type I error rates of .064-.067, a relative inflation of roughly 30% over nominal .05. The text acknowledges this but the abstract and Sec. 5.1 conclude that all three tests 'controlled Type I error close to nominal under substantial weight variability' without stating the boundary of validity. If r_e = 0.2 is within the intended scope, the claim should be qualified; if it is outside, the qualifying condition should be stated. As written, the conclusion is stronger than the table supports.
  4. [Secs. 2.6, 5.3] The recommendation of AD as a general-purpose default is broader than the simulation evidence. The DGP is normal throughout, each discrepancy type is examined at one fixed effect size, and the scenarios are deliberately constructed to correspond to the known theoretical strengths of the three tests (Sec. 2.6). Such scenarios can show that weighting preserves known comparative advantages under exogenous weights, but they do not establish that AD is the best default under realistic imbalance structures, non-normal covariate distributions, or estimated weights. Please temper the applied recommendation or broaden the simulation accordingly.
minor comments (4)
  1. [Eq. (7)] The formula is difficult to parse: the placement of the 'N·2' factor and the square root is unclear. Please rewrite with explicit numerators/denominators and define all symbols at first use.
  2. [Sec. 2.4.2 / Sec. 5.3] The tie multiplier tau_k is retained as the raw unweighted count. For weighted data this is not obviously the right information measure; the caveat in Sec. 5.3 is useful, but the property is not examined in the simulations.
  3. [Introduction] The citation for Somers' D (Newson and Falcaro, 2023) points to a reference titled 'Robit regression in Stata', which appears unrelated. Please verify the reference.
  4. [Sec. 2.5] The definition of the permutation p-value (Eq. 10) uses '#' rather than an explicit count and could be made clearer by writing (1 + sum_j 1(T_j^* >= T_obs))/(R+1).

Circularity Check

0 steps flagged

No significant circularity: the weighted statistics are direct extensions, and the simulation results are empirical rather than built into the definitions.

full rationale

The paper's derivation is self-contained. The weighted KS, CVM, and AD statistics are obtained by substituting weighted ECDFs (and the Kish effective sample size for the AD variance term) into classical unweighted definitions; each reduces to the unweighted statistic when weights are equal, but this is a consistency property of the extension, not a restatement of the target conclusion. The permutation procedure is a standard label-permutation calibration applied to whatever statistic is supplied, and it does not define the test's conclusions. The simulation outcomes could have failed—for example, weighting could have altered the comparative ordering—and the paper in fact reports a partial non-confirmation for the diffuse scenario, since CVM does not uniquely dominate AD. This indicates the results are not forced by construction. The self-citations to the author's own Stata commands are implementation references, not load-bearing evidence for the statistical claims. The main scientific weakness—that the simulation weights are generated independently of treatment and covariates, so the 'no assumption about how the weights were generated' claim is not validated for estimated IPTW weights—is a validity and external-generalizability concern, not a circularity: the permutation p-value is not defined in terms of the conclusion it is used to support, and no specific equation or fitted parameter is renamed as a prediction.

Axiom & Free-Parameter Ledger

4 free parameters · 4 axioms · 0 invented entities

The method introduces no new physical or statistical entities. The load-bearing assumptions are (1) permutation exchangeability for arbitrary weights, (2) the Kish standardization in AD_w, and (3) the ported tie-multiplier rule. The free parameters are simulation tuning choices rather than parameters of the proposed tests.

free parameters (4)
  • Delta_center = 6
    Shift magnitude for the centrally located discrepancy, tuned in preliminary simulations to produce a non-degenerate power curve; a design choice, not a method parameter.
  • Delta_tail = 50
    Shift magnitude for the tail-located discrepancy, also hand-tuned; arbitrary.
  • delta_diffuse = 1
    Uniform shift for the diffuse scenario; chosen by hand.
  • r_e (Kish effective sample size ratio) = 0.6 (main); 0.2, 0.4, 0.8, 0.9 (sensitivity)
    Target weight-variability level in the lognormal weight DGP; chosen as representative of propensity-score-weighted data, not derived from first principles.
axioms (4)
  • domain assumption Under H0, group labels are exchangeable with respect to (value, weight) pairs, so label permutation yields an exact null distribution for any weight-generating mechanism.
    Invoked in Section 2.5; the paper claims no assumption about weight generation, but exchangeability is exactly such an assumption and may fail when weights are IPTW-type functions of treatment and covariates.
  • ad hoc to paper The Kish effective sample size n_e is the correct standardization for the weighted Anderson–Darling statistic under the null.
    Introduced in Eqs. (8)-(9) as 'a principled correction' without derivation; the paper rejects design-based variance formulas but does not prove that n_e yields the null variance of the weighted ECDF difference.
  • domain assumption The discrete AD formulation of Pettitt (1976) and Scholz-Stephens (1987), including the raw tie multiplier tau_k, carries over to weighted ECDFs.
    Used in Eq. (9); ties are multiplied by unweighted counts, which may not reflect weighted information at tied values.
  • domain assumption Weights are strictly positive and fixed attributes of observations.
    Stated in Section 2.1; excludes zero or negative weights and any weight that changes under permutation.

pith-pipeline@v1.3.0-alltime-deepseek · 14179 in / 13867 out tokens · 145292 ms · 2026-08-01T06:42:01.253579+00:00 · methodology

0 comments
read the original abstract

Assessing covariate balance is a core diagnostic step in causal inference, but commonly used summary measures can miss meaningful distributional differences they are not designed to detect. Distributional goodness-of-fit tests, including the Kolmogorov-Smirnov (KS), Anderson-Darling (AD), and Cramer-von Mises (CVM) tests, offer a more complete comparison but have previously been available only for unweighted data. We extend all three to accommodate case weights of any origin, using a shared label-permutation inference procedure that requires no assumption about how the weights were generated. In a four-scenario simulation study, all three weighted tests controlled Type I error close to nominal across sample sizes from 1,000 to 4,000 under substantial weight variability. Each test's known unweighted comparative advantage was preserved under weighting for two of three discrepancy types: KS was most powerful against a centrally located discrepancy, and AD was overwhelmingly most powerful against a tail-located discrepancy, while AD and CVM performed comparably against a diffuse discrepancy, both outperforming KS. These findings support AD as a reasonable general-purpose default for routine covariate balance assessment, while KS retains an advantage when a centrally concentrated imbalance is specifically suspected. The methods are implemented in the Stata commands kstest, adtest, and cvmtest.

Figures

Figures reproduced from arXiv: 2607.21782 by Ariel Linden.

Figure 1
Figure 1. Figure 1: Weighted empirical CDFs of BMI, Coached vs. Controls, following inverse [PITH_FULL_IMAGE:figures/full_fig_p034_1.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

59 extracted references

  1. [1]

    Anderson, T. W. and Darling, D. A. , title =. Annals of Mathematical Statistics , year =

  2. [2]

    Scandinavian Actuarial Journal , year =

    Cram\'er, Harald , title =. Scandinavian Actuarial Journal , year =

  3. [3]

    Goodness-of-Fit Techniques , publisher =

  4. [4]

    , title =

    Fleiss, Joseph L. , title =

  5. [5]

    Kish, Leslie , title =

  6. [6]

    , title =

    Kolmogorov, A. , title =. Giornale dell'Istituto Italiano degli Attuari , year =

  7. [7]

    Pettitt, A. N. , title =. Biometrika , year =

  8. [8]

    Scholz, F. W. and Stephens, M. A. , title =. Journal of the American Statistical Association , year =

  9. [9]

    , title =

    Smirnov, N. , title =. Annals of Mathematical Statistics , year =

  10. [10]

    Stephens, M. A. , title =. Journal of the American Statistical Association , year =

  11. [11]

    von Mises, Richard , title =

  12. [12]

    , title =

    Austin, Peter C. , title =. Statistics in Medicine , year =

  13. [13]

    Stuart, E. A. , title =. Statistical Science , year =

  14. [14]

    Rubin, D. B. , title =. Biometrics , year =

  15. [15]

    Rubin, D. B. , title =. Annals of Applied Statistics , year =

  16. [16]

    and Samuels, S.J

    Linden, A. and Samuels, S.J. , title =. Journal of Evaluation in Clinical Practice , year =

  17. [17]

    and Yarnold, P.R

    Linden, A. and Yarnold, P.R. , title =. Journal of Evaluation in Clinical Practice , year =

  18. [18]

    and Falcaro, M

    Newson, R.B. and Falcaro, M. , title =. The Stata Journal , year =

  19. [19]

    2025 , note =

    Linden, Ariel , title =. 2025 , note =

  20. [20]

    2026 , note =

    Linden, Ariel , title =. 2026 , note =

  21. [21]

    and Prochaska, James O

    Linden, Ariel and Butterworth, Susan W. and Prochaska, James O. , title =. Journal of Evaluation in Clinical Practice , year =

  22. [22]

    and Hernan, M.A

    Robins, J.M. and Hernan, M.A. and Brumback, B. , title =. Epidemiology , year =

  23. [23]

    Journal of Evaluation in Clinical Practice , year =

    Linden, A , title =. Journal of Evaluation in Clinical Practice , year =

  24. [24]

    Political Analysis , year =

    Hainmueller, J , title =. Political Analysis , year =

  25. [25]

    Kuiper, N. H. , title =. Proceedings of the Koninklijke Nederlandse Akademie van Wetenschappen, Series A , year =

  26. [26]

    and Garcia, N

    Ramdas, A. and Garcia, N. and Cuturi, M. , title =. Entropy , year =

  27. [27]

    and Roberts, N

    Linden, A. and Roberts, N. , title =. American Journal of Managed Care , year =

  28. [28]

    2003 , month =

    Linden, Ariel and Adams, John and Roberts, Nancy , title =. 2003 , month =

  29. [29]

    Anderson, T. W. and Darling, D. A. (1952). Asymptotic theory of certain ``goodness of fit'' criteria based on stochastic processes. Annals of Mathematical Statistics , 23(2):193--212

  30. [30]

    Austin, P. C. (2009). Balance diagnostics for comparing the distribution of baseline covariates between treatment groups in propensity-score matched samples. Statistics in Medicine , 28(25):3083--3107

  31. [31]

    Austin, P. C. (2015). Moving towards best practice when using inverse probability of treatment weighting (iptw) using the propensity score to estimate causal treatment effects in observational studies. Statistics in Medicine , 34(28):3661--3679

  32. [32]

    Cram\'er, H. (1928). On the composition of elementary errors. Scandinavian Actuarial Journal , 1928(1):13--74

  33. [33]

    D'Agostino, R. B. and Stephens, M. A., editors (1986). Goodness-of-Fit Techniques . Marcel Dekker, New York

  34. [34]

    Hainmueller, J. (2012). Entropy balancing: A multivariate reweighting method to produce balanced samples in observational studies. Political Analysis , 20:25--46

  35. [35]

    Kish, L. (1965). Survey Sampling . John Wiley & Sons, New York

  36. [36]

    Kolmogorov, A. (1933). Sulla determinazione empirica di una legge di distribuzione. Giornale dell'Istituto Italiano degli Attuari , 4:83--91

  37. [37]

    Kuiper, N. H. (1960). Tests concerning random points on a circle. Proceedings of the Koninklijke Nederlandse Akademie van Wetenschappen, Series A , 63:38--47

  38. [38]

    Linden, A. (2014). Combining propensity score-based stratification and weighting to improve causal inference in the evaluation of health care interventions. Journal of Evaluation in Clinical Practice , 20:1065--1071

  39. [39]

    Linden, A. (2025a). ADTEST : Stata module to perform a two-sample anderson-darling equality-of-distributions test. Statistical Software Components S459559, Boston College Department of Economics

  40. [40]

    Linden, A. (2025b). CVMTEST : Stata module to perform a two-sample cramer-von mises equality-of-distributions test. Statistical Software Components S459558, Boston College Department of Economics

  41. [41]

    Linden, A. (2026). KSTEST : Stata module to perform a weighted two-sample kolmogorov-smirnov equality-of-distributions test. Statistical Software Components s459801, Boston College Department of Economics

  42. [42]

    Linden, A., Adams, J., and Roberts, N. (2003). Evaluation methods in disease management: determining program effectiveness. Position Paper for the Disease Management Association of America (DMAA)

  43. [43]

    W., and Prochaska, J

    Linden, A., Butterworth, S. W., and Prochaska, J. O. (2010). Motivational interviewing-based health coaching as a chronic care intervention. Journal of Evaluation in Clinical Practice , 16(1):166--174

  44. [44]

    and Roberts, N

    Linden, A. and Roberts, N. (2005). A user’s guide to the disease management literature: recommendations for reporting and assessing program outcomes. American Journal of Managed Care , 11:113--120

  45. [45]

    and Samuels, S

    Linden, A. and Samuels, S. (2013). Using balance statistics to determine the optimal number of controls in matching studies. Journal of Evaluation in Clinical Practice , 19(5):968--975

  46. [46]

    and Yarnold, P

    Linden, A. and Yarnold, P. (2016). Using machine learning to assess covariate balance in matching studies. Journal of Evaluation in Clinical Practice , 22:848--854

  47. [47]

    and Yarnold, P

    Linden, A. and Yarnold, P. (2017). Using classification tree analysis to generate propensity score weights. Journal of Evaluation in Clinical Practice , 23:703--712

  48. [48]

    and Yarnold, P

    Linden, A. and Yarnold, P. (2018). Estimating causal effects for survival (time-to-event) outcomes by combining classification tree analysis and propensity score weighting. Journal of Evaluation in Clinical Practice , 24:380--387

  49. [49]

    and Falcaro, M

    Newson, R. and Falcaro, M. (2023). Robit regression in stata. The Stata Journal , 3:658--682

  50. [50]

    Pettitt, A. N. (1976). A two-sample A nderson-- D arling rank statistic. Biometrika , 63(1):161--168

  51. [51]

    Ramdas, A., Garcia, N., and Cuturi, M. (2017). On W asserstein two-sample testing and related families of nonparametric tests. Entropy , 19:47

  52. [52]

    Robins, J., Hernan, M., and Brumback, B. (2000). Marginal structural models and causal inference in epidemiology. Epidemiology , 11:550--560

  53. [53]

    Rubin, D. B. (1973). Matching to remove bias in observational studies. Biometrics , 29:159--184

  54. [54]

    Rubin, D. B. (2008). For objective causal inference, design trumps analysis. Annals of Applied Statistics , 2:808--840

  55. [55]

    Scholz, F. W. and Stephens, M. A. (1987). K-sample A nderson-- D arling tests. Journal of the American Statistical Association , 82(399):918--924

  56. [56]

    Smirnov, N. (1948). Table for estimating the goodness of fit of empirical distributions. Annals of Mathematical Statistics , 19(2):279--281

  57. [57]

    Stephens, M. A. (1974). EDF statistics for goodness of fit and some comparisons. Journal of the American Statistical Association , 69(347):730--737

  58. [58]

    Stuart, E. A. (2010). Matching methods for causal inference: a review and a look forward. Statistical Science , 25(1):1--21

  59. [59]

    von Mises, R. (1931). Wahrscheinlichkeitsrechnung und ihre Anwendung in der Statistik und theoretischen Physik . Deuticke, Leipzig