REVIEW 4 major objections 4 minor 59 references
This paper extends the Kolmogorov-Smirnov, Cramer-von Mises, and Anderson-Darling tests to weighted samples and shows, through a single label-permutation calibration, that all three keep false-positive rates near nominal while preserving th
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 06:42 UTC pith:2F5TJNFJ
load-bearing objection Useful weighted ECDF tests for balance diagnostics, but the 'works for any weights' claim is only validated for exogenous weights—under IPTW the permutation null is not obviously correct. the 4 major comments →
Weighted Extensions of the Kolmogorov-Smirnov, Cramer-von Mises, and Anderson-Darling Tests for Assessing Covariate Balance
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper's central claim is that the Kolmogorov-Smirnov, Cramer-von Mises, and Anderson-Darling statistics can be lifted to any positive case-weight setting by replacing ordinary empirical CDFs with weighted empirical CDFs, and in the Anderson-Darling statistic by replacing the sample size N with the Kish effective sample size inside the variance-standardization term. All three statistics are then calibrated with one label-permutation procedure that shuffles group labels over fixed (value, weight) pairs. The paper argues, and supports by simulation, that this yields tests whose false-positive rates stay near nominal under substantial weight variability, that the classical sensitivity orderi
What carries the argument
The apparatus is (1) the weighted empirical CDF, where each observation contributes w_i/sum(w) instead of 1/n; (2) the Kish effective sample size inserted into AD's variance standardization, which prevents the weighted null variance from being understated when weights vary; and (3) the permutation p-value, computed as the proportion of relabeled datasets whose statistic equals or exceeds the observed one. These pieces make the weighted statistics reduce exactly to their classical forms when all weights are equal, and they make the inference independent of how the weights were generated.
Load-bearing premise
The load-bearing premise is that valid inference follows from permuting treatment labels across fixed (value, weight) pairs under the null; this exchangeability can break when weights are themselves estimated from treatment and covariates, because each weight already encodes information about the label being shuffled.
What would settle it
Generate a dataset in which treatment depends on a covariate through a logistic model, construct inverse-propensity weights, and then apply the three permutation tests to a covariate that is truly balanced after weighting; if the tests reject more often than nominal across sample sizes, the claim that calibration is independent of how weights were generated is falsified.
If this is right
- Applied researchers can now run full-distribution balance checks on weighted data, rather than relying only on standardized mean differences and variance ratios.
- A single permutation calibration works across all three statistics and any positive weighting scheme, so no separate null-distribution formula is needed for each weighting method.
- The known sensitivity profiles carry over: KS is the most powerful test for a centrally located imbalance, AD is dramatically the best for tail imbalance, and AD and CVM are comparably good (and both better than KS) for a diffuse shift.
- The paper's recommendation is to use AD as the routine default for balance diagnostics, to prefer KS when a central imbalance is specifically suspected, and to consider reporting AD and KS together to cover both blind spots.
- At very severe weight variability (effective sample size ratio 0.2), all three tests show a modest false-positive inflation—about 0.064 to 0.067 at nominal 0.05—so calibration is not perfectly invariant to extreme weights.
Where Pith is reading between the lines
- If weights are estimated from treatment and covariates, as with inverse-propensity weighting, each unit's weight already encodes its treatment propensity; the fixed-weight permutation null may then not match the true null. A direct simulation with treatment-dependent weights would settle whether the 'no assumption about weight origin' claim holds beyond exogenously assigned weights.
- The diffuse-scenario result suggests AD and CVM coincide near a uniform shift because AD's variance standardization has nothing to emphasize. Varying the discrepancy's functional form or the base distribution's shape could map the boundary between the tests' regimes.
- Because every formula uses weights only in ratios, the tests are invariant to rescaling weights; this likely makes them applicable to survey weights or post-stratification weights without any extra adjustment.
- The same weighted-ECDF-plus-permutation framework could be extended immediately to the Kuiper or Wasserstein statistics, which the paper itself names as future directions, adding tests sensitive to cyclical or cost-scaled discrepancies.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript extends the two-sample Kolmogorov-Smirnov, Cramér–von Mises, and Anderson-Darling tests to case-weighted data for covariate-balance assessment. It defines weighted ECDFs and weighted statistics, and calibrates all three using a single label-permutation procedure that treats each observation's weight as a fixed attribute. The simulation study uses exogenous lognormal weights at sample sizes 1,000–4,000 and reports Type I error near nominal and preserved comparative advantages (KS for central, AD for tail, AD/CVM for diffuse discrepancies). An IPTW application to BMI is presented, and Stata implementations are referenced. The paper recommends AD as a reasonable general-purpose default.
Significance. If the permutation calibration is valid for weights of the kind used in practice, the paper fills a practical gap: applied covariate-balance assessment after IPTW or entropy balancing currently relies mostly on moment-based summaries, and weighted distributional tests are not routinely available. The weighted KS and CVM forms follow naturally from weighted ECDFs, the shared permutation inference is transparent and reduces to the classical unweighted test, and the simulation includes paired power comparisons and a weight-variability sensitivity check. These are genuine strengths. However, the central generality claim — that the inference 'requires no assumption about how the weights were generated' — is not supported by the simulations, which assign weights exogenously and independently of treatment; the applied example uses IPTW weights. The AD_w standardization through Kish effective sample size is asserted rather than derived. These points affect the validity of the headline claims and require revision.
major comments (4)
- [Sec. 2.5 / Sec. 2.6.1] The claim that the label-permutation procedure 'requires no assumption about how the weights were generated' is broader than what is shown. Permuting labels among fixed (value, weight) pairs generates a null under which (X, W) are jointly exchangeable across treatment labels. For IPTW or entropy-balancing weights, W is a function of G and X; even when weighted covariate distributions are equal, (X, W) is typically not exchangeable with G. The simulation (Eq. 12) assigns weights independently of y, g, and p, so Tables 2 and 5 validate the procedure only under exogenous weights. The applied example (Sec. 4.1) uses IPTW weights, so the p-values in Table 6 inherit this unvalidated assumption. Please either prove/examine validity for estimated weights, add simulations with estimated propensity or entropy weights, or explicitly restrict the claim.
- [Sec. 2.4, Eqs. (7)-(9)] The Anderson-Darling standardization is not derived. Eq. (7) contains an ambiguous factor 'N·2' (and Eq. (9) an analogous 'ne·2') whose origin is unclear; the usual null variance of the ECDF difference involves a term like D(1-D)(1/n1+1/n0), not an unexplained N×2 constant. Since permutation p-values are invariant to a common multiplicative constant, this ambiguity does not invalidate the reported p-values, but it affects the reported statistic and the interpretation of AD_w. More importantly, replacing N by the Kish effective sample size n_e (Eq. 8) is asserted as 'the standard adjustment' without derivation; this choice changes the relative weighting across x and should be justified or supported by simulation under estimated weights.
- [Table 5 / Sec. 3.6] At the most severe weight variability examined (r_e = 0.2), all three tests show Type I error rates of .064-.067, a relative inflation of roughly 30% over nominal .05. The text acknowledges this but the abstract and Sec. 5.1 conclude that all three tests 'controlled Type I error close to nominal under substantial weight variability' without stating the boundary of validity. If r_e = 0.2 is within the intended scope, the claim should be qualified; if it is outside, the qualifying condition should be stated. As written, the conclusion is stronger than the table supports.
- [Secs. 2.6, 5.3] The recommendation of AD as a general-purpose default is broader than the simulation evidence. The DGP is normal throughout, each discrepancy type is examined at one fixed effect size, and the scenarios are deliberately constructed to correspond to the known theoretical strengths of the three tests (Sec. 2.6). Such scenarios can show that weighting preserves known comparative advantages under exogenous weights, but they do not establish that AD is the best default under realistic imbalance structures, non-normal covariate distributions, or estimated weights. Please temper the applied recommendation or broaden the simulation accordingly.
minor comments (4)
- [Eq. (7)] The formula is difficult to parse: the placement of the 'N·2' factor and the square root is unclear. Please rewrite with explicit numerators/denominators and define all symbols at first use.
- [Sec. 2.4.2 / Sec. 5.3] The tie multiplier tau_k is retained as the raw unweighted count. For weighted data this is not obviously the right information measure; the caveat in Sec. 5.3 is useful, but the property is not examined in the simulations.
- [Introduction] The citation for Somers' D (Newson and Falcaro, 2023) points to a reference titled 'Robit regression in Stata', which appears unrelated. Please verify the reference.
- [Sec. 2.5] The definition of the permutation p-value (Eq. 10) uses '#' rather than an explicit count and could be made clearer by writing (1 + sum_j 1(T_j^* >= T_obs))/(R+1).
Circularity Check
No significant circularity: the weighted statistics are direct extensions, and the simulation results are empirical rather than built into the definitions.
full rationale
The paper's derivation is self-contained. The weighted KS, CVM, and AD statistics are obtained by substituting weighted ECDFs (and the Kish effective sample size for the AD variance term) into classical unweighted definitions; each reduces to the unweighted statistic when weights are equal, but this is a consistency property of the extension, not a restatement of the target conclusion. The permutation procedure is a standard label-permutation calibration applied to whatever statistic is supplied, and it does not define the test's conclusions. The simulation outcomes could have failed—for example, weighting could have altered the comparative ordering—and the paper in fact reports a partial non-confirmation for the diffuse scenario, since CVM does not uniquely dominate AD. This indicates the results are not forced by construction. The self-citations to the author's own Stata commands are implementation references, not load-bearing evidence for the statistical claims. The main scientific weakness—that the simulation weights are generated independently of treatment and covariates, so the 'no assumption about how the weights were generated' claim is not validated for estimated IPTW weights—is a validity and external-generalizability concern, not a circularity: the permutation p-value is not defined in terms of the conclusion it is used to support, and no specific equation or fitted parameter is renamed as a prediction.
Axiom & Free-Parameter Ledger
free parameters (4)
- Delta_center =
6
- Delta_tail =
50
- delta_diffuse =
1
- r_e (Kish effective sample size ratio) =
0.6 (main); 0.2, 0.4, 0.8, 0.9 (sensitivity)
axioms (4)
- domain assumption Under H0, group labels are exchangeable with respect to (value, weight) pairs, so label permutation yields an exact null distribution for any weight-generating mechanism.
- ad hoc to paper The Kish effective sample size n_e is the correct standardization for the weighted Anderson–Darling statistic under the null.
- domain assumption The discrete AD formulation of Pettitt (1976) and Scholz-Stephens (1987), including the raw tie multiplier tau_k, carries over to weighted ECDFs.
- domain assumption Weights are strictly positive and fixed attributes of observations.
read the original abstract
Assessing covariate balance is a core diagnostic step in causal inference, but commonly used summary measures can miss meaningful distributional differences they are not designed to detect. Distributional goodness-of-fit tests, including the Kolmogorov-Smirnov (KS), Anderson-Darling (AD), and Cramer-von Mises (CVM) tests, offer a more complete comparison but have previously been available only for unweighted data. We extend all three to accommodate case weights of any origin, using a shared label-permutation inference procedure that requires no assumption about how the weights were generated. In a four-scenario simulation study, all three weighted tests controlled Type I error close to nominal across sample sizes from 1,000 to 4,000 under substantial weight variability. Each test's known unweighted comparative advantage was preserved under weighting for two of three discrepancy types: KS was most powerful against a centrally located discrepancy, and AD was overwhelmingly most powerful against a tail-located discrepancy, while AD and CVM performed comparably against a diffuse discrepancy, both outperforming KS. These findings support AD as a reasonable general-purpose default for routine covariate balance assessment, while KS retains an advantage when a centrally concentrated imbalance is specifically suspected. The methods are implemented in the Stata commands kstest, adtest, and cvmtest.
Figures
Reference graph
Works this paper leans on
-
[1]
Anderson, T. W. and Darling, D. A. , title =. Annals of Mathematical Statistics , year =
-
[2]
Scandinavian Actuarial Journal , year =
Cram\'er, Harald , title =. Scandinavian Actuarial Journal , year =
-
[3]
Goodness-of-Fit Techniques , publisher =
-
[4]
, title =
Fleiss, Joseph L. , title =
-
[5]
Kish, Leslie , title =
-
[6]
, title =
Kolmogorov, A. , title =. Giornale dell'Istituto Italiano degli Attuari , year =
-
[7]
Pettitt, A. N. , title =. Biometrika , year =
-
[8]
Scholz, F. W. and Stephens, M. A. , title =. Journal of the American Statistical Association , year =
-
[9]
, title =
Smirnov, N. , title =. Annals of Mathematical Statistics , year =
-
[10]
Stephens, M. A. , title =. Journal of the American Statistical Association , year =
-
[11]
von Mises, Richard , title =
-
[12]
, title =
Austin, Peter C. , title =. Statistics in Medicine , year =
-
[13]
Stuart, E. A. , title =. Statistical Science , year =
-
[14]
Rubin, D. B. , title =. Biometrics , year =
-
[15]
Rubin, D. B. , title =. Annals of Applied Statistics , year =
-
[16]
and Samuels, S.J
Linden, A. and Samuels, S.J. , title =. Journal of Evaluation in Clinical Practice , year =
-
[17]
and Yarnold, P.R
Linden, A. and Yarnold, P.R. , title =. Journal of Evaluation in Clinical Practice , year =
-
[18]
and Falcaro, M
Newson, R.B. and Falcaro, M. , title =. The Stata Journal , year =
-
[19]
2025 , note =
Linden, Ariel , title =. 2025 , note =
2025
-
[20]
2026 , note =
Linden, Ariel , title =. 2026 , note =
2026
-
[21]
and Prochaska, James O
Linden, Ariel and Butterworth, Susan W. and Prochaska, James O. , title =. Journal of Evaluation in Clinical Practice , year =
-
[22]
and Hernan, M.A
Robins, J.M. and Hernan, M.A. and Brumback, B. , title =. Epidemiology , year =
-
[23]
Journal of Evaluation in Clinical Practice , year =
Linden, A , title =. Journal of Evaluation in Clinical Practice , year =
-
[24]
Political Analysis , year =
Hainmueller, J , title =. Political Analysis , year =
-
[25]
Kuiper, N. H. , title =. Proceedings of the Koninklijke Nederlandse Akademie van Wetenschappen, Series A , year =
-
[26]
and Garcia, N
Ramdas, A. and Garcia, N. and Cuturi, M. , title =. Entropy , year =
-
[27]
and Roberts, N
Linden, A. and Roberts, N. , title =. American Journal of Managed Care , year =
-
[28]
2003 , month =
Linden, Ariel and Adams, John and Roberts, Nancy , title =. 2003 , month =
2003
-
[29]
Anderson, T. W. and Darling, D. A. (1952). Asymptotic theory of certain ``goodness of fit'' criteria based on stochastic processes. Annals of Mathematical Statistics , 23(2):193--212
1952
-
[30]
Austin, P. C. (2009). Balance diagnostics for comparing the distribution of baseline covariates between treatment groups in propensity-score matched samples. Statistics in Medicine , 28(25):3083--3107
2009
-
[31]
Austin, P. C. (2015). Moving towards best practice when using inverse probability of treatment weighting (iptw) using the propensity score to estimate causal treatment effects in observational studies. Statistics in Medicine , 34(28):3661--3679
2015
-
[32]
Cram\'er, H. (1928). On the composition of elementary errors. Scandinavian Actuarial Journal , 1928(1):13--74
1928
-
[33]
D'Agostino, R. B. and Stephens, M. A., editors (1986). Goodness-of-Fit Techniques . Marcel Dekker, New York
1986
-
[34]
Hainmueller, J. (2012). Entropy balancing: A multivariate reweighting method to produce balanced samples in observational studies. Political Analysis , 20:25--46
2012
-
[35]
Kish, L. (1965). Survey Sampling . John Wiley & Sons, New York
1965
-
[36]
Kolmogorov, A. (1933). Sulla determinazione empirica di una legge di distribuzione. Giornale dell'Istituto Italiano degli Attuari , 4:83--91
1933
-
[37]
Kuiper, N. H. (1960). Tests concerning random points on a circle. Proceedings of the Koninklijke Nederlandse Akademie van Wetenschappen, Series A , 63:38--47
1960
-
[38]
Linden, A. (2014). Combining propensity score-based stratification and weighting to improve causal inference in the evaluation of health care interventions. Journal of Evaluation in Clinical Practice , 20:1065--1071
2014
-
[39]
Linden, A. (2025a). ADTEST : Stata module to perform a two-sample anderson-darling equality-of-distributions test. Statistical Software Components S459559, Boston College Department of Economics
-
[40]
Linden, A. (2025b). CVMTEST : Stata module to perform a two-sample cramer-von mises equality-of-distributions test. Statistical Software Components S459558, Boston College Department of Economics
-
[41]
Linden, A. (2026). KSTEST : Stata module to perform a weighted two-sample kolmogorov-smirnov equality-of-distributions test. Statistical Software Components s459801, Boston College Department of Economics
2026
-
[42]
Linden, A., Adams, J., and Roberts, N. (2003). Evaluation methods in disease management: determining program effectiveness. Position Paper for the Disease Management Association of America (DMAA)
2003
-
[43]
W., and Prochaska, J
Linden, A., Butterworth, S. W., and Prochaska, J. O. (2010). Motivational interviewing-based health coaching as a chronic care intervention. Journal of Evaluation in Clinical Practice , 16(1):166--174
2010
-
[44]
and Roberts, N
Linden, A. and Roberts, N. (2005). A user’s guide to the disease management literature: recommendations for reporting and assessing program outcomes. American Journal of Managed Care , 11:113--120
2005
-
[45]
and Samuels, S
Linden, A. and Samuels, S. (2013). Using balance statistics to determine the optimal number of controls in matching studies. Journal of Evaluation in Clinical Practice , 19(5):968--975
2013
-
[46]
and Yarnold, P
Linden, A. and Yarnold, P. (2016). Using machine learning to assess covariate balance in matching studies. Journal of Evaluation in Clinical Practice , 22:848--854
2016
-
[47]
and Yarnold, P
Linden, A. and Yarnold, P. (2017). Using classification tree analysis to generate propensity score weights. Journal of Evaluation in Clinical Practice , 23:703--712
2017
-
[48]
and Yarnold, P
Linden, A. and Yarnold, P. (2018). Estimating causal effects for survival (time-to-event) outcomes by combining classification tree analysis and propensity score weighting. Journal of Evaluation in Clinical Practice , 24:380--387
2018
-
[49]
and Falcaro, M
Newson, R. and Falcaro, M. (2023). Robit regression in stata. The Stata Journal , 3:658--682
2023
-
[50]
Pettitt, A. N. (1976). A two-sample A nderson-- D arling rank statistic. Biometrika , 63(1):161--168
1976
-
[51]
Ramdas, A., Garcia, N., and Cuturi, M. (2017). On W asserstein two-sample testing and related families of nonparametric tests. Entropy , 19:47
2017
-
[52]
Robins, J., Hernan, M., and Brumback, B. (2000). Marginal structural models and causal inference in epidemiology. Epidemiology , 11:550--560
2000
-
[53]
Rubin, D. B. (1973). Matching to remove bias in observational studies. Biometrics , 29:159--184
1973
-
[54]
Rubin, D. B. (2008). For objective causal inference, design trumps analysis. Annals of Applied Statistics , 2:808--840
2008
-
[55]
Scholz, F. W. and Stephens, M. A. (1987). K-sample A nderson-- D arling tests. Journal of the American Statistical Association , 82(399):918--924
1987
-
[56]
Smirnov, N. (1948). Table for estimating the goodness of fit of empirical distributions. Annals of Mathematical Statistics , 19(2):279--281
1948
-
[57]
Stephens, M. A. (1974). EDF statistics for goodness of fit and some comparisons. Journal of the American Statistical Association , 69(347):730--737
1974
-
[58]
Stuart, E. A. (2010). Matching methods for causal inference: a review and a look forward. Statistical Science , 25(1):1--21
2010
-
[59]
von Mises, R. (1931). Wahrscheinlichkeitsrechnung und ihre Anwendung in der Statistik und theoretischen Physik . Deuticke, Leipzig
1931
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.