REVIEW 3 major objections 5 minor 1 cited by
Covariate-adjusted win statistics in randomized clinical trials with ordinal outcomes
T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read Covariate-adjusted win statistics keep consistency and never increase asymptotic variance.
desk verdict A substantive extension of covariate-adjusted win statistics to ordinal RCT outcomes, with a plausible variance-reduction theorem; needs the web-appendix proofs made visible and the finite-sample variance caveats taken seriously. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the win estimand $\tau_1=P\{Y_i(1)>Y_j(0)\}$, the probability that a randomly chosen treated outcome beats a randomly chosen control outcome, alongside its loss counterpart and their ratio and difference. The machinery is the U-statistic representation of weighted pairwise comparisons: an estimator is a normalized sum over treatment–control pairs of a symmetric kernel that records which patient wins, with weights determined by the fitted propensity score. The load-bearing identity is Theorem 1(b): for any balancing-weight estimator, the influence function is the unadjusted influence function minus its projection on the tangent space of the propensity model, which is exactly what makes the adjusted asymptotic variance no larger. Closed-form variance estimates follow from taking the sample second moment of the estimated influence function, so inference does not require resampling.
What would settle it
Simulate a randomized trial with treatment probability $\pi=0.5$, a strong prognostic covariate, and a logistic propensity model that omits the intercept so it cannot represent $\pi$; if the IPW or OW win-ratio estimator shows bias that persists as $n$ grows to 10,000, the claimed model-robust consistency is false.
Extended reading notes
Core claim
On the paper's own terms, the central discovery is that the class of balancing-weight win estimators shares one influence-function structure: a normalized weighted comparison of every treatment–control pair, with weights built from a fitted propensity score. Under randomization, the weighted estimand reduces to the unadjusted win estimand whenever the propensity model can represent the constant treatment probability. The paper's Theorem 1 then shows the influence function of any such estimator equals the unadjusted influence function minus its projection on the propensity-score tangent space, so the asymptotic variance is no larger. Theorem 2 shows that adding an ordinal outcome regression through augmented weighting keeps consistency even when that regression is misspecified, and achieves local efficiency when it is correct. The win ratio and win difference inherit these properties by the delta method.
Load-bearing premise
The load-bearing premise is that the fitted treatment-assignment model can represent the true constant randomization probability; if it cannot, the weighted estimator targets a different quantity and the consistency claim does not hold.
Editorial extensions
If this is right
- Covariate adjustment for win ratio and win difference can be prespecified in a trial protocol with a guarantee of no asymptotic precision loss relative to the unadjusted analysis.
- Analysts can report confidence intervals from closed-form variance estimators rather than bootstrap resampling.
- Augmented weighting estimators provide an additional efficiency gain in most simulation settings, and remain consistent when the outcome regression is misspecified.
- Overlap weighting is the recommended default when sample sizes are small or randomization is unbalanced, because it tends to be at least as efficient as IPW and avoids extreme weights.
- The ORCHID reanalysis illustrates that adjustment can reduce standard errors by about 30–40% while leaving the clinical conclusion unchanged.
Reading between the lines
- Editorial inference: because the variance-reduction argument only uses the structure of pairwise comparisons and a propensity tangent space, the same projection identity should extend to time-to-event win ratios and to other pairwise estimands such as net benefit with ties handled explicitly.
- Editorial inference: the model robustness result is asymmetric: it protects against misspecifying the outcome model, not the propensity model, so in practice the propensity model must include an intercept or otherwise be able to represent the constant randomization probability.
- Editorial inference: the simulations show variance estimates falling below Monte Carlo variance when the augmented outcome model has many parameters at n around 200, so small-sample corrections or sample-splitting are a natural next test.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper develops covariate-adjusted estimators for win statistics (win ratio and win difference) with ordinal outcomes in randomized clinical trials. The authors propose inverse probability weighting (IPW), overlap weighting (OW), and their augmented versions (AIPW, AOW), building on U-statistic theory. They state two central theoretical results: (i) the influence function of any balancing-weighting estimator equals the unadjusted estimator's influence function minus its projection onto the propensity-score tangent space, implying no larger asymptotic variance; and (ii) all estimators are consistent for the target estimand even when working models are misspecified, i.e., model-robust. The paper provides closed-form variance estimators, simulation studies, an application to the ORCHID trial, and an R package.
Significance. If the results are correct, the paper fills a gap by formally extending covariate adjustment for average treatment effects to pairwise win estimands, with a clean influence-function projection argument and practical software. The simulation study is broad and the GitHub code enables reproducibility. However, the central model-robustness claim is overstated: for IPW and OW, consistency to the unadjusted win estimand requires the propensity model to be able to represent the constant propensity, a condition not stated in the abstract or theorems. The variance-reduction claim likewise relies on correct specification of the propensity model. These issues affect the paper's central message and need to be addressed with explicit conditions or revised claims.
major comments (3)
- [Abstract; Section 3.2, Eqs. (3.4)-(3.5); Section 7] The abstract claims that 'all of the covariate-adjusted estimators do not compromise consistency for the target estimand even when the associated working models are incorrectly specified.' This is not correct for the IPW estimator (3.4) and the OW estimator (3.5) unless the propensity-score model family contains the constant propensity. Under randomization the true propensity is a constant π; consistency to τ1 requires the estimated propensity score to converge to π so that the weights become constants. If the working model cannot represent a constant (e.g., a logistic model without an intercept when π≠0.5), the estimated weights converge to a nonconstant function and the estimator targets a different weighted estimand. Theorem 1 does not state this condition, and the proof in Web Appendix B is not available to verify. Please add the required assumption or qualify the model-robustness claim to apply only to outcome-model misspecification and to propensity models that contain the constant.
- [Section 3.3, Theorem 1(b); Section 7] The statement that the influence function φ(O) equals χh(O) minus its projection on the tangent space of the propensity model, and the consequent conclusion that 'the asymptotic variance of the propensity score weighting estimators is no larger than that of the unadjusted estimator,' presuppose that the propensity model is correctly specified (or at least that the probability limit of the estimated propensity is the true constant). If the model is misspecified, the projection is not an orthogonal projection in the sense needed for the variance decomposition, and the variance-reduction claim is not established. The theorem and the discussion should explicitly state this condition.
- [Section 3.3 (Theorem 1) and Section 4.3 (Theorem 2)] The proofs of Theorems 1 and 2 are stated to be in Web Appendices B and D, but these appendices are not included in the submitted manuscript. Because these theorems are load-bearing for the paper's central claims of consistency, variance reduction, and local efficiency, the full proofs must be provided to the reviewers.
minor comments (5)
- [Section 1] The word 'recommendped' in the paragraph about regulatory guidance should be 'recommended'.
- [Section 3.3] The phrase 'to obtain consistant variance estimators' contains a typo: 'consistant' should be 'consistent'. Similarly, Section 3.4 has 'asymtoptic' which should be 'asymptotic'.
- [Section 7] The final paragraph contains a sentence fragment: 'Our work thus illustrates the operational steps involved in estimating win estimands with covariate adjustment. tics as well as the una djusted win statistics in the winPSW R package...' This appears to be a corruption and should be rewritten.
- [Section 3.1] The statement 'under randomization, τ h 1 = τ1 as long as h(Xi,X j) is at most a function of covariates only through the propensity score' is potentially confusing because the true propensity score is constant under randomization; a nonconstant function of covariates is not a function of the propensity score. Please clarify what is intended.
- [General] The references to supporting material are inconsistent: the text mentions 'Web Appendix A-D' and 'Web Appendices A–E'. Please standardize.
Circularity Check
No material circularity: the variance-reduction claim follows from an influence-function projection identity, and self-citations are contextual rather than load-bearing.
full rationale
The derivation chain is self-contained rather than circular. Win estimands are defined independently in Section 2.1, and the unadjusted U-statistic estimator follows Bebu and Lachin (2016). The covariate-adjusted estimators in Section 3 are weighted versions of the same pairwise kernel. Theorem 1(a) obtains the influence function from U-statistic theory, and Theorem 1(b) is the semiparametric identity that a balancing-weighting estimator's influence function equals the unadjusted influence function minus its projection on the propensity-score tangent space. The variance inequality then follows from the orthogonality of that projection, not from fitting or renaming the target. Theorem 2's double robustness and local efficiency for the augmented estimators is standard semiparametric theory, and the simulation section validates the estimators against independently computed Monte Carlo truth, so the empirical results are not fitted to the target. The paper cites the authors' own prior work (e.g., Zeng et al. 2021, Li et al. 2018, Cao et al. 2025), but these citations are motivational and contextual; no load-bearing theorem or variance-reduction claim reduces to those citations, and the proofs are supplied in the paper or in standard external references such as Tsiatis (2006) and Mao (2018). One precision caveat is worth noting but is not circularity: the abstract's blanket 'model-robust' wording for IPW/OW implicitly assumes the propensity model family can represent the constant propensity (e.g., a logistic model with an intercept). If the model family cannot represent the constant, the weights do not collapse to constants and the estimator targets a different weighted estimand; this is a model-specification and correctness caveat, not a circular dependence of the derivation on its own outputs.
Assumptions & free parameters
assumptions (6)
- domain assumption Randomization: treatment assignment Z is independent of potential outcomes and baseline covariates, with constant propensity pi.
- domain assumption Positivity: 0 < pi < 1.
- domain assumption SUTVA (consistency and no interference): Y = Z Y(1) + (1-Z) Y(0).
- standard math Regularity conditions for U-statistics and U-processes: finite moments and Donsker-type conditions on the kernel classes.
- domain assumption Logistic propensity model includes an intercept so that the constant propensity pi is in the model family.
- domain assumption Double robustness: at least one of the propensity model or the outcome regression model is correctly specified.
Cite this review
Pith. "Pith review of Covariate-adjusted win statistics in randomized clinical trials with ordinal outcomes." pith.science (2026). https://pith.science/paper/P3J722VJ
@misc{pith2026250820349,
author = {Pith},
title = {Pith review of: Covariate-adjusted win statistics in randomized clinical trials with ordinal outcomes},
year = {2026},
howpublished = {\url{https://pith.science/paper/P3J722VJ}},
note = {Machine review of arXiv:2508.20349}
}
read the original abstract
Ordinal outcomes are common in clinical settings where they often represent increasing levels of disease progression or different levels of functional impairment. In this article, we focus on representing the average treatment effect for ordinal outcomes via intrinsic pairwise outcome comparisons captured through win estimands, such as the win ratio and win difference. Recognizing the value of baseline covariate adjustment toward enhanced precision, we first develop propensity score weighting estimators, including both inverse probability weighting (IPW) and overlap weighting (OW), tailored to estimating win estimands. Furthermore, we develop augmented weighting estimators that leverage an additional ordinal outcome regression to potentially improve efficiency over weighting alone. Leveraging the theory of U-statistics, we establish the asymptotic theory for all estimators, and derive closed-form variance estimators to support statistical inference. We also prove that all of the covariate-adjusted estimators do not compromise consistency for the target estimand even when the associated working models are incorrectly specified; hence these covariate-adjusted estimators are model-robust. Through simulations we demonstrate the enhanced efficiency of the weighted estimators over the unadjusted estimator, with the augmented weighting estimators showing a further improvement in efficiency except for extreme cases. Finally, we illustrate our proposed methods with the ORCHID trial, and implement our covariate adjustment methods in an R package winPSW.
Forward citations
Cited by 1 Pith paper
-
Estimation and Inference for Win Measures with Multiple Ordinal Endpoints Subject to Missingness
Develops and validates IPW and AIPW estimators for win measures with missing hierarchical ordinal endpoints, including variance estimation via influence functions.
Reference graph
Works this paper leans on
-
[1]
Arcones, M. and Gin \'e , E. (1993), Limit theorems for U-processes, Ann Prob\/ , 21, 1494--1542
work page 1993
-
[2]
Bebu, I. and Lachin, J. (2016), Large sample inference for a win ratio analysis of a composite outcome based on prioritized components, Biostatistics\/ , 17, 178--187
work page 2016
-
[3]
Benkeser, D., Díaz, I., Luedtke, A., Segal, J., Scharfstein, D., and Rosenblum, M. (2021), Improving precision and power in randomized trials for COVID-19 treatments using covariate adjustment, for binary, ordinal, and time-to-event outcomes, Biometrics\/ , 77, 1467--1481
work page 2021
-
[4]
Cao, Z., Ghazi, L., Mastrogiacomo, C., Forastiere, L., Wilson, F. P., and Li, F. (2025), Using overlap weights to address extreme propensity scores in estimating restricted mean counterfactual survival times, Am J Epidemiol\/ , 194, 2402--2411
work page 2025
-
[5]
Cheng, C., Li, F., Thomas, L., and Li, F. (2022), Addressing extreme propensity scores in estimating counterfactual survival functions via the overlap weights, Am J Epidemiol\/ , 191, 1140--1151
work page 2022
-
[6]
Cummings, M., Baldwin, M., Abrams, D., Jacobson, S., Meyer, B., Balough, E., Aaron, J., Claassen, J., Rabbani, L., Hastie, J., Hochman, B., Salazar-Schicchi, J., Yip, N., Brodie, D., and O'Donnell, M. (2020), Epidemiology, clinical course, and outcomes of critically ill adults with COVID-19 in New York City: a prospective cohort study, Lancet\/ , 395, 1763--1770
work page 2020
-
[7]
Evans, S., Knutsson, M., Amarenco, P., Albers, G., Bath, P., Denison, H., Ladenvall, P., Jonasson, J., Easton, J., Minematsu, K., Molina, C., Y, W., Wong, K., and Johnston, S. (2020), Methodologies for pragmatic and efficient assessment of benefits and harms: Application to the SOCRATES trial, Clin Trials\/ , 17, 617--626
work page 2020
-
[8]
Hoeffding, W. (1948), A class of statistics with asymptotically normal distribution, Ann Math Stat\/ , 19, 293--352
work page 1948
Show all 42 references
-
[9]
ICH (2021), ICH E9(R1) Statistical Principles for Clinical Trials: Addendum: Estimands and Sensitivity Analysis in Clinical Trials\/ , European Medicines Evaluation Agency
2021
-
[10]
and Siegerink, B
Klok, F. and Siegerink, B. (2023), Ordinal outcomes add value to clinical trials, Lancet\/ , 401, 995
2023
-
[11]
and Li, F
Li, F. and Li, F. (2019), Propensity score weighting for causal inference with multiple treatments, Ann Appl Stat\/ , 13, 2389--2415
2019
-
[12]
(2018), Balancing covariate via propensity score weighting, J Am Stat Assoc\/ , 113, 390--400
Li, F., Morgan, K., and Zaslavsky, A. (2018), Balancing covariate via propensity score weighting, J Am Stat Assoc\/ , 113, 390--400
2018
-
[13]
(2019), Addressing extreme propensity scores via the overlap weights, Am J Epidemiol\/ , 188, 250--257
Li, F., Thomas, L., and Li, F. (2019), Addressing extreme propensity scores via the overlap weights, Am J Epidemiol\/ , 188, 250--257
2019
-
[14]
(2019), Propensity score weighting analysis and treatment effect discovery, Stat Methods Med Res\/ , 28, 2439--2454
Mao, H., Li, L., and Greene, T. (2019), Propensity score weighting analysis and treatment effect discovery, Stat Methods Med Res\/ , 28, 2439--2454
2019
-
[15]
(2018), On the propensity score weighting analysis with survival outcome: Estimands, estimation, and inference, Stat Med\/ , 37, 3745--3763
Mao, H., Li, L., Yang, W., and Shen, Y. (2018), On the propensity score weighting analysis with survival outcome: Estimands, estimation, and inference, Stat Med\/ , 37, 3745--3763
2018
-
[16]
(2018), On causal estimation using U-statistics, Biometrika\/ , 105, 215--220
Mao, L. (2018), On causal estimation using U-statistics, Biometrika\/ , 105, 215--220
2018
-
[17]
(1923), On the application of probability theory to agricultural experiments: essay on principles, section 9
Neyman, J. (1923), On the application of probability theory to agricultural experiments: essay on principles, section 9. Masters Thesis portions translated into English by D Dabrowska and T Speed (1990), Stat Sci\/ , 5, 465--472
1923
-
[18]
(2017), Analysis of an ordinal endpoint for use in evaluating treatments for severe influenza requiring hospitalization, Clin Trials\/ , 14, 264--276
Peterson, R., Vock, D., Powers, J., Emery, S., Cruz, E., Hunsberger, S., Jain, M., Pett, S., and Neaton, J. (2017), Analysis of an ordinal endpoint for use in evaluating treatments for severe influenza requiring hospitalization, Clin Trials\/ , 14, 264--276
2017
-
[19]
(2012), The win ratio: a new approach to the analysis of composite endpoints in clinical trials based on clinical priorities, Eur Heart J\/ , 33, 176--182
Pocock, S., Ariti, C., Collier, T., and Wang, D. (2012), The win ratio: a new approach to the analysis of composite endpoints in clinical trials based on clinical priorities, Eur Heart J\/ , 33, 176--182
2012
-
[20]
and Rubin, D
Rosenbaum, P. and Rubin, D. (1983), The central role of the propensity score in observational studies for causal effects, Biometrika\/ , 70, 41--55
1983
-
[21]
H., Semler, M
Self, W. H., Semler, M. W., Leither, L. M., Casey, J. D., Angus, D. C., Brower, R. G., Chang, S. Y., Collins, S. P., Eppensteiner, J. C., Filbin, M. R., et al. (2020), Effect of hydroxychloroquine on clinical status at 14 days in hospitalized patients with COVID-19: a randomiz...
2020
-
[22]
(2023), Statistical analyses of ordinal outcomes in randomised controlled trials: protocol for a scoping review, Trials\/ , 24, 286
Selman, C., Lee, K., Whitehead, C., Manley, B., and Mahar, R. (2023), Statistical analyses of ordinal outcomes in randomised controlled trials: protocol for a scoping review, Trials\/ , 24, 286
2023
-
[23]
(2014), Inverse probability weighting for covariate adjustment in randomized studies, Stat Med\/ , 33, 555--568
Shen, C., Li, X., and Li, L. (2014), Inverse probability weighting for covariate adjustment in randomized studies, Stat Med\/ , 33, 555--568
2014
-
[24]
(2020), Using propensity score methods to create target populations in observational clinical research, J Am Med Assoc\/ , 323, 466--467
Thomas, L., Li, F., and Pencina, M. (2020), Using propensity score methods to create target populations in observational clinical research, J Am Med Assoc\/ , 323, 466--467
2020
-
[25]
(2006), Semiparametric Theory and Missing Data\/ , New York: Springer
Tsiatis, A. (2006), Semiparametric Theory and Missing Data\/ , New York: Springer
2006
-
[26]
(2017), Review of recent methodological developments in group-randomized trials: part 1--design, Am J Publ Health\/ , 107, 907--915
Turner, E., Li, F., Gallis, J., Prague, M., and Murray, D. (2017), Review of recent methodological developments in group-randomized trials: part 1--design, Am J Publ Health\/ , 107, 907--915
2017
-
[27]
Z., and Kahan, B
Uddin, M., Bashir, N. Z., and Kahan, B. C. (2024), Evaluating whether the proportional odds models to analyse ordinal outcomes in COVID-19 clinical trials is providing clinically interpretable treatment effects: A systematic review, Clin Trials\/ , 21, 363--370
2024
-
[28]
U.S. Food and Drug Administration (2023 a ), Adjusting for Covariates in Randomized Clinical Trials for Drugs and Biologics Products , https://www.fda.gov/regulatory-information/search-fda-guidance-documents/adjusting-covariates-randomized-clinical-trials-drugs-and-biological-products
2023
-
[29]
Guidance for Industry , https://www.fda.gov/regulatory-information/search-fda-guidance-documents/covid-19-developing-drugs-and-biological-products-treatment-or-prevention
--- (2023 b ), COVID-19: Developing Drugs and Biological Products for Treatment or Prevention. Guidance for Industry , https://www.fda.gov/regulatory-information/search-fda-guidance-documents/covid-19-developing-drugs-and-biological-products-treatment-or-prevention
2023
-
[30]
and Wellner, J
Van Der Vaart, A. and Wellner, J. (1996), Weak Convergence and Empirical Processes\/ , New York, NY: Springer
1996
-
[31]
Venables, W. N. and Ripley, B. D. (2002), Modern Applied Statistics with S\/ , New York: Springer, fourth edition, ://www.stats.ox.ac.uk/pub/MASS4/. ISBN 0-387-95457-0
2002
-
[32]
(2015), Increasing the power of the Mann-Whitney test in randomized experiments through flexible covariate adjustment, Stat Med\/ , 34, 1012--1030
Vermeulen, K., Thas, O., and Vansteelandt, S. (2015), Increasing the power of the Mann-Whitney test in randomized experiments through flexible covariate adjustment, Stat Med\/ , 34, 1012--1030
2015
-
[33]
(2023 a ), Model-robust inference for clinical trials that improve precision by stratified randomization and covariate adjustment, J Am Stat Assoc\/ , 118, 1152--1163
Wang, B., Susukida, R., Mojtabai, R., Amin-Esmaeili, M., and Rosenblum, M. (2023 a ), Model-robust inference for clinical trials that improve precision by stratified randomization and covariate adjustment, J Am Stat Assoc\/ , 118, 1152--1163
2023
-
[34]
and Pocock, S
Wang, D. and Pocock, S. (2016), A win ratio approach to comparing continuous non-normal outcomes in clinical trials, Pharm Stat\/ , 15, 238--245
2016
-
[35]
(2023 b ), Adjusted win ratio using the inverse probability of treatment weighting, J Biopharm Stat\/ , 35, 21--36
Wang, D., Zheng, S., Cui, Y., He, N., Chen, T., and Huang, B. (2023 b ), Adjusted win ratio using the inverse probability of treatment weighting, J Biopharm Stat\/ , 35, 21--36
2023
-
[36]
(2014), Variance reduction in randomised trials by inverse probability weighting using the propensity score, Stat Med\/ , 33, 721--737
Williams, E., Forbes, A., and White, I. (2014), Variance reduction in randomised trials by inverse probability weighting using the propensity score, Stat Med\/ , 33, 721--737
2014
-
[37]
(2022), Trial of Erythropoietin for Hypoxic-Ischemic Encephalopathy in Newborns, N Engl J Med\/ , 387, 148--159
Wu, Y., Comstock, B., Gonzalez, F., Mayock, D., Goodman, A., Maitre, N., Chang, T., Meurs, K., AL, L., Bendel-Stenzel, E., Mathur, A., Wu, T., Riley, D., Mietzsch, U., Baserga, M., Poindexter, B., Rogers, E., Lowe, J., Kuban, K., O'Shea, T., Wisnowski, J., Mckinstry, R., Bluml...
2022
-
[38]
(2021), Propensity score weighting for causal subgroup analysis, Stat Med\/ , 40, 4294--4309
Yang, S., Lorenzi, S., Papadogeorgou, G., Wojdyla, D., Li, F., and Thomas, L. (2021), Propensity score weighting for causal subgroup analysis, Stat Med\/ , 40, 4294--4309
2021
-
[39]
(2021), Propensity score weighting for covariate adjustment in randomized clinical trials, Stat Med\/ , 40, 842--858
Zeng, S., Li, F., Wang, R., and Li, F. (2021), Propensity score weighting for covariate adjustment in randomized clinical trials, Stat Med\/ , 40, 842--858
2021
-
[40]
(2019), Estimating Mann–Whitney-type Causal Effects, Int Stat Rev\/ , 87, 514--530
Zhang, Z., Ma, S., Shen, C., and Liu, C. (2019), Estimating Mann–Whitney-type Causal Effects, Int Stat Rev\/ , 87, 514--530
2019
-
[41]
Y., Mitra, N., Hemming, K., Harhay, M
Zhu, A. Y., Mitra, N., Hemming, K., Harhay, M. O., and Li, F. (2024), Leveraging baseline covariates to analyze small cluster-randomized trials with a rare binary outcome, Biometrical J\/ , 66, 2200135
2024
-
[42]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.