{"id":"47417468-3836-4ea1-a836-b335e5fd5e0e","arxiv_id":"2508.20349","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Adjusting win statistics for baseline covariates with propensity and outcome weighting is consistent, can reduce asymptotic variance, and is implemented in the R package winPSW.","lead":"This paper develops covariate-adjusted estimators for win statistics (win ratio and win difference) with ordinal outcomes in randomized trials, using inverse probability, overlap, and augmented weighting. The methods improve precision over unadjusted analyses, come with closed-form variance formulas and R software, and are illustrated on the ORCHID COVID-19 trial.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Model-robustness claim in the abstract is overstated: IPW/OW consistency requires the propensity model to be able to represent the constant propensity.","rationale":"The reader's weakest_assumption identifies exactly this issue: consistency of IPW and OW relies on the estimated logistic propensity score converging to the constant pi, which requires the propensity model family to contain the constant. This is a real limitation of the abstract's model-robustness wording, but it is a qualification rather than a fundamental flaw. Standard logistic regression includes an intercept, so the proposed methods work in their intended implementations; the concern is that the paper overclaims robustness across arbitrary misspecification. The reader's CONDITIONAL verdict remains appropriate: the abstract and theorem statements should be revised to state the intercept/constant-propensity condition explicitly. No change to the verdict is needed.","tokens_in":30800,"tokens_out":59646,"duration_ms":501452,"concrete_test":"Simulate a two-arm RCT with pi=0.7 and an ordinal outcome independent of covariates. Fit the IPW and OW estimators using a logistic propensity model with no intercept (e.g., in R, glm(Z ~ -1 + X1 + X2, family=binomial)). Compute Monte Carlo bias of the resulting win ratio and win difference estimators against the true tau_1 as n grows. If the bias does not vanish, the model-robustness claim fails without the intercept qualification; rerunning with an intercept should eliminate the bias.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract claims 'all of the covariate-adjusted estimators do not compromise consistency for the target estimand even when the associated working models are incorrectly specified; hence these covariate-adjusted estimators are model-robust.' For IPW (3.4) and OW (3.5), consistency to tau_1 requires the estimated propensity score to converge to the true constant pi, because only then do the weights reduce to constants and the weighted estimand (3.1) collapse to tau_1. This convergence is guaranteed only if the working propensity model family contains the constant propensity, as a logistic model with an intercept does. If the model is specified without an intercept, or in any other way that cannot represent e(X)=pi, the estimated propensity converges to the best-fitting nonconstant approximation, the weights do not become constant, and the estimator targets a different weighted estimand instead of tau_1. Section 3.1 notes that tau^h_1=tau_1 'as long as h is at most a function of covariates only through the propensity score,' but neither the abstract nor the theorems state the required condition that the propensity model family can represent the constant propensity. The blanket model-robustness claim is therefore too strong without this unstated qualification.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper develops covariate-adjusted estimators for win statistics (win ratio and win difference) with ordinal outcomes in randomized clinical trials. The authors propose inverse probability weighting (IPW), overlap weighting (OW), and their augmented versions (AIPW, AOW), building on U-statistic theory. They state two central theoretical results: (i) the influence function of any balancing-weighting estimator equals the unadjusted estimator's influence function minus its projection onto the propensity-score tangent space, implying no larger asymptotic variance; and (ii) all estimators are consistent for the target estimand even when working models are misspecified, i.e., model-robust. The paper provides closed-form variance estimators, simulation studies, an application to the ORCHID trial, and an R package.","tokens_in":31017,"tokens_out":9249,"duration_ms":85091,"significance":"If the results are correct, the paper fills a gap by formally extending covariate adjustment for average treatment effects to pairwise win estimands, with a clean influence-function projection argument and practical software. The simulation study is broad and the GitHub code enables reproducibility. However, the central model-robustness claim is overstated: for IPW and OW, consistency to the unadjusted win estimand requires the propensity model to be able to represent the constant propensity, a condition not stated in the abstract or theorems. The variance-reduction claim likewise relies on correct specification of the propensity model. These issues affect the paper's central message and need to be addressed with explicit conditions or revised claims.","major_comments":[{"comment":"The abstract claims that 'all of the covariate-adjusted estimators do not compromise consistency for the target estimand even when the associated working models are incorrectly specified.' This is not correct for the IPW estimator (3.4) and the OW estimator (3.5) unless the propensity-score model family contains the constant propensity. Under randomization the true propensity is a constant π; consistency to τ1 requires the estimated propensity score to converge to π so that the weights become constants. If the working model cannot represent a constant (e.g., a logistic model without an intercept when π≠0.5), the estimated weights converge to a nonconstant function and the estimator targets a different weighted estimand. Theorem 1 does not state this condition, and the proof in Web Appendix B is not available to verify. Please add the required assumption or qualify the model-robustness claim to apply only to outcome-model misspecification and to propensity models that contain the constant.","section":"Abstract; Section 3.2, Eqs. (3.4)-(3.5); Section 7"},{"comment":"The statement that the influence function φ(O) equals χh(O) minus its projection on the tangent space of the propensity model, and the consequent conclusion that 'the asymptotic variance of the propensity score weighting estimators is no larger than that of the unadjusted estimator,' presuppose that the propensity model is correctly specified (or at least that the probability limit of the estimated propensity is the true constant). If the model is misspecified, the projection is not an orthogonal projection in the sense needed for the variance decomposition, and the variance-reduction claim is not established. The theorem and the discussion should explicitly state this condition.","section":"Section 3.3, Theorem 1(b); Section 7"},{"comment":"The proofs of Theorems 1 and 2 are stated to be in Web Appendices B and D, but these appendices are not included in the submitted manuscript. Because these theorems are load-bearing for the paper's central claims of consistency, variance reduction, and local efficiency, the full proofs must be provided to the reviewers.","section":"Section 3.3 (Theorem 1) and Section 4.3 (Theorem 2)"}],"minor_comments":[{"comment":"The word 'recommendped' in the paragraph about regulatory guidance should be 'recommended'.","section":"Section 1"},{"comment":"The phrase 'to obtain consistant variance estimators' contains a typo: 'consistant' should be 'consistent'. Similarly, Section 3.4 has 'asymtoptic' which should be 'asymptotic'.","section":"Section 3.3"},{"comment":"The final paragraph contains a sentence fragment: 'Our work thus illustrates the operational steps involved in estimating win estimands with covariate adjustment. tics as well as the una djusted win statistics in the winPSW R package...' This appears to be a corruption and should be rewritten.","section":"Section 7"},{"comment":"The statement 'under randomization, τ h 1 = τ1 as long as h(Xi,X j) is at most a function of covariates only through the propensity score' is potentially confusing because the true propensity score is constant under randomization; a nonconstant function of covariates is not a function of the propensity score. Please clarify what is intended.","section":"Section 3.1"},{"comment":"The references to supporting material are inconsistent: the text mentions 'Web Appendix A-D' and 'Web Appendices A–E'. Please standardize.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The paper is a potentially useful contribution to the clinical trials literature, but the abstract's model-robustness claim is broader than what the theory supports. The authors should either add the explicit condition that the propensity model contains the constant propensity for the IPW/OW consistency and variance-reduction claims, or moderate the abstract. The omission of the Web Appendices from the submission is a practical obstacle for review; they should be included. The self-citations are not excessive, but the novelty relative to Wang et al. (2023b) and Vermeulen et al. (2015) should be clearly delineated in the revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear [name],\n\nThis paper does something genuinely useful: it extends covariate adjustment with IPW/OW and augmented versions to win ratio and win difference for ordinal outcomes in randomized trials, and it ships an R package. The key theoretical claim — that a balancing-weight estimator's influence function equals the unadjusted influence function minus its projection on the propensity-score tangent space — is a natural extension of the logic in Zeng et al., and it gives trialists a principled reason to expect variance reduction without changing the estimand. If the web-appendix proofs are as advertised, that claim is the main reason to take the paper seriously.\n\nWhat is new: OW and AOW estimators for pairwise win estimands, closed-form variance estimators, local efficiency/robustness results for the augmented versions, simulation evidence across balanced/unbalanced designs and misspecified outcome models, and a clean ORCHID reanalysis. The simulation design is reasonable and the relative efficiency gains support the theoretical case. This is an extension of prior work, but a substantive one.\n\nNow the soft spots, in proportion. First, the abstract's blanket 'model-robust' wording is slightly overbroad. For IPW and OW, consistency for the unadjusted target requires the fitted propensity model to contain the constant propensity: if you fit a logistic model with an intercept, the MLE converges to pi under randomization, so the standard implementation is fine. The paper should just say that. The stress-test concern about no-intercept specifications is real but not damaging for normal practice. Second, the proofs live in web appendices not in the arXiv text, so the asymptotic claims cannot be independently checked from what's here. Third, Table 2 shows the augmented estimators' variance estimators are markedly anti-conservative under the interaction DGP and unbalanced design at n=200-400 (variance ratios around 0.5-0.8, coverage sometimes below 90%). The authors acknowledge this, but it matters because that is exactly the trial-size range. Fourth, the efficiency comparisons do not include existing covariate-adjusted win or Mann-Whitney estimators, so the relative performance claims are narrower than the framing suggests.\n\nCitation pattern looks fine; the reliance on Zeng et al. is legitimate since they are extending that framework.\n\nMy recommendation: send it to a serious referee. The central argument is coherent, the contribution is real, and the main concerns are addressable: make the proofs available, add the intercept condition, improve the finite-sample variance discussion, and add at least one existing adjusted estimator as a baseline. I would cite this if I worked on trial-based win statistics.","headline":"A substantive extension of covariate-adjusted win statistics to ordinal RCT outcomes, with a plausible variance-reduction theorem; needs the web-appendix proofs made visible and the finite-sample variance caveats taken seriously.","tokens_in":31533,"tokens_out":4265,"would_cite":true,"duration_ms":43898,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62G20","62P10"],"pacs":[],"model":"deepseek-v4-flash","headline":"Covariate-adjusted win statistics keep consistency and never increase asymptotic variance.","keywords":["ordinal outcomes","win ratio","win difference","covariate adjustment","propensity score weighting","overlap weighting","augmented weighting","U-statistics"],"falsifier":"Simulate a randomized trial with treatment probability $\\pi=0.5$, a strong prognostic covariate, and a logistic propensity model that omits the intercept so it cannot represent $\\pi$; if the IPW or OW win-ratio estimator shows bias that persists as $n$ grows to 10,000, the claimed model-robust consistency is false.","tokens_in":30630,"feed_emoji":"📊","tokens_out":7554,"duration_ms":71983,"temperature":0.7,"pith_summary":"Randomized trials with ordinal outcomes often report win ratios or win differences: pairwise comparisons of whether a treated patient beats a control patient. This paper claims that these estimands can be covariate-adjusted with inverse probability weighting, overlap weighting, or their augmented versions without sacrificing consistency, and that the adjusted estimators have asymptotic variance no larger than the unadjusted one. The carrying argument is an influence-function decomposition: each adjusted estimator's influence function is the unadjusted influence function minus its projection onto the propensity-score tangent space. If correct, the result gives trialists a formal reason to use prespecified baseline covariates for win statistics, together with closed-form variance formulas that avoid resampling. The paper also shows in simulations and in a completed COVID-19 hydroxychloroquine trial that adjustment can cut standard errors by roughly a third.","feed_headline":"Covariate-adjusted win ratios match or beat unadjusted precision","feed_subtitle":"New weighting estimators for ordinal trial outcomes have closed-form variances and stay consistent under model misspecification.","key_machinery":"The central object is the win estimand $\\tau_1=P\\{Y_i(1)>Y_j(0)\\}$, the probability that a randomly chosen treated outcome beats a randomly chosen control outcome, alongside its loss counterpart and their ratio and difference. The machinery is the U-statistic representation of weighted pairwise comparisons: an estimator is a normalized sum over treatment–control pairs of a symmetric kernel that records which patient wins, with weights determined by the fitted propensity score. The load-bearing identity is Theorem 1(b): for any balancing-weight estimator, the influence function is the unadjusted influence function minus its projection on the tangent space of the propensity model, which is exactly what makes the adjusted asymptotic variance no larger. Closed-form variance estimates follow from taking the sample second moment of the estimated influence function, so inference does not require resampling.","core_discovery":"On the paper's own terms, the central discovery is that the class of balancing-weight win estimators shares one influence-function structure: a normalized weighted comparison of every treatment–control pair, with weights built from a fitted propensity score. Under randomization, the weighted estimand reduces to the unadjusted win estimand whenever the propensity model can represent the constant treatment probability. The paper's Theorem 1 then shows the influence function of any such estimator equals the unadjusted influence function minus its projection on the propensity-score tangent space, so the asymptotic variance is no larger. Theorem 2 shows that adding an ordinal outcome regression through augmented weighting keeps consistency even when that regression is misspecified, and achieves local efficiency when it is correct. The win ratio and win difference inherit these properties by the delta method.","pith_inferences":["Editorial inference: because the variance-reduction argument only uses the structure of pairwise comparisons and a propensity tangent space, the same projection identity should extend to time-to-event win ratios and to other pairwise estimands such as net benefit with ties handled explicitly.","Editorial inference: the model robustness result is asymmetric: it protects against misspecifying the outcome model, not the propensity model, so in practice the propensity model must include an intercept or otherwise be able to represent the constant randomization probability.","Editorial inference: the simulations show variance estimates falling below Monte Carlo variance when the augmented outcome model has many parameters at n around 200, so small-sample corrections or sample-splitting are a natural next test."],"forward_implications":["Covariate adjustment for win ratio and win difference can be prespecified in a trial protocol with a guarantee of no asymptotic precision loss relative to the unadjusted analysis.","Analysts can report confidence intervals from closed-form variance estimators rather than bootstrap resampling.","Augmented weighting estimators provide an additional efficiency gain in most simulation settings, and remain consistent when the outcome regression is misspecified.","Overlap weighting is the recommended default when sample sizes are small or randomization is unbalanced, because it tends to be at least as efficient as IPW and avoids extreme weights.","The ORCHID reanalysis illustrates that adjustment can reduce standard errors by about 30–40% while leaving the clinical conclusion unchanged."],"supporting_citations":[{"why":"Supplies the balancing-weight framework and the IPW-versus-OW comparison for average treatment effects that this paper extends to win estimands.","marker":"Zeng et al. (2021)"},{"why":"Provides the U-statistic formulation of causal pairwise estimands and the augmented IPW estimator that the paper generalizes.","marker":"Mao (2018)"},{"why":"Gives the unadjusted win-ratio estimator and its large-sample variance, the baseline all adjusted estimators are compared against.","marker":"Bebu and Lachin (2016)"},{"why":"Establishes that inverse probability weighting with an estimated propensity score does not inflate variance for average treatment effects, a result the paper transfers to pairwise comparisons.","marker":"Shen et al. (2014)"},{"why":"Shows the variance-reduction property of IPW in randomized trials, another foundation for the claimed precision gain.","marker":"Williams et al. (2014)"},{"why":"Defines balancing weights and the exact-balance and minimum-variance properties of overlap weighting used to motivate OW.","marker":"Li et al. (2018)"},{"why":"Supplies the semiparametric tangent-space projection theory used in Theorem 1(b).","marker":"Tsiatis (2006)"},{"why":"Provides the uniform laws for U-processes used to prove consistency of the weighted estimators.","marker":"Arcones and Giné (1993)"}],"fun_headline_variants":["Covariate-adjusted win stats boost precision","Win statistics sharpen with covariate adjustment","Adjusted win estimators stay robust to model errors","Closed-form variance for adjusted win statistics","Propensity weighting improves win ratio precision"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the fitted treatment-assignment model can represent the true constant randomization probability; if it cannot, the weighted estimator targets a different quantity and the consistency claim does not hold.","fun_headline_variants_meta":{"raw":{"variants":["Covariate-adjusted win stats boost precision","Win statistics sharpen with covariate adjustment","Adjusted win estimators stay robust to model errors","Closed-form variance for adjusted win statistics","Propensity weighting improves win ratio precision"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000349,"raw_usage":{"total_tokens":1895,"prompt_tokens":921,"completion_tokens":974,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":537,"completion_tokens_details":{"reasoning_tokens":910}},"tokens_in":537,"tokens_out":974,"duration_ms":9808,"temperature":1.0,"reasoning_tokens":910,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T16:47:31.966412+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate a randomized trial with treatment probability $\\pi=0.5$, a strong prognostic covariate, and a logistic propensity model that omits the intercept so it cannot represent $\\pi$; if the IPW or OW win-ratio estimator shows bias that persists as $n$ grows to 10,000, the claimed model-robust consistency is false.","supporting_citations":[{"cited_title":"(2006), Semiparametric Theory and Missing Data\\/ , New York: Springer","cited_arxiv_id":null,"evidence_quote":"Supplies the semiparametric tangent-space projection theory used in Theorem 1(b)."},{"cited_title":"(2014), Variance reduction in randomised trials by inverse probability weighting using the propensity score, Stat Med\\/ , 33, 721--737","cited_arxiv_id":null,"evidence_quote":"Shows the variance-reduction property of IPW in randomized trials, another foundation for the claimed precision gain."},{"cited_title":"(2021), Propensity score weighting for covariate adjustment in randomized clinical trials, Stat Med\\/ , 40, 842--858","cited_arxiv_id":null,"evidence_quote":"Supplies the balancing-weight framework and the IPW-versus-OW comparison for average treatment effects that this paper extends to win estimands."},{"cited_title":"and Lachin, J","cited_arxiv_id":null,"evidence_quote":"Gives the unadjusted win-ratio estimator and its large-sample variance, the baseline all adjusted estimators are compared against."},{"cited_title":"(2014), Inverse probability weighting for covariate adjustment in randomized studies, Stat Med\\/ , 33, 555--568","cited_arxiv_id":null,"evidence_quote":"Establishes that inverse probability weighting with an estimated propensity score does not inflate variance for average treatment effects, a result the paper transfers to pairwise comparisons."},{"cited_title":"and Gin \\'e , E","cited_arxiv_id":null,"evidence_quote":"Provides the uniform laws for U-processes used to prove consistency of the weighted estimators."}],"review_version":2}