REVIEW 2 major objections 5 minor 37 references
The Win Ratio at the Design Stage of Clinical Trials
T0 review · 2 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read The win ratio can lift clinical-trial power by up to 50% over time-to-first-event analysis, and a new formula sizes precision-based trials from the desired confidence-interval width.
desk verdict Useful simulation study of the win ratio, but the 'novel' sample size formula is a direct inversion of Yu et al.'s variance approximation, which the paper's own case study shows to be unreliable; needs revision, not rejection. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying object is the unmatched win ratio: all $N_T \times N_C$ patient pairs are compared hierarchically, each pair is scored as a win, loss, or tie on the first outcome in the hierarchy that distinguishes the two patients, and the win ratio is $P(\text{win})/P(\text{loss})$. The design machinery is the approximate variance of the log-transformed win ratio from [19], $\operatorname{Var}(\log \widehat{\mathrm{WR}}) \approx \frac{4(1+p_{\mathrm{tie}})}{3 p_T(1-p_T)(1-p_{\mathrm{tie}}) N_{\mathrm{total}}}$, which depends only on allocation and tie probability. The paper's new sample-size formula is the algebraic inverse of this variance expression solved for $N_{\mathrm{total}}$ at a target Wald confidence-interval width; the same variance approximation underlies the power comparisons and the sensitivity of power to the tie proportion.
What would settle it
Run the paper's IPHAK scenario in a Monte Carlo study: with the expected win ratio of 1.32 and the observed tie proportion, the analytic variance formula gives power 0.764 while simulation gives 0.872; reproducing and explaining that gap would settle whether the variance approximation supporting both the power comparison and the new sample-size formula is accurate enough.
Extended reading notes
Core claim
On its own terms, the paper establishes two results. First, for composite endpoints with a clinically meaningful hierarchy, the win ratio can deliver materially higher statistical power than single-outcome analysis or time-to-first-event analysis: in the planned IPHAK trial setting the win ratio reached simulated power 0.872 versus 0.298 for the binary endpoint alone and 0.726 for the continuous endpoint alone, and in survival simulations the win ratio beat time-to-first-event analysis by more than 50% when the hazard ratio on the top-ranked outcome was strong. Second, by inverting the approximate variance formula $\operatorname{Var}(\log \widehat{\mathrm{WR}}) \approx \frac{4(1+p_{\mathrm{tie}})}{3 p_T (1-p_T)(1-p_{\mathrm{tie}}) N_{\mathrm{total}}}$, the paper derives a closed-form total sample size for precision-based trials that depends only on the desired confidence-interval width, the allocation proportion, and the expected tie proportion. The qualifying condition is that these gains appear when the highest-ranked outcome carries a non-negligible effect; when a continuous outcome sits at the top of the hierarchy, it tends to decide almost all pairwise comparisons and the lower-ranked outcomes add little.
Load-bearing premise
The design tool rests on the approximate variance formula for the log win ratio being accurate enough; the new sample-size formula is just that formula solved for N, so if the approximation mis-states the variance, the required sample size is wrong.
Editorial extensions
If this is right
- If the power gains hold, a trial that would be underpowered under time-to-first-event analysis can become adequately powered at the same sample size by switching to the win ratio, provided the hierarchy reflects clinical priorities and the top outcome is not nearly null.
- The sample-size formula gives a way to plan precision-based win-ratio trials without simulating a full treatment-effect scenario; the inputs are the desired confidence-interval width, the treatment allocation, and an anticipated tie proportion.
- The win ratio's advantage over single-endpoint analysis is selective: it appears when lower-ranked outcome effects are moderate relative to the higher-ranked outcome, and it largely disappears when a continuous outcome at the top of the hierarchy dominates the pairwise decisions.
- Adding a continuous outcome at the bottom of the hierarchy can break ties and raise power, even if it attenuates the estimated win ratio.
- Precision-based design is more forgiving than power-based design: a slight shortfall in precision is less consequential than labeling a trial unsuccessful, which supports using the approximation-based formula in exploratory settings.
Reading between the lines
- An implicit consequence of the variance formula is that the required sample size is highly sensitive to the anticipated tie proportion, so a sensitivity analysis over tie proportions should accompany any use of the new formula.
- The case-study gap between the analytic approximation (0.764) and simulation (0.872) suggests the approximate variance may understate power; if so, the sample-size formula would tend to under-size trials, and designs should verify target width by simulation.
- The same plug-in logic could be adapted to stratified or covariate-adjusted win ratio analyses, where tie probabilities would be stratum-specific; the paper's references already supply the needed variance formulas.
- For ordinal endpoints such as the modified Rankin Scale, the win ratio's nonparametric nature suggests the power gains may transfer, but the tie structure differs and would need its own simulation checks.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper addresses the use of the win ratio (WR) at the design stage of clinical trials. It contains (i) a case study of the planned IPHAK trial comparing unmatched WR analysis with single-endpoint tests for a composite of a binary and a continuous outcome; (ii) two simulation studies, one comparing WR with t-test and Fisher's exact test for a binary-plus-continuous composite and one comparing WR with time-to-first-event (TTFE) Cox/log-rank analysis for two time-to-event outcomes; and (iii) a proposed sample size formula for precision-based trials obtained by inverting the Yu et al. approximate variance of log(WR). The main claims are that WR analysis can provide materially higher power than single-endpoint or TTFE analysis, with increases up to 50% in some scenarios, and that the proposed formula allows trials to be sized for a target confidence-interval width.
Significance. If the power gains and the precision formula hold up, the paper would offer practical guidance for designing WR trials and a simple alternative to fully simulation-based planning. The paper has notable strengths: the simulations are structured according to the ADEMP framework, Monte Carlo standard errors are reported, code is posted on OSF, and the authors candidly report a discrepancy between analytic and simulated power in the case study. However, the new design formula inherits a variance approximation whose error is empirically visible in the paper's own case study, and the inferential procedure for the unmatched WR is not described in enough detail to audit the reported power estimates. These issues are load-bearing for the paper's central claims.
major comments (2)
- [Section 3; Section 5, Eq. (Width)] The precision-based sample size formula in Section 5 is obtained by algebraically solving Width = 2 Z_{1-α/2} se(log WR) for N_total using the variance approximation of Yu et al. In Section 3, the same approximation gives a power of 0.764 for the IPHAK scenario while the simulation gives 0.872 (MCSE 0.0106), and the authors state that the 'exact reason for this substantial difference is not fully clear'. Because a precision-based formula propagates any error in Var(log WR) directly to the width and hence to N_total, the paper needs a dedicated simulation study checking achieved confidence-interval width and coverage across a range of WR values, tie probabilities, and sample sizes before the formula can be recommended for design. The current text supplies no such validation.
- [Section 2.2; Section 4.1; Section 4.2] The p-value computation for the unmatched WR is never specified. Section 2.2 notes that variance estimation for the unmatched WR is complex and that bootstrap resampling was proposed, but the simulation studies in Sections 4.1 and 4.2 do not report the test statistic, the resampling procedure, the number of bootstrap samples, or how ties were handled. Without this information the reported power estimates are not reproducible from the text. In addition, no null-case simulations are reported, so the type I error rate of the procedure is unknown; the apparent power gains could in part reflect anti-conservatism of the test rather than a true efficiency advantage. The authors should specify the inference procedure and report type I error calibration.
minor comments (5)
- [Section 1] The phrase 'importance for patience' appears to be a typo and should read 'importance for patients'.
- [Figures 3 and 4 captions] The word 'Errobars' in the captions of Figures 3 and 4 should be corrected to 'Error bars'.
- [Appendix A.1] The notation 'NT otal' should be 'N_total'; the same typo appears in the variance formula display.
- [Section 5] The proposed formula is an algebraic rearrangement of the variance approximation in Yu et al. [19], not a new variance result; the wording 'novel formula' should be softened, and the text should state explicitly that the formula's accuracy depends entirely on the approximation in [19].
- [Section 2.3] The statement that the available implementation of the rank-based simulation approach [20] 'consistently returns power estimates of either zero or one' is a strong claim about existing software and should be documented with version information and a reproducible example.
Circularity Check
No significant circularity: the sample size formula is an explicit algebraic inversion of the externally cited Yu et al. variance approximation, and the power claims rest on Monte Carlo simulation rather than on fitting or self-citation.
full rationale
The derivation chain is self-contained against external benchmarks and contains no load-bearing self-citation. The Section 5 precision-based sample size formula is obtained by solving Width = 2 Z_{1-α/2} se(log WR) using the variance approximation Var(log WR) = 4(1+p_tie)/(3 p_T (1-p_T)(1-p_tie) N_total), which the paper explicitly attributes to Yu et al. [19]; this is an algebraic rearrangement of a cited external formula, not a target defined in terms of its own output. The power claims in Sections 3 and 4 are generated by Monte Carlo simulation under stated data-generating mechanisms (binomial/normal and Weibull), with code publicly available on OSF, and are compared with the t-test, Fisher's exact test, and Cox log-rank test, so they do not reduce to fitted inputs. The acknowledged discrepancy in Section 3 (analytic power 0.764 vs simulated power 0.872) is a limitation of the accuracy of the Yu et al. approximation and is explicitly flagged by the authors; it is an external-validity concern, not a circularity. No self-citation is used to justify the central premise, and no uniqueness theorem is imported from the authors' own prior work.
Assumptions & free parameters
free parameters (3)
- anticipated tie probability p_tie =
0.02 in the worked example; about 0.013 in the IPHAK case study (1.3% ties)
- Weibull shape parameters for the TTE simulation =
kappa = 4 for time to death, kappa = 2 for time to hospitalization
- IPHAK case study scenario inputs =
EBP rates 68% vs 85% among PA patients, DDD mean change -1.364, SD 1.8, correlation 0.5
assumptions (4)
- domain assumption The approximate variance formula Var(log WR) = 4(1 + p_tie) / (3 p_T (1 - p_T) (1 - p_tie) N_total) of Yu et al. [19] is accurate enough for trial design.
- standard math log(WR) is approximately normally distributed with the cited variance (CLT).
- domain assumption The unmatched WR test with resampling-based inference has correct type I error at alpha = 5% in the simulated settings.
- domain assumption Independence of simulated death and hospitalization times, and equal censoring across treatment groups.
Cite this review
Pith. "Pith review of The Win Ratio at the Design Stage of Clinical Trials." pith.science (2026). https://pith.science/paper/6ZQ2XFGN
@misc{pith2026250715685,
author = {Pith},
title = {Pith review of: The Win Ratio at the Design Stage of Clinical Trials},
year = {2026},
howpublished = {\url{https://pith.science/paper/6ZQ2XFGN}},
note = {Machine review of arXiv:2507.15685}
}
read the original abstract
The win ratio offers a flexible approach to incorporate the hierarchy of clinical outcomes into the analysis of a composite endpoint, enabling simultaneous consideration of multiple outcome types, unlike traditional time-to-first-event (TTFE) analysis or focus on a single outcome. We examined the statistical power of the win ratio compared to single-endpoint analyses and TTFE analysis through a case study and simulation studies. Furthermore, we provide a novel formula to estimate the required sample size for win ratio analysis based on the desired width of its confidence interval, facilitating precision-based trial design. Our results indicate that win ratio analysis generally outperforms single-endpoint analyses when treatment effects on lower-ranked outcomes are moderate compared to those on higher-ranked outcomes. The win ratio can provide greater power than TTFE analysis, especially when the effect on the highest-ranked outcome is substantial, reaching increases in power up to 50\%. Further, even for moderate treatment effects on the highest-ranked outcome, win ratio analysis achieved higher power. Future work should expand our simulations to additional data-generating mechanisms and outcome types, particularly ordinal outcomes, where the win ratio provides an alternative to existing non-parametric and parametric methods. Our findings highlight the potential of the win ratio to improve statistical efficiency in pharmaceutical and other clinical trial designs using composite endpoints, particularly when no single component dominates the treatment effect. However, when continuous outcomes occupy the top of the hierarchy, these tend to drive overall analysis, sidelining contributions of lower-ranked outcomes and limiting benefits of hierarchical win ratio analysis.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[21]
Statistical power considerations in the use of win ratio in cardiovascular outcome trials
Wang B, Zhou D, Zhang J, et al. Statistical power considerations in the use of win ratio in cardiovascular outcome trials. Contemporary Clinical Trials 2023; 124: 107040. doi: https://doi.org/10.1016/j.cct.2022.107040
arXiv 2023
-
[1]
The win ratio approach for composite endpoints: practi- cal guidance based on previous experience
Redfors B, Gregson J, Crowley A, et al. The win ratio approach for composite endpoints: practi- cal guidance based on previous experience. European Heart Journal 2020; 41(46): 4391–4399. doi: https://doi.org/10.1093/eurheartj/ehaa665
-
[2]
Use of the win ratio analysis in critical care trials
Monzo L, Levy B, Duarte K, et al. Use of the win ratio analysis in critical care trials. American Journal of Respiratory and Critical Care Medicine 2024; 209(7): 798–804. doi: https://doi.org/10.1164/rccm.202309-1644CP
-
[3]
Regression models and life-tables
Cox D. Regression models and life-tables. Journal of the Royal Statistical Society: Se- ries B (Methodological) 1972; 34(2): 187–202. doi: https://doi.org/10.1111/j.2517- 6161.1972.tb00899.x
arXiv 1972
-
[4]
Nonparametric estimation from incomplete observations
Kaplan E, Meier P. Nonparametric estimation from incomplete observations. Journal of the American statistical association 1958; 53(282): 457–481
work page 1958
-
[5]
Nonparametric estimation of a survivorship function with doubly censored data
Turnbull B. Nonparametric estimation of a survivorship function with doubly censored data. Journal of the American statistical association 1974; 69(345): 169–173
work page 1974
-
[6]
A note on compet- ing risks in survival data analysis
Satagopan J, Ben-Porat L, Berwick M, Robson M, Kutler D, Auerbach A. A note on compet- ing risks in survival data analysis. British journal of cancer 2004; 91(7): 1229–1235. doi: https://doi.org/10.1038/sj.bjc.6602102
-
[7]
Fine J, Gray R. A proportional hazards model for the subdistribution of a competing risk.Journal of the American statistical association 1999; 94(446): 496–509
work page 1999
Show all 37 references
-
[8]
A win ratio approach to comparing continuous non-normal outcomes in clinical trials
Wang D, Pocock S. A win ratio approach to comparing continuous non-normal outcomes in clinical trials. Pharmaceutical Statistics 2016; 15(3): 238–245. doi: 10.1002/pst.1743
2016 doi
-
[9]
Win ratio in biomedical science: a bibliometric analysis
Li Z, Izumi A, Vervoort D, Ranadive A, Verma S, Fremes S. Win ratio in biomedical science: a bibliometric analysis. CJC Open 2025. doi: https://doi.org/10.1016/j.cjco.2025.05.006
2025 doi
-
[10]
Combining mortality and longitudinal measures in clinical trials
Finkelstein D, Schoenfeld D. Combining mortality and longitudinal measures in clinical trials. Statistics in medicine 1999; 18(11): 1341–1354
1999
-
[11]
Generalized pairwise comparisons of prioritized outcomes in the two-sample problem
Buyse M. Generalized pairwise comparisons of prioritized outcomes in the two-sample problem. Statistics in medicine 2010; 29(30): 3245–3257. doi: https://doi.org/10.1002/sim.3923
2010 doi
-
[12]
The win ratio: a new approach to the analysis of composite endpoints in clinical trials based on clinical priorities
Pocock S, Ariti C, Collier T, Wang D. The win ratio: a new approach to the analysis of composite endpoints in clinical trials based on clinical priorities. European Heart Journal 2011; 33(2): 176–182. doi: 10.1093/eurheartj/ehr352 25
2011 doi
-
[13]
Use of the win ratio in cardiovascular trials
Ferreira J, Jhund P, Duarte K, et al. Use of the win ratio in cardiovascular trials. Heart failure 2020; 8(6): 441–450. doi: https://doi.org/10.1016/j.jchf.2020.02.010
2020 doi
-
[14]
Comparing minimally invasive and open pancreaticoduo- denectomy for the treatment of pancreatic cancer: a win ratio analysis
Beal E, Dalmacy D, Paro A, et al. Comparing minimally invasive and open pancreaticoduo- denectomy for the treatment of pancreatic cancer: a win ratio analysis. Journal of Gastrointesti- nal Surgery 2022; 26(8): 1697–1704. doi: https://doi.org/10.1007/s11605-022-05380-3
2022 doi
-
[15]
Berwanger O, Pfeffer M, Claggett B, et al. Sacubitril/valsartan versus ramipril for patients with acute myocardial infarction: win-ratio analysis of the PARADISE-MI trial.European journal of heart failure 2022; 24(10): 1918–1927. doi: https://doi.org/10.1002/ejhf.2663
2022 doi
-
[16]
Mobile health-technology-integrated care for atrial fibrillation: a win ratio analysis from the mAFA-II randomized clinical trial
Romiti G, Guo Y , Corica B, et al. Mobile health-technology-integrated care for atrial fibrillation: a win ratio analysis from the mAFA-II randomized clinical trial. Thrombosis and haemostasis 2023; 123(11): 1042–1048. doi: 10.1055/s-0043-1769612
2023 doi
-
[17]
Win ratio analysis of short-term clinical outcomes of focal therapy and robot-assisted radical prostatectomy for the patients with localized prostate cancer
Teramoto A, Sakamaki K, Shoji S, Uemura K. Win ratio analysis of short-term clinical outcomes of focal therapy and robot-assisted radical prostatectomy for the patients with localized prostate cancer. Scientific Reports 2024; 14(1): 17019. doi: https://doi.org/10.1038/s41598-0...
2024 doi
-
[18]
Sample size formula for general win ratio analysis
Mao L, Kim K, Miao X. Sample size formula for general win ratio analysis. Biometrics 2021; 78(3): 1257–1268. doi: 10.1111/biom.13501
2021 doi
-
[19]
Sample size formula for a win ratio endpoint
Yu R, Ganju J. Sample size formula for a win ratio endpoint. Statistics in Medicine 2022; 41(6): 950–963. doi: 10.1002/sim.9297
2022 doi
-
[20]
Power considerations for the win ratio: A rank-based simulation approach
Bonner L, Ciolino J, Kaye K, Wunderink R, Scholtens D. Power considerations for the win ratio: A rank-based simulation approach. Contemporary Clinical Trials 2025: 107937. doi: https://doi.org/10.1016/j.cct.2025.107937
2025
-
[22]
Win odds: an adaptation of the win ratio to include ties
Brunner E, Vandemeulebroecke M, Mütze T. Win odds: an adaptation of the win ratio to include ties. Statistics in Medicine 2021; 40(14): 3367–3384. doi: https://doi.org/10.1002/sim.8967
2021 doi
-
[23]
The stratified win ratio.Journal of biopharma- ceutical statistics 2018; 28(4): 778–796
Dong G, Qiu J, Wang D, Vandemeulebroecke M. The stratified win ratio.Journal of biopharma- ceutical statistics 2018; 28(4): 778–796. doi: https://doi.org/10.1080/10543406.2017.1397007
2018
-
[24]
The stratified win statistics (win ratio, win odds, and net benefit)
Dong G, Hoaglin D, Huang B, et al. The stratified win statistics (win ratio, win odds, and net benefit). Pharmaceutical Statistics 2023; 22(4): 748–756. doi: https://doi.org/10.1002/pst.2293 26
2023 doi
-
[25]
Adjusted win ratio with strat- ification: calculation methods and interpretation
Gasparyan S, Folkvaljon F, Bengtsson O, Buenconsejo J, Koch G. Adjusted win ratio with strat- ification: calculation methods and interpretation. Statistical Methods in Medical Research2021; 30(2): 580–611. doi: https://doi.org/10.1177/0962280220942558
-
[26]
Adjusted win ratio using the inverse proba- bility of treatment weighting
Wang D, Zheng S, Cui Y , He N, Chen T, Huang B. Adjusted win ratio using the inverse proba- bility of treatment weighting. Journal of biopharmaceutical statistics 2025; 35(1): 21–36. doi: https://doi.org/10.1080/10543406.2023.2275759
2025
-
[27]
Using simulation studies to evaluate statistical methods
Morris T, White I, Crowther M. Using simulation studies to evaluate statistical methods. Statis- tics in medicine 2019; 38(11): 2074–2102. doi: https://doi.org/10.1002/sim.8086
2019 doi
-
[28]
Generating survival times to simulate Cox pro- portional hazards models
Bender R, Augustin T, Blettner M. Generating survival times to simulate Cox pro- portional hazards models. Statistics in medicine 2005; 24(11): 1713–1723. doi: https://doi.org/10.1002/sim.2059
2005 doi
-
[29]
The tyranny of power: is there a better way to calculate sample size?
Bland J. The tyranny of power: is there a better way to calculate sample size?. Bmj 2009; 339. doi: https://doi.org/10.1136/bmj.b3985
2009 doi
-
[30]
The ongoing tyranny of statistical significance testing in biomedical research
Stang A, Poole C, Kuss O. The ongoing tyranny of statistical significance testing in biomedical research. European journal of epidemiology 2010; 25(4): 225–230. doi: https://doi.org/10.1007/s10654-010-9440-x
2010 doi
-
[31]
The p-value Function and Statistical Inference
Fraser D. The p-value Function and Statistical Inference. The American Statistician 2019; 73(sup1): 135–147. doi: 10.1080/00031305.2018.1556735
2019 arXiv
-
[32]
Van Rijn M, Bech A, Bouyer J, Brand v. dJ. Statistical significance versus clini- cal relevance. Nephrology Dialysis Transplantation 2017; 32(suppl_2): ii6–ii12. doi: https://doi.org/10.1093/ndt/gfw385
2017 doi
-
[33]
Sample Sizes for Clinical Trials
Julious S. Sample Sizes for Clinical Trials. Chapman and Hall/CRC . 2023
2023
-
[34]
Simulating biologically plausible complex survival data
Crowther M, Lambert P. Simulating biologically plausible complex survival data. Statistics in medicine 2013; 32(23): 4118–4134. doi: https://doi.org/10.1002/sim.5823
2013 doi
-
[35]
A comparative study to alterna- tives to the log-rank test
Dormuth I, Liu T, Xu J, Pauly M, Ditzhaus M. A comparative study to alterna- tives to the log-rank test. Contemporary clinical trials 2023; 128: 107165. doi: https://doi.org/10.1016/j.cct.2023.107165
2023
-
[36]
Pragmatic outcomes for stroke research
Ospel J, Ganesh A, Goyal M. Pragmatic outcomes for stroke research. The Lancet Neurology 2024; 23(9): 860–862
2024
-
[37]
ICH Harmonised 27 Tripartite Guideline: Statistical Principles for Clinical Trials (E9)
International Council for Harmonisation of Technical Requirements for Pharmaceu- ticals for Human Use (ICH) Efficacy Expert Working Group . ICH Harmonised 27 Tripartite Guideline: Statistical Principles for Clinical Trials (E9). Tech. Rep. E9, International Council for Harmoni...
1998
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.