{"id":"84054f12-d4a7-4d5d-b0d2-ce6033226285","arxiv_id":"2507.15685","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A case study and simulation studies find the win ratio can beat single-outcome and time-to-first-event analysis in power, and a new formula sizes trials from a target confidence interval width.","lead":"This paper uses a case study and Monte Carlo simulations to show that the win ratio, a hierarchical pairwise comparison method for composite clinical endpoints, can deliver substantially more statistical power than single-outcome or time-to-first-event analysis. It also offers a formula for planning precision-based trials by sizing a study from a target confidence interval width.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The Section 5 precision-based sample size formula inherits the Yu et al. variance approximation that the paper's own case study (Section 3) shows to be unreliable (analytic power 0.764 vs simulated 0.872); a direct width/coverage check is needed before the formula is used for design.","rationale":"I agree with the reader's weakest-assumption identification. The central claim has two parts; the power-gain part is supported by internally consistent simulations and is consistent with earlier work [21], but the novel sample size formula is the part that most needs to be true for the paper to make a new design contribution, and it rests on a variance approximation that the paper's own Section 3 case study shows can be off by roughly 11 percentage points in power. The Section 5 formula is a direct algebraic inversion of that approximation, so any inaccuracy in the variance has a first-order effect on the computed N_total. The abstract and Section 5 frame the formula as 'novel' and useful for precision-based trials; the verification in the paper is only an arithmetic consistency check (N=134 from the given inputs), not a check that a Wald interval for log(WR) actually achieves the target width or coverage under a known data-generating process. The missing null-case simulation and the unspecified p-value procedure are additional audit gaps, but they are secondary to the variance-formula issue because the precision-based design tool is the paper's most distinctive contribution. I do not see this as grounds for rejection: the simulation code is shared, the ADEMP structure is clear, the MCSEs are reported, and the authors explicitly acknowledge the limitation in Section 6 and Appendix A.1. The appropriate disposition is to require the authors to validate the Section 5 formula by simulation (or to add a caveat that it is only as reliable as the Yu et al. approximation). That matches the reader's CONDITIONAL verdict, so I leave the verdict unchanged.","tokens_in":14986,"tokens_out":5614,"duration_ms":57501,"concrete_test":"Use the publicly available IPHAK simulation code (OSF https://osf.io/6qjup/) and its parameter specifications to generate the planned analysis set (255 per group). Across at least 2500 replications, compute the empirical mean and distribution of the width of the 95% Wald interval for log(WR), and its coverage of the true WR. If the empirical mean width deviates by more than 10% from the width implied by the Section 5 formula (with the simulated p_tie), or if coverage is not close to 95%, the precision-based formula is miscalibrated at the paper's own design point. As a complementary audit, run the Section 4.1 simulation grid under the global null (no treatment effect) using the same WR p-value procedure and verify that the empirical rejection rate at α=0.05 is 5% within MCSE; this settles whether the reported power gains are interpretable.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's new contribution in Section 5 is derived by solving Width = 2 Z_{1-α/2} se(log WR) for N_total using the Yu et al. [19] approximation Var(log WR) = 4(1+p_tie)/(3 p_T (1-p_T)(1-p_tie) N_total). The load-bearing premise is that this variance approximation is accurate enough for design. The paper itself supplies evidence against that premise in Section 3: with WR=1.32 and the simulated tie fraction, the [19] power calculation gives 0.764 while the simulation gives 0.872 (MCSE 0.0106); the authors write that the 'exact reason for this substantial difference is not fully clear' and attribute it to the variance approximation. For a precision-based formula, an error in Var(log WR) propagates directly to the width and therefore to N_total; the formula also omits any dependence on the true WR, and Appendix A.1 itself cautions that the approximation 'may be inaccurate when the true WR is either very large or small and tends to underestimate in small samples'. Because no null-case simulation is reported and the p-value procedure for the unmatched WR is never specified, the power claims are also not independently auditable from the text. The concern is thus not that the formula is approximate, but that its error is empirically visible and unexplained in the paper's own worked setting and is used without a dedicated validation for the width application.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper addresses the use of the win ratio (WR) at the design stage of clinical trials. It contains (i) a case study of the planned IPHAK trial comparing unmatched WR analysis with single-endpoint tests for a composite of a binary and a continuous outcome; (ii) two simulation studies, one comparing WR with t-test and Fisher's exact test for a binary-plus-continuous composite and one comparing WR with time-to-first-event (TTFE) Cox/log-rank analysis for two time-to-event outcomes; and (iii) a proposed sample size formula for precision-based trials obtained by inverting the Yu et al. approximate variance of log(WR). The main claims are that WR analysis can provide materially higher power than single-endpoint or TTFE analysis, with increases up to 50% in some scenarios, and that the proposed formula allows trials to be sized for a target confidence-interval width.","tokens_in":15085,"tokens_out":6739,"duration_ms":76998,"significance":"If the power gains and the precision formula hold up, the paper would offer practical guidance for designing WR trials and a simple alternative to fully simulation-based planning. The paper has notable strengths: the simulations are structured according to the ADEMP framework, Monte Carlo standard errors are reported, code is posted on OSF, and the authors candidly report a discrepancy between analytic and simulated power in the case study. However, the new design formula inherits a variance approximation whose error is empirically visible in the paper's own case study, and the inferential procedure for the unmatched WR is not described in enough detail to audit the reported power estimates. These issues are load-bearing for the paper's central claims.","major_comments":[{"comment":"The precision-based sample size formula in Section 5 is obtained by algebraically solving Width = 2 Z_{1-α/2} se(log WR) for N_total using the variance approximation of Yu et al. In Section 3, the same approximation gives a power of 0.764 for the IPHAK scenario while the simulation gives 0.872 (MCSE 0.0106), and the authors state that the 'exact reason for this substantial difference is not fully clear'. Because a precision-based formula propagates any error in Var(log WR) directly to the width and hence to N_total, the paper needs a dedicated simulation study checking achieved confidence-interval width and coverage across a range of WR values, tie probabilities, and sample sizes before the formula can be recommended for design. The current text supplies no such validation.","section":"Section 3; Section 5, Eq. (Width)"},{"comment":"The p-value computation for the unmatched WR is never specified. Section 2.2 notes that variance estimation for the unmatched WR is complex and that bootstrap resampling was proposed, but the simulation studies in Sections 4.1 and 4.2 do not report the test statistic, the resampling procedure, the number of bootstrap samples, or how ties were handled. Without this information the reported power estimates are not reproducible from the text. In addition, no null-case simulations are reported, so the type I error rate of the procedure is unknown; the apparent power gains could in part reflect anti-conservatism of the test rather than a true efficiency advantage. The authors should specify the inference procedure and report type I error calibration.","section":"Section 2.2; Section 4.1; Section 4.2"}],"minor_comments":[{"comment":"The phrase 'importance for patience' appears to be a typo and should read 'importance for patients'.","section":"Section 1"},{"comment":"The word 'Errobars' in the captions of Figures 3 and 4 should be corrected to 'Error bars'.","section":"Figures 3 and 4 captions"},{"comment":"The notation 'NT otal' should be 'N_total'; the same typo appears in the variance formula display.","section":"Appendix A.1"},{"comment":"The proposed formula is an algebraic rearrangement of the variance approximation in Yu et al. [19], not a new variance result; the wording 'novel formula' should be softened, and the text should state explicitly that the formula's accuracy depends entirely on the approximation in [19].","section":"Section 5"},{"comment":"The statement that the available implementation of the rank-based simulation approach [20] 'consistently returns power estimates of either zero or one' is a strong claim about existing software and should be documented with version information and a reproducible example.","section":"Section 2.3"}],"recommendation":"major_revision","confidential_remarks":"The manuscript would benefit from a statistical reviewer with access to the OSF code to verify the p-value implementation, since the text alone does not permit replication of the power estimates. The 'novel formula' is an algebraic inversion of an existing approximation, so the contribution should be reframed as an applied design tool. If the authors cannot add a validation study of the width formula, the sample-size contribution should be substantially downgraded or removed. The paper's applied simulation results are useful, but the missing inferential details and lack of type I error calibration are important."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe paper gives trial designers something useful, and the simulation work is mostly solid. What's genuinely new: the power comparisons for a binary-plus-continuous composite and for TTE endpoints extend earlier TTE-focused work, and the worked example of the precision-based sample size formula is internally consistent. The code is on OSF, the simulations follow the ADEMP structure, and MCSEs are reported. That's real evidence and should be credited.\n\nThe soft spots are in the new contribution. Section 5's 'novel formula' is an algebraic inversion of the Yu et al. variance approximation for log(WR). The paper itself shows in Section 3 that this approximation gives a power estimate of 0.764 versus a simulated 0.872 (MCSE 0.0106), and the authors say the reason is 'not fully clear.' The formula inherits that approximation. It's true that in this case the approximation seems to overestimate the variance (the simulated power is higher), so the sample size formula would tend to over-size rather than under-size — but that is not a reason to trust it across the settings the paper recommends. Appendix A.1 also warns the approximation 'may be inaccurate when the true WR is either very large or small and tends to underestimate in small samples.' So a direct check of the width and coverage performance of the proposed formula is needed before anyone uses it to size a real trial.\n\nTwo other gaps: no null-case simulations are reported, so the actual type I error rate of the unmatched WR test in these settings is unknown; power defined as P(p <= 0.05) is only meaningful if the test is calibrated. And the paper never states which test statistic and resampling procedure were used for the WR p-values, which blocks independent audit from the text. The OSF code helps, but the methods section should say.\n\nThe central power finding — WR can beat single endpoints and TTFE — is plausible and consistent with Wang et al. [21]. I don't think the paper's core claim collapses, but the design-tool claim is oversold relative to the evidence.\n\nWho is this for? Trial statisticians considering the win ratio at the planning stage. It deserves a serious peer review, but it needs revision before publication: add null simulations, specify the test procedure, and validate the sample size formula against simulated confidence interval widths and coverage.\n\nRecommendation: send to review, require major revision.","headline":"Useful simulation study of the win ratio, but the 'novel' sample size formula is a direct inversion of Yu et al.'s variance approximation, which the paper's own case study shows to be unreliable; needs revision, not rejection.","tokens_in":15851,"tokens_out":2745,"would_cite":false,"duration_ms":29651,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62P10","62G10","62N03"],"pacs":[],"model":"deepseek-v4-flash","headline":"The win ratio can lift clinical-trial power by up to 50% over time-to-first-event analysis, and a new formula sizes precision-based trials from the desired confidence-interval width.","keywords":["win ratio","composite endpoint","statistical power","sample size","precision-based trial design","time-to-first-event analysis","Monte Carlo simulation","clinical trial design"],"falsifier":"Run the paper's IPHAK scenario in a Monte Carlo study: with the expected win ratio of 1.32 and the observed tie proportion, the analytic variance formula gives power 0.764 while simulation gives 0.872; reproducing and explaining that gap would settle whether the variance approximation supporting both the power comparison and the new sample-size formula is accurate enough.","tokens_in":14564,"feed_emoji":"📊","tokens_out":8662,"duration_ms":73813,"temperature":0.7,"pith_summary":"The paper argues that the win ratio—a hierarchical pairwise comparison of patients on a composite endpoint—is not only a more clinically meaningful summary than time-to-first-event analysis, but often a more powerful one. In simulations, switching from time-to-first-event analysis to the win ratio increased power by up to 50% when the effect on the highest-ranked outcome was substantial, and the win ratio usually beat single-endpoint analyses when lower-ranked outcomes carried moderate effects. The paper also proposes a precision-based sample-size formula, $N_{\\mathrm{total}} = \\frac{16 Z_{1-\\alpha/2}^2 (1+p_{\\mathrm{tie}})}{3 p_T (1-p_T)(1-p_{\\mathrm{tie}}) \\, \\mathrm{Width}^2}$, which lets a trial be sized from the desired width of the confidence interval for $\\log(\\mathrm{WR})$ instead of from a full treatment-effect scenario. A sympathetic reader would take away that the win ratio can make composite-endpoint trials more efficient, and that precision-based win-ratio trials are feasible to plan.","feed_headline":"Win ratio lifts trial power by up to 50% over standard analysis","feed_subtitle":"A new formula sizes precision-based trials directly from the desired confidence-interval width.","key_machinery":"The carrying object is the unmatched win ratio: all $N_T \\times N_C$ patient pairs are compared hierarchically, each pair is scored as a win, loss, or tie on the first outcome in the hierarchy that distinguishes the two patients, and the win ratio is $P(\\text{win})/P(\\text{loss})$. The design machinery is the approximate variance of the log-transformed win ratio from [19], $\\operatorname{Var}(\\log \\widehat{\\mathrm{WR}}) \\approx \\frac{4(1+p_{\\mathrm{tie}})}{3 p_T(1-p_T)(1-p_{\\mathrm{tie}}) N_{\\mathrm{total}}}$, which depends only on allocation and tie probability. The paper's new sample-size formula is the algebraic inverse of this variance expression solved for $N_{\\mathrm{total}}$ at a target Wald confidence-interval width; the same variance approximation underlies the power comparisons and the sensitivity of power to the tie proportion.","core_discovery":"On its own terms, the paper establishes two results. First, for composite endpoints with a clinically meaningful hierarchy, the win ratio can deliver materially higher statistical power than single-outcome analysis or time-to-first-event analysis: in the planned IPHAK trial setting the win ratio reached simulated power 0.872 versus 0.298 for the binary endpoint alone and 0.726 for the continuous endpoint alone, and in survival simulations the win ratio beat time-to-first-event analysis by more than 50% when the hazard ratio on the top-ranked outcome was strong. Second, by inverting the approximate variance formula $\\operatorname{Var}(\\log \\widehat{\\mathrm{WR}}) \\approx \\frac{4(1+p_{\\mathrm{tie}})}{3 p_T (1-p_T)(1-p_{\\mathrm{tie}}) N_{\\mathrm{total}}}$, the paper derives a closed-form total sample size for precision-based trials that depends only on the desired confidence-interval width, the allocation proportion, and the expected tie proportion. The qualifying condition is that these gains appear when the highest-ranked outcome carries a non-negligible effect; when a continuous outcome sits at the top of the hierarchy, it tends to decide almost all pairwise comparisons and the lower-ranked outcomes add little.","pith_inferences":["An implicit consequence of the variance formula is that the required sample size is highly sensitive to the anticipated tie proportion, so a sensitivity analysis over tie proportions should accompany any use of the new formula.","The case-study gap between the analytic approximation (0.764) and simulation (0.872) suggests the approximate variance may understate power; if so, the sample-size formula would tend to under-size trials, and designs should verify target width by simulation.","The same plug-in logic could be adapted to stratified or covariate-adjusted win ratio analyses, where tie probabilities would be stratum-specific; the paper's references already supply the needed variance formulas.","For ordinal endpoints such as the modified Rankin Scale, the win ratio's nonparametric nature suggests the power gains may transfer, but the tie structure differs and would need its own simulation checks."],"forward_implications":["If the power gains hold, a trial that would be underpowered under time-to-first-event analysis can become adequately powered at the same sample size by switching to the win ratio, provided the hierarchy reflects clinical priorities and the top outcome is not nearly null.","The sample-size formula gives a way to plan precision-based win-ratio trials without simulating a full treatment-effect scenario; the inputs are the desired confidence-interval width, the treatment allocation, and an anticipated tie proportion.","The win ratio's advantage over single-endpoint analysis is selective: it appears when lower-ranked outcome effects are moderate relative to the higher-ranked outcome, and it largely disappears when a continuous outcome at the top of the hierarchy dominates the pairwise decisions.","Adding a continuous outcome at the bottom of the hierarchy can break ties and raise power, even if it attenuates the estimated win ratio.","Precision-based design is more forgiving than power-based design: a slight shortfall in precision is less consequential than labeling a trial unsuccessful, which supports using the approximation-based formula in exploratory settings."],"supporting_citations":[{"why":"Introduces the win ratio as a hierarchical pairwise comparison of composite endpoints and defines the estimand and test that the paper's design work builds on.","marker":"[12]"},{"why":"Supplies the approximate variance formula for the log-transformed win ratio that the paper uses for power comparison and inverts into its new sample-size formula.","marker":"[19]"},{"why":"Earlier simulation evidence that the win ratio can outperform the Cox model in composite time-to-event endpoints, which the paper extends to gains up to 50%.","marker":"[21]"},{"why":"Provides the pilot-data-based sample-size method for the win ratio that the paper reviews and contrasts with its own formula.","marker":"[18]"},{"why":"Rank-based simulation approach for win-ratio power that the paper compares with the variance-approximation method and finds unreliable in practice.","marker":"[20]"}],"fun_headline_variants":["Win ratio boosts power by 50% in composite trials","Precision-based sample size formula for win ratio trials","Win ratio outperforms time-to-first-event by up to 50%","New formula sizes win-ratio trials from CI width","Win ratio gains power, then sizes trials precisely"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The design tool rests on the approximate variance formula for the log win ratio being accurate enough; the new sample-size formula is just that formula solved for N, so if the approximation mis-states the variance, the required sample size is wrong.","fun_headline_variants_meta":{"raw":{"variants":["Win ratio boosts power by 50% in composite trials","Precision-based sample size formula for win ratio trials","Win ratio outperforms time-to-first-event by up to 50%","New formula sizes win-ratio trials from CI width","Win ratio gains power, then sizes trials precisely"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000884,"raw_usage":{"total_tokens":3876,"prompt_tokens":1060,"completion_tokens":2816,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":676,"completion_tokens_details":{"reasoning_tokens":2736}},"tokens_in":676,"tokens_out":2816,"duration_ms":21723,"temperature":1.0,"reasoning_tokens":2736,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T15:26:31.459603+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the paper's IPHAK scenario in a Monte Carlo study: with the expected win ratio of 1.32 and the observed tie proportion, the analytic variance formula gives power 0.764 while simulation gives 0.872; reproducing and explaining that gap would settle whether the variance approximation supporting both the power comparison and the new sample-size formula is accurate enough.","supporting_citations":[{"cited_title":"Power considerations for the win ratio: A rank-based simulation approach","cited_arxiv_id":null,"evidence_quote":"Rank-based simulation approach for win-ratio power that the paper compares with the variance-approximation method and finds unreliable in practice."}],"review_version":1}