{"id":"ad0e1895-2e55-473f-9926-d887a86dc35b","arxiv_id":"2608.10728","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A doubly robust, efficient win-ratio regression estimator that replaces censored future pairwise scores with their conditional expectation given observed history.","lead":"This paper develops a new estimator for win-ratio regression in observational studies, using future-score correction to recover information lost to censoring. The method combines inverse probability weighting with outcome augmentation to achieve double robustness and, under correct nuisance models, semiparametric efficiency.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Future-score correction's practical value rests on at least one of the censoring model or the Markov transition model being correctly specified; the EHR application likely violates both, and the simulations do not test this double-misspecification scenario.","rationale":"I read the paper in good faith and focused on the central claim that the AIPW-FC estimator is doubly robust, asymptotically normal, and efficient when all nuisance functions are correctly specified. The main theorems (3.1, 3.2, 4.3, 4.5) appear coherent, and the U-statistic inference machinery is standard. The weak point is practical: the future-score projection Phi is estimated by a Markov recursion that is likely misspecified in EHR data, and the double-robustness guarantee does not cover simultaneous misspecification of G and Phi. The simulations do not test this case, so the claimed efficiency gain may not hold under realistic conditions. This is not a mathematical inconsistency but a load-bearing limitation for the application and for the practical relevance of the efficiency gain. The reader's weakest_assumption identified exactly this concern, and the conditional verdict is appropriate. A concrete double-misspecification simulation would settle whether the concern lands.","tokens_in":33355,"tokens_out":29429,"duration_ms":334926,"concrete_test":"Run a Monte Carlo simulation under the Section C.2 DGM at 50% censoring, but with both the censoring model and the death/hospitalization transition models misspecified by omitting Z_cvd (a predictor of both censoring and transitions). Estimate the AIPW-FC coefficients and compare bias to the true beta0 = (0.584, -0.538, -0.274). If any component's bias exceeds the Monte Carlo standard error (about 0.02) or 95% coverage drops below 0.90, the double-robustness guarantee is not operational when both nuisances are wrong, and the EHR results should be interpreted as dependent on the Markov and censoring models being adequate.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing concern is that the double-robustness guarantee in Theorem 3.2 requires at least one of the fitted censoring survival G or the fitted future-score projection Phi to be correctly specified. In the implementation (Section 3.3, Eqs. (3.4)-(3.5); Appendix D.2), Phi is computed by deterministic recursion under one-step Markov transition models for death and hospitalization fitted to the same data. In the OneFlorida application, both G and Phi are working models estimated from the same cohort; the Markov assumption (state (D,H) plus baseline covariates) is likely violated, so neither may be correctly specified. When both are misspecified, E{K12(beta0;eta)} != 0 with O(1) bias because the second-order drift bound in Theorem 3.2 does not vanish. The simulation study (Section C.3) tests misspecification of G alone or Phi alone, but not the double-misspecification scenario that is most relevant to the EHR analysis. The headline efficiency gain (RE up to 1.50 under 65% censoring) is demonstrated only under correctly specified transition models, so it does not establish that the gain survives realistic misspecification.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops a win-ratio regression framework for prioritized composite outcomes under right censoring, centered on a future-score correction that replaces unobserved future pairwise score contributions by their conditional expectation given observed pair histories. A complete-data target is defined through an integrated residual process (Eq. 2.5), and an observed-data estimating equation is built by combining inverse-probability-of-censoring weighting, inverse-propensity treatment weighting, and baseline outcome augmentation (Eq. 3.2). The main theoretical claims are double robustness for treatment assignment and censoring (Theorems 3.1 and 3.2), consistency and asymptotic normality of the cross-fitted U-statistic estimator (Proposition 4.2 and Theorem 4.3), and semiparametric efficiency when all nuisance functions are correctly specified (Corollary 4.5). The paper reports simulations at 30%, 50%, and 65% censoring, with relative efficiency gains up to 1.50 for AIPW-FC, and an application to OneFlorida breast cancer data comparing adjuvant chemotherapy groups under a death-before-hospitalization priority rule.","tokens_in":33503,"tokens_out":12396,"duration_ms":120075,"significance":"If the theoretical claims are fully justified, the paper makes a useful contribution to win-ratio methodology for observational studies with censored prioritized outcomes: it defines a clean complete-data estimand, provides a principled way to use information from unresolved pairs, and supplies a full set of proofs in the appendix. The simulation design is carefully benchmarked against an independently computed target, and the application illustrates the method on a real EHR cohort. The main caveats are that the efficiency claim rests on an unverified tangent-space assertion, the consistency proof for estimated nuisance functions has a gap as written, and the practical double-robustness guarantee is not examined under the double-misspecification scenario most relevant to the EHR application.","major_comments":[{"comment":"The semiparametric efficiency result hinges on the sentence \"For the nonparametric observed-data model considered here, the closure of the observed-data tangent space at P0 is L2_0(P0)\" (Section B.4). This assertion is stated without proof or reference. Because the model is defined under sequential independent censoring and treatment exchangeability, it is not obvious whether these are restrictions on the statistical model or merely identifying assumptions on P0; the tangent space could be a proper subspace in the former case. Please provide a rigorous justification or a precise citation to standard nonparametric censoring-model tangent-space results.","section":"Section B.4, Proposition 4.1 and Corollary 4.5"},{"comment":"The consistency proof asserts that because the evaluation pairs are independent of the cross-fitted nuisance fits, one may replace \\hat\\eta by \\eta_{P0}(\\beta) directly: the displayed equation \"\\hat U_n(\\beta) = 1/(n(n-1)) \\sum K_{ij}{\\beta;\\eta_{P0}(\\beta)} + o_p(1)\" does not follow from independence alone. One needs an additional argument, for example using the second-order drift bound in Theorem 3.2 together with uniform L2 consistency of the nuisance estimators and the Lipschitz condition in Assumption B.8, to show that the random-nuisance term is o_p(1) uniformly on \\Theta. As written this is a gap in a central proof.","section":"Section B.5, proof of Proposition 4.2"},{"comment":"The double-robustness guarantee in Theorem 3.2 requires at least one of the censoring survival G and the future-score projection \\Phi to be correctly specified. In the implementation, \\Phi is obtained from one-step Markov transition models for death and hospitalization fitted to the same data (Section 3.3, Eqs. (3.4)-(3.5); Appendix D.2). The misspecification study in Section C.3 varies G alone or \\Phi alone, but not both. When both are misspecified, the second-order drift bound in Theorem 3.2 is O(||G-G0|| ||\\Phi-\\Phi0||), which need not vanish, so the bias can be O(1). The paper should either add a double-misspecification simulation, especially under the kind of realistic misspecification expected in the OneFlorida application, or explicitly state this limitation and temper the wording of the robustness claims.","section":"Section 3.3 and Appendix C.3"}],"minor_comments":[{"comment":"The phrase \"double robustness for treatment assignment and censoring\" could be more precise: it means either the propensity or the outcome model is correct, and either the censoring model or the future-score projection is correct, not robustness to arbitrary simultaneous misspecification.","section":"Abstract and Section 3.2"},{"comment":"For AIPW-FC, the table reports two ASE and coverage values (model-based and bootstrap) without an explicit column label in the header; a footnote such as \"model/bootstrap\" would make the table easier to read.","section":"Table 1 and Section 5"},{"comment":"The sentence reporting \"the common benchmark value ... for all three censoring scenarios\" should note that this equality is by construction, since the complete-data target is defined free of censoring.","section":"Section C.2"},{"comment":"The notation \\hat\\eta_{ij} is introduced only in the sentence following Eq. (4.1); consider defining it immediately before the estimating equation to avoid ambiguity.","section":"Section 4, after Eq. (4.1)"},{"comment":"A few reference formatting issues, such as \"WANG Hongyue\" in the text and the reference list, should be cleaned up for journal style.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is in scope for a statistics methodology journal and the core idea is interesting. The main work for a revised version is to fill the proof gaps around the tangent-space assertion and the estimated-nuisance step in the consistency proof, and to address the untested double-misspecification scenario. If the authors can add those justifications and a corresponding simulation or explicit limitation, the paper would likely be acceptable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The core idea here is genuinely new: when censoring cuts off the future portion of a pairwise score, replace it with its conditional expectation given the observed histories. That future-score correction, combined with inverse treatment weighting and outcome augmentation, gives a doubly robust estimating equation with a clean U-statistic theory. The proofs in the appendix are thorough and, as far as I can tell, internally consistent. The simulation results showing relative efficiency gains up to 1.50 under 65% censoring are plausible and the bootstrap inference matches the plug-in variance estimator well in the correctly specified scenarios.\n\nThe soft spot is one the authors actually half-acknowledge but do not confront. Theorem 3.2 gives double robustness in the usual sense: you need either the censoring survival G or the future-score projection Phi to be correct, and separately either the propensity e or the outcome regression m to be correct. In the EHR application, both G and Phi are working models estimated from the same cohort, and the transition model for Phi is a simple one-step Markov recursion over four states. Nothing in the paper suggests either is correctly specified. When both are wrong, the product of misspecification errors is not negligible; the bias is first order. The simulation study only tests one-side misspecification at a time, not the double-misspecification that is most relevant to the application. That is a genuine gap in evidence for the practical claims.\n\nA second, related limitation is that the efficiency gain (RE up to 1.50) is demonstrated only under correct nuisance models. Under misspecification we do not know how much of that precision benefit survives, or whether the point estimates still target the intended quantity. This is not a fatal flaw in the methodology, but it should be stated more prominently than the abstract suggests.\n\nMinor points: the semiparametric efficiency claim relies on the standard tangent-space closure assertion, which is accepted but not proved; and no code or data are provided, so independent replication is currently impossible. The treatment-misspecification simulations showed AIPW-FC correcting bias from a bad propensity model, which is reassuring.\n\nThe paper is a real methodological contribution for win-ratio regression in the presence of censoring and confounding. The formal theory is a step beyond the existing IPCW literature, and the future-score construction is worth taking seriously. The main fixes needed are a double-misspecification simulation, a more cautious framing of the efficiency gains, and release of code. I would send it to a serious statistics journal for peer review; with those additions it could be a solid publication.","headline":"Future-score correction is a real new idea with solid theory, but the paper underestimates how much of its practical value hinges on at least one of the censoring or transition models being right.","tokens_in":34092,"tokens_out":1778,"would_cite":true,"duration_ms":21937,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62G05","62G20","62P10"],"pacs":[],"model":"deepseek-v4-flash","headline":"Future-score correction lets censored pairs keep contributing to win-ratio regression.","keywords":["win-ratio regression","prioritized composite outcomes","future-score correction","double robustness","inverse probability of censoring weighting","semiparametric efficiency","U-statistics","electronic health records"],"falsifier":"Simulate a treated-versus-control cohort with 65% censoring where censoring and the future transition model both omit a strong shared predictor (e.g., a frail subgroup with higher death, hospitalization, and censoring rates), fit AIPW-FC with both misspecified, and check whether the estimated coefficient drifts from the known complete-data target; if the bias does not vanish, the double-robustness claim fails in that regime.","tokens_in":33102,"feed_emoji":"⚖️","tokens_out":8974,"duration_ms":73587,"temperature":0.7,"pith_summary":"This paper proposes a regression method for prioritized composite outcomes—outcomes like death-before-hospitalization—in observational studies where follow-up is censored. The target is a time-constant logit coefficient that balances a duration-weighted pairwise residual over follow-up, and the observed-data estimating equation is built so that a pair continues to contribute information after censoring. The mechanism is future-score correction: when censoring stops a pair's later comparisons from being observed, the missing portion of the pairwise score is replaced by its conditional expectation given the observed pair histories. Treatment weighting and baseline outcome augmentation are added, and the paper proves the estimator is doubly robust for both treatment assignment and censoring, asymptotically normal, and semiparametrically efficient when nuisance functions are correct. Simulations show the correction's efficiency gain grows with censoring, reaching about 1.5-fold at 65% censoring, and an electronic health record application illustrates the method.","feed_headline":"Future-score correction lifts win-ratio efficiency 1.5-fold","feed_subtitle":"New estimator recovers information from unresolved comparisons; the gain grows with censoring.","key_machinery":"The load-bearing object is the future-score projection $$\\Phi_{ij,\\ell}(\\$\\beta$) = E\\left\\{\\sum_{r>\\ell} S^\\circ_{ij,r}(\\$\\beta$) \\,\\middle|\\, H^+_{ij,\\ell}\\right\\},$$ the $L^2$-best predictable predictor of the remaining pairwise score after interval $\\ell$. It enters the censoring-corrected score $$$S^{{FC}}$_{ij}(\\$\\beta$;G,\\Phi)=\\sum_{\\ell=1}^M \\frac{Y_{ij}(t_{\\ell-1})}{G_{ij,\\ell-1}}S^\\circ_{ij,\\ell}(\\$\\beta$)+\\sum_{\\ell=1}^{M-1}\\frac{Y_{ij}(t_{\\ell-1})}{G_{ij,\\ell}}\\Phi_{ij,\\ell}(\\$\\beta$)\\,dM^C_{ij,\\ell}(G),$$ whose first term is standard pairwise inverse-probability-of-censoring weighting and whose second term is the future-score correction. The augmented kernel multiplies the centered censoring-corrected score by the inverse-propensity treatment weight and adds back the baseline outcome regression $m$, which is what produces double robustness for treatment and censoring. The projection is computed either by deterministic forward recursion under fitted one-step death and hospitalization transition models on a finite state space, or by Monte Carlo draws from the fitted future-trajectory distributions.","core_discovery":"The paper's central claim is that the censoring-corrected estimating equation based on the augmented kernel $$K_{ij}(\\$\\beta$;\\eta) = \\frac{A_i(1-A_j)}{e(X_i)(1-e(X_j))}\\{$S^{{FC}}$_{ij}(\\$\\beta$;G,\\Phi)-m(X_i,X_j;\\$\\beta$)\\}+m(X_i,X_j;\\$\\beta$)$$ has the same population root $\\beta_0$ as the complete-data target $E\\{\\varphi^\\circ_{12}(\\beta)\\}=0$. When censoring removes a pair from future risk, the correction term $\\Phi_{ij,\\ell}$, the conditional expectation of the remaining complete-data score given the observed pair history, fills in the lost contribution through a telescoping inverse-weighting identity. As a result, the estimator remains consistent if either the censoring survival model or the future-score projection is correct, and if either the propensity score or the baseline outcome regression is correct; at the true nuisance functions the estimator is regular and attains the semiparametric efficiency bound. In simulations the correction reduces variance relative to inverse-probability weighting alone, with relative efficiency up to 1.50 under 65% censoring, while point estimates stay near the complete-data target.","pith_inferences":["A testable extension is to apply the same future-score correction to other win statistics, such as win odds or net benefit, wherever the unobserved future part of a comparison has a predictable conditional mean given the observed pair history.","Because the correction exploits predictability of the future, the efficiency gain should shrink when the transition models lose predictive power; comparing AIPW-FC against AIPW under deliberately weak transition models would quantify how much of the reported 1.5-fold gain depends on the quality of the fitted recursion.","If a future study has access to auxiliary longitudinal markers that predict hospitalization, conditioning the future-score projection on them should amplify the efficiency gain, since the projection becomes more accurate; this is a direct implication of the mechanism that the paper does not simulate."],"forward_implications":["Under heavy censoring, unresolved pairs no longer have to be discarded: simulations show the correction's efficiency gain grows from about 1.1 at 30% censoring to 1.5 at 65% censoring, with near-nominal coverage.","The same regression coefficient can be estimated from observational EHR-style data with a causal treatment interpretation, because the estimator is consistent when either propensity or outcome regression is correct and either censoring model or future-score projection is correct.","Wald inference from the empirical Hoeffding projection gives valid confidence intervals, so bootstrap resampling is not the only option for standard errors in win-ratio regression.","In the OneFlorida breast cancer analysis, future-score correction cuts the standard error of the treatment coefficient from 0.21 to 0.18 at three years and leaves point estimates essentially unchanged, indicating the gain comes from recovered information rather than a shifted estimand."],"supporting_citations":[{"why":"Supplies the pairwise-comparison approach to prioritized composite outcomes that the paper generalizes to regression.","marker":"Finkelstein and Schoenfeld, 1999"},{"why":"Defines generalized pairwise comparisons, the framework from which the paper's win, loss, and tie processes are drawn.","marker":"Buyse, 2010"},{"why":"Introduces the win ratio and the death-before-hospitalization priority rule used in the paper's simulations and application.","marker":"Pocock et al., 2012"},{"why":"Provides the inverse-probability-of-censoring weighted win statistic that the future-score correction extends.","marker":"Dong et al., 2020"},{"why":"Defines proportional win-fractions regression, whose coefficient the complete-data target recovers under a time-constant proportional structure.","marker":"Mao and Wang, 2021"},{"why":"Introduces augmented inverse probability weighting, the source of the treatment-weighting and outcome-augmentation components.","marker":"Robins et al., 1994"},{"why":"Establishes doubly robust estimation with estimated nuisance functions, the template for the AIPW-FC kernel.","marker":"Bang and Robins, 2005"},{"why":"Provides the U-statistic inference framework used to derive asymptotic normality and variance estimation.","marker":"Kowalski and Tu, 2008"},{"why":"Demonstrates that the unadjusted win ratio depends on censoring, motivating the complete-data target over follow-up.","marker":"Oakes, 2016"},{"why":"Supplies the uniform U-process limit theorems used to establish consistency of the regression estimator.","marker":"Arcones and Giné, 1993"}],"fun_headline_variants":["Win-ratio gains 50% efficiency via future-score correction","Censored wins recovered: new estimator gains 1.5x","Future-score correction boosts win-ratio efficiency at high censoring","Recover lost pair info: win-ratio efficiency up 50% at 65% censoring"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The estimator is consistent only if at least one of the fitted censoring model or the fitted future-score projection is correct, and at least one of the treatment propensity model or the outcome regression is correct; in the paper's implementation the future-score projection is built from one-step death and hospitalization transition models fitted to the same data, so if censoring and those transition models are both wrong, the estimating equation can be biased.","fun_headline_variants_meta":{"raw":{"variants":["Win-ratio gains 50% efficiency via future-score correction","Censored wins recovered: new estimator gains 1.5x","Future-score correction boosts win-ratio efficiency at high censoring","Recover lost pair info: win-ratio efficiency up 50% at 65% censoring"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000217,"raw_usage":{"total_tokens":1475,"prompt_tokens":1023,"completion_tokens":452,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":639,"completion_tokens_details":{"reasoning_tokens":371}},"tokens_in":639,"tokens_out":452,"duration_ms":4913,"temperature":1.0,"reasoning_tokens":371,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T18:33:53.827513+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate a treated-versus-control cohort with 65% censoring where censoring and the future transition model both omit a strong shared predictor (e.g., a frail subgroup with higher death, hospitalization, and censoring rates), fit AIPW-FC with both misspecified, and check whether the estimated coefficient drifts from the known complete-data target; if the bias does not vanish, the double-robustness claim fails in that regime.","supporting_citations":[],"review_version":1}