REVIEW 3 major objections 4 minor 2 references
Using Statistical Precision Medicine to Identify Optimal Treatments in a Heart Failure Setting
T0 review · 3 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read This paper claims that data-driven individualized treatment rules for choosing between furosemide and torsemide can give heart failure patients roughly 8 to 11 more days alive within a year than a one-drug-fits-all rule.
desk verdict Solid, honest application of established precision medicine methods to furosemide vs torsemide in Medicare; the days-alive result is consistently positive, but the abstract's 9-day claim is for the wrong outcome and unmeasured confounding remains the key threat. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the treatment rule $d(x)$, a map from baseline covariates $X$ to a treatment $A \in \{-1,1\}$ (torsemide vs. furosemide). Each learner optimizes a value function $V(d) = E[\,Y\, I\{A=d(X)\}/\pi(A;X)\,]$, the inverse-probability-weighted expected reward (here, days to death or to the composite of death and heart failure readmission), with $\pi(A;X)$ the propensity score estimated by random forests. The three learners are random forests, residual weighted learning (which weights outcomes by residuals from a fitted outcome model), and efficient augmentation and relaxation learning (which adds a doubly robust correction). Censoring is handled by recursively imputed survival trees, and dimensionality is reduced by a variable-importance step using out-of-bag Gini impurity.
What would settle it
Measure an unmeasured confounder—say, a validated frailty index or physician-documented prognosis—in a validation cohort with the same inclusion criteria, re-estimate the propensity score with it included, and recompute the value-function differences. If the added days shrink toward zero or the optimal rule changes materially, the central claim is falsified as stated; alternatively, a randomized SMART trial assigning patients by the learned rule could settle the causal question directly.
Extended reading notes
Core claim
On the paper's own terms, the central discovery is that an optimal treatment rule—a map from each patient's clinical history to a choice of furosemide or torsemide—can be learned from observational claims data and outperforms a universal rule. In the real-world Medicare cohort, all three learners found positive gains in days alive within one year (7.8 to 10.7 days, with confidence intervals excluding zero), and the abstract highlights roughly 9 additional days as the headline improvement. The toy examples show that, when effect heterogeneity is present, a tailored rule can lower absolute mortality risk far below either one-treatment-fits-all option. The paper frames this as evidence that precision medicine methods built on causal inference can identify useful individualized treatment decisions in pharmacoepidemiology.
Load-bearing premise
The load-bearing premise is that the Medicare claims covariates capture all factors that affect both which loop diuretic a patient receives and how long they survive, so the inverse-probability-weighted value function identifies a true causal effect of treatment choice rather than a reflection of unmeasured differences such as frailty or physician prognosis.
Editorial extensions
If this is right
- If the estimated rules were implemented at discharge, clinicians would switch some patients from furosemide to torsemide (or vice versa) according to claims-based variables such as age, COPD, atrial fibrillation, and prior medications, with an expected gain of about a week of life in the following year.
- The similarity of results across three different learners suggests the tailoring signal is not an artifact of one algorithm, strengthening the case that the benefit is real.
- For the composite outcome of readmission or death, gains were smaller (up to 7.6 days) and two of the three confidence intervals included zero, so the strongest evidence is for marginal survival rather than readmission avoidance.
- A prospective SMART randomized trial that assigns patients using the learned rule would be the natural next test of whether these observational gains are causally reproducible.
- Because the value function requires estimating the propensity score, the rule is only as credible as the covariates used; adding richer clinical data such as measured ejection fraction or laboratory values may change the rule.
Reading between the lines
- The paper's abstract emphasizes a roughly 9-day gain for 'survival time free of heart failure readmission,' but Table 5 shows the larger and more statistically reliable gains are in days alive; readers should treat the composite-readmission result as less settled than the mortality result.
- A 7-to-11-day gain in a 365-day horizon is modest, and whether it is worth acting on depends on the cost, side effects, and patient preferences around switching diuretics, which the paper does not quantify.
- Because only furosemide and torsemide were compared and the inclusion criteria required no prior loop diuretic use in six months, the rule may not generalize to patients already on a diuretic or to those for whom bumetanide is appropriate; extending to three treatments would require multi-arm optimal-rule methods.
- The toy examples show tailoring helps most when interactions are unknown; in the real data, the ten covariates with the highest variable-importance scores (with age the most important) form a candidate list of effect modifiers that future studies could measure prospectively.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes to identify optimal individualized loop diuretic (furosemide vs. torsemide) prescribing rules for heart failure patients using three precision medicine learners: random forests, residual weighted learning, and efficient augmentation relaxed learning. After two hypothetical examples illustrating effect heterogeneity, it applies the learners to 66,741 Medicare fee-for-service beneficiaries with a recent heart failure hospitalization and a new loop diuretic prescription. The value of a treatment rule is estimated through inverse probability weighting with a propensity score estimated by random forests, and rules are evaluated with 10-fold cross-validation against a universal torsemide rule. The real-world results (Table 5) report gains of 7.8-10.7 additional days alive and 4.7-7.6 additional days alive free of heart failure readmission, with confidence intervals for the composite outcome including zero for two of the three learners. The abstract, however, states that the largest gain is approximately 9 days for survival time free of heart failure readmission, which is inconsistent with Table 5.
Significance. If the central claims were fully supported, this would be a useful demonstration of how modern individualized treatment rule methods can be applied to claims-based pharmacoepidemiology, potentially informing future decision support tools. The use of three distinct learners, real Medicare data, and a clinically meaningful outcome is a strength. However, the paper's current form contains a direct internal inconsistency between the abstract and the principal results table, an apparently erroneous duplication in the toy-example table, and no quantitative treatment of unmeasured confounding—the last being essential for any causal interpretation of the estimated value-function contrasts. These issues currently prevent the paper from supporting its headline conclusions.
major comments (3)
- [Abstract and Table 5] The abstract claims that "the improvement under the optimal treatment rule in the real-world setting is greatest (additional ~9 days under the tailored rule) for survival time free of heart failure readmission." This is contradicted by Table 5, where the largest composite-outcome gain is 7.6 days (Residual Weighted Learner, 95% CI 2.2-13.0), and the largest overall gain is 10.7 days for days alive (Efficient Augmentation Relaxed Learner). Please correct the abstract so that it accurately represents the reported results, and ensure that all text descriptions of the main findings match Table 5.
- [Table 3 (toy example)] The two panels of Table 3, labeled "Has T2DM" and "Does not have T2DM," contain identical numbers. As a result, the three-modifier toy example does not actually demonstrate any additional effect modification by T2DM, and the claimed reduction in absolute risk from 0.17 (Table 2) to 0.15 (Table 4) is not supported by the displayed data. A direct recomputation from Table 3 (treating the two panels as distinct only in name) yields a tailored-rule risk different from 0.15. Please correct the table or the associated text and recompute the corresponding risks.
- [Treatment Rule Estimation and Covariates] The causal interpretation of the value-function contrast V(d*) - V(universal torsemide) relies on the assumption that the propensity score π(A|X) includes all common causes of loop diuretic choice and the survival/composite reward. The covariates listed in the Covariates section are derived from six months of Medicare claims and do not include measured ejection fraction, renal function laboratories (creatinine/eGFR), electrolytes, BNP, functional status, or clinician's prognostic judgment. These are plausibly related to both diuretic selection and the outcome. Without a negative-control outcome, a quantitative sensitivity analysis, or at minimum a prominent and detailed discussion of this limitation, the estimated 5-11 day gains may reflect residual confounding rather than a causal effect of adopting the tailored rule. Please add such an analysis or substantially soften the causal language throughout.
minor comments (4)
- [Methods] In the 'Methods' paragraph, the text says 'We evaluated the expected value function (time to each outcome) for each precision medicine method (Table 3).' This should refer to Table 5, which contains the real-world results.
- [Treatment Rule Estimation] The displayed value function V(d) = E[Y * I{A=d(X)} / π(A;X)] contains a trailing '2' that appears to be a typographical artifact. If the intent is to use the inverse probability weighted estimator adjusted for estimated propensity scores, please write the formula without the extraneous character.
- [Censoring] The description of the RIST imputation states that the imputation was repeated 2 times; it would be helpful to briefly justify this choice, since multiple imputation typically requires more repetitions to reflect between-imputation variability.
- [Table 5] The table would be easier to interpret if it also reported the proportion of patients for whom each learner recommended torsemide rather than furosemide, as the clinical meaning of a tailored rule depends on which patients would be switched.
Circularity Check
No significant circularity: the tailored-rule comparison is an out-of-sample, cross-validated estimate, not an identity or a fitted parameter renamed as a prediction.
full rationale
The paper's central derivation is an empirical estimation pipeline. The optimal treatment rule is fit on training data and evaluated on held-out test folds ('The tailored rule is generated on the training data, and evaluated on the test data'; 'We used 10-fold cross validation'). Thus the reported value-function contrasts between the tailored rule and the universal torsemide rule are out-of-sample estimates, and the composite-outcome confidence intervals in Table 5 include zero, confirming the comparison is not forced by construction. The IPW value function V(d) is a standard causal estimator, and the paper explicitly treats the no-unmeasured-confounding requirement as an assumption of the causal framework rather than deriving it from the cited methods. The learners (RWL, EARL, random forest, RIST) are published algorithms; although several references are co-authored by the senior author, they are used as external tools, and the paper's empirical result does not reduce to those citations. The hypothetical toy example is constructed, not predicted. No equation in the paper defines an input in terms of the output it is used to derive. A separate reporting discrepancy exists—the abstract attributes the ~9-day gain to HF-readmission-free survival while Table 5 places the largest such gain at 7.6 days and the 9-11 day gains under days alive—but this is a consistency error, not circularity. The derivation chain is therefore self-contained in the sense required here.
Assumptions & free parameters
free parameters (3)
- Number of covariates retained by variable importance =
10
- RIST imputation repeats =
2
- Number of trees in RIST model =
50
assumptions (5)
- domain assumption No unmeasured confounding: all common causes of loop diuretic choice and survival are included in the Medicare claims covariates.
- domain assumption Positivity: every patient has a positive probability of receiving either furosemide or torsemide.
- standard math Consistency and no interference between patients.
- domain assumption Claims-based ejection fraction algorithm and loop diuretic dose equivalency are valid.
- domain assumption Censoring is non-informative after imputation by RIST.
Cite this review
Pith. "Pith review of Using Statistical Precision Medicine to Identify Optimal Treatments in a Heart Failure Setting." pith.science (2026). https://pith.science/paper/STZYJMRZ
@misc{pith2026250107789,
author = {Pith},
title = {Pith review of: Using Statistical Precision Medicine to Identify Optimal Treatments in a Heart Failure Setting},
year = {2026},
howpublished = {\url{https://pith.science/paper/STZYJMRZ}},
note = {Machine review of arXiv:2501.07789}
}
read the original abstract
Identifying optimal medical treatments to improve survival has long been a critical goal of pharmacoepidemiology. Traditionally, we use an average treatment effect measure to compare outcomes between treatment plans. However, new methods leveraging advantages of machine learning combined with the foundational tenets of causal inference are offering an alternative to the average treatment effect. Here, we use three unique, precision medicine algorithms (random forests, residual weighted learning, efficient augmentation relaxed learning) to identify optimal treatment rules where patients receive the optimal treatment as indicated by their clinical history. First, we present a simple hypothetical example and a real-world application among heart failure patients using Medicare claims data. We next demonstrate how the optimal treatment rule improves the absolute risk in a hypothetical, three-modifier setting. Finally, we identify an optimal treatment rule that optimizes the time to outcome in a real-world heart failure setting. In both examples, we compare the average time to death under the optimized, tailored treatment rule with the average time to death under a universal treatment rule to show the benefit of precision medicine methods. The improvement under the optimal treatment rule in the real-world setting is greatest (additional ~9 days under the tailored rule) for survival time free of heart failure readmission.
Reference graph
Works this paper leans on
-
[9]
Torsemide Versus Furosemide in Patients With Acute Heart Failure (from the ASCEND-HF Trial)
Mentz RJ, Hasselblad V, DeVore AD, et al. Torsemide Versus Furosemide in Patients With Acute Heart Failure (from the ASCEND-HF Trial). Am J Cardiol. 2016. 10. Luckett DJ, Laber EB, Kahkoska AR, Maahs DM, Mayer-Davis E, Kosorok MR. Estimating Dynamic Treatment Regimes in Mobile Health Using V-Learning. J Am Stat Assoc. 2020;115(530):692–706. 11. Luckett DJ...
work page 2016
-
[31]
Laber EB, Qian M. Generalization Error for Decision Problems. Wiley StatsRef Stat Ref Online. 2019:1–15. 32. Wager S, Athey S. Estimation and Inference of Heterogeneous Treatment Effects using Random Forests. J Am Stat Assoc. 2018;113(523):1228–1242. 33. Jiang X, Nelson AE, Cleveland RJ, et al. Precision Medicine Approach to Develop and Internally Validat...
work page 2019
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.