REVIEW 5 major objections 4 minor 21 references
In an emulated DAPA-HF trial, aggregate analysis shows no dapagliflozin mortality benefit, but HTE-guided stratification identifies a subgroup with significant harm (HR 6.68) and another with significant benefit (HR 0.20).
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 19:29 UTC pith:WOQCJBAS
load-bearing objection Real workflow, but the headline subgroup HRs are selection artifacts — a hold-out validation or sample-splitting is needed before believing them. the 5 major comments →
Optimizing Clinical Trial Protocols Using EHR-Derived Heterogeneous Treatment Effects
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central claim is that heterogeneous treatment effect (HTE) estimation can convert an apparently neutral trial emulation into clinically actionable subgroups. In a real-world emulation of DAPA-HF built from EHR data, dapagliflozin showed no significant all-cause mortality effect in the full propensity-matched cohort (HR 1.681, p=0.1507). After computing individual HTEs with a Meta-S learner and cutting them at a data-driven decision-tree threshold (-0.03501), patients below the threshold had a large protective association (HR 0.203, 95% CI 0.087–0.476) while those above had a large harmful association (HR 6.680, 95% CI 2.759–16.171). The paper reports that these directionally distinct eff
What carries the argument
The load-bearing mechanism is a two-step pipeline: (1) a Meta-S learner, i.e., a single prediction model with treatment as a covariate, generating patient-specific HTE estimates from EHR covariates; (2) a shallow decision-tree regression on a transformed outcome Y*(T - p_treatment) that learns an HTE cutoff separating predicted-benefit from predicted-harm patients. Subgroups are then validated with propensity-score matched Cox models, and the subgroup labels are converted into interpretable eligibility criteria through decision-tree rule discovery.
Load-bearing premise
The decision-tree HTE cutoff is learned from the same patients whose outcomes are then used to test it, so the subgroup hazard ratios are not independent discoveries.
What would settle it
Apply the same threshold learned in half the cohort to the other half: if the extreme opposite hazard ratios (≈0.20 and ≈6.68) do not reproduce in the held-out half, the reported heterogeneity is an artifact of selection; a simpler signal would be the subgroup p-values drifting toward the overall null under cross-fitting.
If this is right
- If correct, aggregate trial results can hide subgroups with opposite treatment effects, so neutral trials should routinely be examined for HTE before concluding no benefit.
- HTE-derived strata could guide eligibility criteria for future trials, potentially enriching for responders and excluding those likely harmed.
- The same pipeline can be applied to other completed trials or real-world data to generate hypotheses about which baseline characteristics define benefit or harm.
- Observed subgroup HRs (0.203 and 6.680) imply effect sizes much larger than typical average trial effects, suggesting high clinical relevance if replicated.
Where Pith is reading between the lines
- The threshold learned on the full data and then used to split the same patients for validation means the reported subgroup p-values are not independent; they are likely inflated by selection on outcomes.
- A stronger test would be pre-specifying the cutoff on a training set and confirming on a held-out cohort, or using cross-fitting and sample splitting to avoid circularity.
- The harmful subgroup's HR 6.68 with a wide CI rests on only 102 dapagliflozin patients; small absolute counts make the estimate unstable and sensitive to a few events.
- If replicated in external EHR or prospective data, the approach could be used to redesign trials by excluding potential-harm patients; if the selection artifact dominates, the apparent heterogeneity will shrink or vanish upon validation.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper emulates the DAPA-HF trial using Mayo Clinic EHR data, estimates heterogeneous treatment effects with a Meta-S learner, and uses a decision-tree threshold on a transformed outcome to split patients into 'beneficial' and 'harmful' subgroups. The authors report that, although the overall emulated cohort shows no significant dapagliflozin effect (HR 1.681, p=0.1507), the HTE-defined beneficial subgroup has HR 0.203 (p=0.0002) and the harmful subgroup HR 6.680 (p<0.0001). They also propose an eligibility-criteria optimization pipeline that yields directionally opposite adjusted HRs (0.175 and 2.125). The central claim is that HTE-guided stratification recovers clinically meaningful heterogeneity hidden in aggregate trial emulation.
Significance. If the results were valid, the approach would be a useful contribution to precision trial design and real-world evidence analysis. The paper also has a sensible motivation: moving from average effects to interpretable subgroups. However, the significance is severely limited by the fact that the headline subgroup effects are not independent tests: the same data are used to learn the HTE threshold and to estimate the within-subgroup Cox HRs. The extreme treatment imbalance (149 versus 8,865) and the absence of propensity-matching details further undermine the reliability of the point estimates. No code, no hold-out validation, and no external replication are provided, so the paper does not currently support its claim of identifying clinically actionable subgroups.
major comments (5)
- [§4.4 and §2.2.2/2.2.4] The decision-tree threshold (-0.03501) is learned by training on the transformed outcome Y* = Y × (T/p − (1−T)/(1−p)), which directly contains the mortality outcome Y and treatment assignment T. The same patients' outcomes are then used in §2.2.4 to fit Cox models within the resulting subgroups. The reported HRs (0.203 and 6.680) and p-values are therefore conditional on a data-chosen split and are not valid unconditional tests. A hold-out validation or sample-splitting procedure is required before any inferential claim can be made.
- [Table 1 and §4.2] The treatment groups are extremely unbalanced: 149 dapagliflozin patients versus 8,865 placebo patients. The beneficial and harmful subgroups contain only 47 and 102 treated patients, respectively, so the subgroup HRs are driven by a very small number of events. The manuscript does not report the propensity-score model specification, matching ratio, caliper, or balance diagnostics. These omissions are load-bearing because the survival estimates depend entirely on the quality of confounding adjustment in this highly imbalanced setting.
- [§1 vs §4.1] The paper contains an internal contradiction about the DAPA-HF trial's all-cause mortality result. §1 states that dapagliflozin was associated with a significant reduction in all-cause mortality (HR 0.83, p=0.0217), whereas §4.1 states that the original trial 'found no statistically significant difference between the two treatments.' This inconsistency affects the interpretation of the emulation benchmark and needs to be resolved.
- [§2.2.5 and §4.5] The eligibility-criteria optimization introduces a second, additional layer of selection. The optimized beneficial and harmful cohorts are chosen using HTE-derived labels as reference targets and are then re-evaluated through the full emulation workflow. The resulting HRs (0.175 and 2.125) are selected estimates, subject to the same circularity as the threshold-based subgroups, with no correction for multiple comparisons or independent validation. The claim that these are 'clinically actionable' criteria is not supported by the analysis.
- [Discussion (§3)] The paper acknowledges that external validation is necessary, but it does not temper the abstract and results sections, which present the directionally distinct subgroup HRs as established findings. Given the selection-circularity, the extreme imbalance, the internal contradiction about DAPA-HF, and the lack of any hold-out or external validation, the central claim is not supported as stated.
minor comments (4)
- [Throughout] Several grammatical issues: 'one independent modelling strategies' in §4.3, 'Meta-S' capitalization inconsistency, and the transformed-outcome formula in §4.4 appears garbled (should read Y* = Y × (T/p − (1−T)/(1−p))).
- [§2.2.3] The text mentions 'two distinct HTE estimation strategies,' but only the Meta-S learner is described in the methods. Clarify whether a second learner was used or remove the reference.
- [Figure 1] The figure caption says panels (b) and (c) present stratified curves, but the labels in the figure are not described precisely enough to map them to 'beneficial' and 'harmful.'
- [Table 1] Several rows report '9,014 (100.00%)' for variables like 'Elevated NT-proBNP,' which is likely an artifact of how the EHR eligibility criteria were operationalized. This should be explained, as it seems implausible that all patients had elevated NT-proBNP.
Circularity Check
Subgroup HRs are fit on the same outcome data used to choose the HTE threshold, so the headline effects are selected, not predicted.
specific steps
-
fitted input called prediction
[§2.2.1–2.2.2, §2.2.4, §4.4]
"we applied a decision tree–based thresholding approach, in which a shallow regression tree was trained on the transformed outcome, incorporating both mortality events and treatment assignment, to learn a data-driven HTE cutoff. ... The final threshold is -0.03501 (HTE value). ... In contrast, the beneficial subgroup ... (HR = 0.203, 95% CI, 0.087–0.476, p = 0.0002), whereas the harmful subgroup ... (HR = 6.680, 95% CI, 2.759–16.171, p < 0.0001)."
The decision-tree cutoff is fit to a pseudo-outcome that explicitly contains mortality Y and treatment assignment T, so the split at -0.03501 is a function of the same full-cohort outcomes that are later used to compute the subgroup Cox HRs. The 'beneficial' and 'harmful' labels are therefore not independent predictions: they are partitions selected by an outcome-weighted transform, and the reported HRs and p-values are conditional on that data-chosen split. No held-out or external sample is used before the headline p-values are presented.
-
fitted input called prediction
[§4.5; results in §2.2.5]
"Qualified candidates were then evaluated using the same downstream analysis framework, including propensity-score matching and adjusted Cox proportional-hazards modeling. Final beneficial candidates were required to have HR < 1 and (p < 0.05), whereas final harmful candidates were required to have HR > 1 and (p < 0.05)."
The optimized eligibility criteria are selected precisely because they produce a significant Cox HR on the same mortality endpoint used to build the HTE threshold. Reporting the selected candidates' HRs (0.175 and 2.125) as validation re-uses the outcome to choose the cohort and then to estimate the effect in that cohort. This is outcome-based selection, not an independent confirmation of the optimized subgroups.
full rationale
Most of the pipeline is standard and the comparison to the DAPA-HF benchmark is a useful external reference, but the two headline subgroup findings are not independent of their fitting inputs. In §4.4 the HTE cutoff is learned from a pseudo-outcome that explicitly embeds the mortality outcome and treatment indicator; the same data are then used in §2.2.4 to compute Cox HRs inside the resulting splits. This makes the reported HRs (0.203 and 6.680) selected statistics: the split has already been optimized against a function of Y and T. The eligibility-criteria optimization in §4.5 is doubly selective, since candidate cohorts are retained only if their Cox HR is significant on the same endpoint. I am not treating the self-citation to [20] as load-bearing for the subgroup claim, and the inconsistent description of DAPA-HF's own mortality result is a correctness/evidence concern rather than circularity. The paper's own limitation statement—'Prospective evaluation and external validation across other data sources are necessary to determine whether these HTE-derived partitions truly define groups'—confirms that the in-sample results are not yet independently validated. The score is 7 rather than 8–10 because the Cox HR magnitudes are not algebraically identical to the fitted threshold; a properly designed out-of-sample or external validation could, in principle, support the claim. As written, however, the central claim reduces to in-sample outcome-based selection.
Axiom & Free-Parameter Ledger
free parameters (4)
- HTE threshold =
-0.03501
- Meta-S learner hyperparameters =
not reported
- Propensity score model specification =
not reported
- Eligibility rule filtering thresholds =
not reported
axioms (5)
- domain assumption Electronic health records can emulate a randomized trial with valid causal interpretation.
- domain assumption Unconfoundedness after propensity-score matching: treatment assignment is independent of potential outcomes given measured covariates.
- standard math The transformed outcome Y* = Y × (T/p_treat - (1-T)/(1-p_treat)) is a valid pseudo-outcome for treatment-effect estimation.
- ad hoc to paper Subgroup assignment based on HTE estimates learned from the same data can be used to validate treatment effects within those subgroups.
- domain assumption All-cause mortality is captured in the EHR without informative censoring or outcome misclassification.
read the original abstract
Traditional randomized trials often obscure clinically meaningful heterogeneity in treatment response by focusing on average effects. Leveraging real-world data to emulate clinical trials and estimate heterogeneous treatment effects (HTEs) offers a promising path toward more precise and efficient trial design. In this study, we emulate the DAPA-HF trial using electronic health records from the Mayo Clinic Cloud (MCC) to investigate whether HTE-guided stratification can identify patient subgroups with distinct treatment responses to dapagliflozin versus placebo in patients with heart failure with reduced ejection fraction. All-cause mortality was evaluated using Cox proportional hazards models, with HTEs estimated using a Meta-S learner and subgroups defined using a decision tree-based thresholding approach. In the overall cohort of the emulation, no significant treatment difference was observed (HR, 1.681; 95% CI, 0.828-3.413; p = 0.1507). However, compared with the overall emulated cohort, in which dapagliflozin showed no statistically significant survival benefit, HTE-driven stratification identified subgroups with significant and directionally distinct treatment effects. The beneficial (low-HTE) subgroup showed a significant survival benefit from dapagliflozin (HR = 0.203, 95% CI, 0.087-0.476, p = 0.0002), whereas the harmful (high-HTE) subgroup showed a significant harmful association with markedly increased mortality risk (HR = 6.680, 95% CI, 2.759-16.171, p < 0.0001). These findings indicate that HTE-guided stratification can uncover clinically meaningful beneficial and harmful treatment-effect patterns that are masked in the full-cohort emulation.
Reference graph
Works this paper leans on
-
[1]
Assessing the gold standard—lessons from the history of RCTs
Bothwell LE, Greene JA, Podolsky SH, Jones DS. Assessing the gold standard—lessons from the history of RCTs. N engl j med. 2016;374(22):2175-81
2016
-
[2]
Why do we need some large, simple randomized trials? Statistics in medicine
Yusuf S, Collins R, Peto R. Why do we need some large, simple randomized trials? Statistics in medicine. 1984;3(4):409-20
1984
-
[3]
Challenges and lessons learned from COVID-19 trials: should we be doing clinical trials differently? Canadian Journal of Cardiology
Janiaud P, Hemkens LG, Ioannidis JP. Challenges and lessons learned from COVID-19 trials: should we be doing clinical trials differently? Canadian Journal of Cardiology. 2021;37(9):1353-64
2021
-
[4]
A literature review on the representativeness of randomized controlled trial samples and implications for the external validity of trial results
Kennedy-Martin T, Curtis S, Faries D, Robinson S, Johnston J. A literature review on the representativeness of randomized controlled trial samples and implications for the external validity of trial results. Trials. 2015;16(1):495
2015
-
[5]
to whom do the results of this trial apply?
Rothwell PM. External validity of randomised controlled trials:“to whom do the results of this trial apply?”. The Lancet. 2005;365(9453):82-93
2005
-
[6]
Evaluating eligibility criteria of oncology trials using real-world data and AI
Liu R, Rizzo S, Whipple S, Pal N, Pineda AL, Lu M, et al. Evaluating eligibility criteria of oncology trials using real-world data and AI. Nature. 2021;592(7855):629-33
2021
-
[7]
Broadening eligibility criteria and diversity among patients for cancer clinical trials
Kaur M, Frahm F, Lu Y, Ascha MS, Guadamuz JS, Dotan E, et al. Broadening eligibility criteria and diversity among patients for cancer clinical trials. NEJM evidence. 2024;3(4):EVIDoa2300236
2024
-
[8]
GIST 2.0: A scalable multi-trait metric for quantifying population representativeness of individual clinical studies
Sen A, Chakrabarti S, Goldstein A, Wang S, Ryan PB, Weng C. GIST 2.0: A scalable multi-trait metric for quantifying population representativeness of individual clinical studies. Journal of biomedical informatics. 2016;63:325-36
2016
-
[9]
Using big data to emulate a target trial when a randomized trial is not available
Hernán MA, Robins JM. Using big data to emulate a target trial when a randomized trial is not available. American journal of epidemiology. 2016;183(8):758-64
2016
-
[10]
Emulating randomized clinical trials with nonrandomized real-world evidence studies: first results from the RCT DUPLICATE initiative
Franklin JM, Patorno E, Desai RJ, Glynn RJ, Martin D, Quinto K, et al. Emulating randomized clinical trials with nonrandomized real-world evidence studies: first results from the RCT DUPLICATE initiative. Circulation. 2021;143(10):1002-13
2021
-
[11]
Real-world evidence and real-world data for evaluating drug safety and effectiveness
Corrigan-Curay J, Sacks L, Woodcock J. Real-world evidence and real-world data for evaluating drug safety and effectiveness. Jama. 2018;320(9):867-8
2018
-
[12]
Emulation of randomized clinical trials with nonrandomized database analyses: results of 32 clinical trials
Wang SV, Schneeweiss S, Franklin JM, Desai RJ, Feldman W, Garry EM, et al. Emulation of randomized clinical trials with nonrandomized database analyses: results of 32 clinical trials. Jama. 2023;329(16):1376-85
2023
-
[13]
Emulate randomized clinical trials using heterogeneous treatment effect estimation for personalized treatments: Methodology review and benchmark
Ling Y, Upadhyaya P, Chen L, Jiang X, Kim Y. Emulate randomized clinical trials using heterogeneous treatment effect estimation for personalized treatments: Methodology review and benchmark. Journal of biomedical informatics. 2023;137:104256
2023
-
[14]
Simulating Colorectal Cancer Trials Using Real-World Data
Chen Z, Zhang H, George TJ, Guo Y, Prosperi M, Guo J, et al. Simulating Colorectal Cancer Trials Using Real-World Data. JCO Clinical Cancer Informatics. 2022;6:e2100195
2022
-
[15]
Electronic medical records can be used to emulate target trials of sustained treatment strategies
Danaei G, Rodríguez LAG, Cantero OF, Logan RW, Hernán MA. Electronic medical records can be used to emulate target trials of sustained treatment strategies. Journal of clinical epidemiology. 2018;96:12-22
2018
-
[16]
Estimation and inference of heterogeneous treatment effects using random forests
Wager S, Athey S. Estimation and inference of heterogeneous treatment effects using random forests. Journal of the American Statistical Association. 2018;113(523):1228-42
2018
-
[17]
Metalearners for estimating heterogeneous treatment effects using machine learning
Künzel SR, Sekhon JS, Bickel PJ, Yu B. Metalearners for estimating heterogeneous treatment effects using machine learning. Proceedings of the national academy of sciences. 2019;116(10):4156-65
2019
-
[18]
Learning representations for counterfactual inference
Johansson F, Shalit U, Sontag D, editors. Learning representations for counterfactual inference. International conference on machine learning; 2016: PMLR
2016
-
[19]
Dapagliflozin in patients with heart failure and reduced ejection fraction
McMurray JJ, Solomon SD, Inzucchi SE, Køber L, Kosiborod MN, Martinez FA, et al. Dapagliflozin in patients with heart failure and reduced ejection fraction. New England Journal of Medicine. 2019;381(21):1995-2008
2019
-
[21]
Rapid and intensive guideline-directed medical therapy for heart failure: 5 core principles
Greene SJ, Butler J, Fonarow GC. Rapid and intensive guideline-directed medical therapy for heart failure: 5 core principles. Circulation. 2024;150(6):422-4
2024
-
[22]
Machine learning methods for estimating heterogeneous causal effects
Athey S, Imbens GW. Machine learning methods for estimating heterogeneous causal effects. stat. 2015;1050(5):1-26. 21
2015
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.