Pith. sign in

REVIEW 5 major objections 4 minor 21 references

In an emulated DAPA-HF trial, aggregate analysis shows no dapagliflozin mortality benefit, but HTE-guided stratification identifies a subgroup with significant harm (HR 6.68) and another with significant benefit (HR 0.20).

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 19:29 UTC pith:WOQCJBAS

load-bearing objection Real workflow, but the headline subgroup HRs are selection artifacts — a hold-out validation or sample-splitting is needed before believing them. the 5 major comments →

arxiv 2607.16934 v1 pith:WOQCJBAS submitted 2026-07-18 stat.AP cs.AIcs.LG

Optimizing Clinical Trial Protocols Using EHR-Derived Heterogeneous Treatment Effects

classification stat.AP cs.AIcs.LG MSC 62P1062N0262N03
keywords heterogeneous treatment effectstrial emulationreal-world evidenceelectronic health recordsheart failuredapagliflozinsubgroup analysissurvival analysis
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Emulating the DAPA-HF trial with electronic health records, the paper finds no significant overall mortality difference between dapagliflozin and placebo, then argues this aggregate null is misleading. Using a Meta-S learner to estimate individual heterogeneous treatment effects and a decision-tree threshold to split patients, the authors identify a low-HTE subgroup with a large protective association (HR 0.203) and a high-HTE subgroup with a large harmful association (HR 6.680). The central thesis is that HTE-guided stratification can recover clinically meaningful beneficial and harmful treatment patterns that average effects hide, and that these patterns can be translated into refined eligibility criteria for future trials. A sympathetic reader would care because this offers a data-driven route toward precision trial design and more individualized treatment decisions, provided the subgroup findings survive independent validation.

Core claim

The central claim is that heterogeneous treatment effect (HTE) estimation can convert an apparently neutral trial emulation into clinically actionable subgroups. In a real-world emulation of DAPA-HF built from EHR data, dapagliflozin showed no significant all-cause mortality effect in the full propensity-matched cohort (HR 1.681, p=0.1507). After computing individual HTEs with a Meta-S learner and cutting them at a data-driven decision-tree threshold (-0.03501), patients below the threshold had a large protective association (HR 0.203, 95% CI 0.087–0.476) while those above had a large harmful association (HR 6.680, 95% CI 2.759–16.171). The paper reports that these directionally distinct eff

What carries the argument

The load-bearing mechanism is a two-step pipeline: (1) a Meta-S learner, i.e., a single prediction model with treatment as a covariate, generating patient-specific HTE estimates from EHR covariates; (2) a shallow decision-tree regression on a transformed outcome Y*(T - p_treatment) that learns an HTE cutoff separating predicted-benefit from predicted-harm patients. Subgroups are then validated with propensity-score matched Cox models, and the subgroup labels are converted into interpretable eligibility criteria through decision-tree rule discovery.

Load-bearing premise

The decision-tree HTE cutoff is learned from the same patients whose outcomes are then used to test it, so the subgroup hazard ratios are not independent discoveries.

What would settle it

Apply the same threshold learned in half the cohort to the other half: if the extreme opposite hazard ratios (≈0.20 and ≈6.68) do not reproduce in the held-out half, the reported heterogeneity is an artifact of selection; a simpler signal would be the subgroup p-values drifting toward the overall null under cross-fitting.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • If correct, aggregate trial results can hide subgroups with opposite treatment effects, so neutral trials should routinely be examined for HTE before concluding no benefit.
  • HTE-derived strata could guide eligibility criteria for future trials, potentially enriching for responders and excluding those likely harmed.
  • The same pipeline can be applied to other completed trials or real-world data to generate hypotheses about which baseline characteristics define benefit or harm.
  • Observed subgroup HRs (0.203 and 6.680) imply effect sizes much larger than typical average trial effects, suggesting high clinical relevance if replicated.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The threshold learned on the full data and then used to split the same patients for validation means the reported subgroup p-values are not independent; they are likely inflated by selection on outcomes.
  • A stronger test would be pre-specifying the cutoff on a training set and confirming on a held-out cohort, or using cross-fitting and sample splitting to avoid circularity.
  • The harmful subgroup's HR 6.68 with a wide CI rests on only 102 dapagliflozin patients; small absolute counts make the estimate unstable and sensitive to a few events.
  • If replicated in external EHR or prospective data, the approach could be used to redesign trials by excluding potential-harm patients; if the selection artifact dominates, the apparent heterogeneity will shrink or vanish upon validation.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

5 major / 4 minor

Summary. The paper emulates the DAPA-HF trial using Mayo Clinic EHR data, estimates heterogeneous treatment effects with a Meta-S learner, and uses a decision-tree threshold on a transformed outcome to split patients into 'beneficial' and 'harmful' subgroups. The authors report that, although the overall emulated cohort shows no significant dapagliflozin effect (HR 1.681, p=0.1507), the HTE-defined beneficial subgroup has HR 0.203 (p=0.0002) and the harmful subgroup HR 6.680 (p<0.0001). They also propose an eligibility-criteria optimization pipeline that yields directionally opposite adjusted HRs (0.175 and 2.125). The central claim is that HTE-guided stratification recovers clinically meaningful heterogeneity hidden in aggregate trial emulation.

Significance. If the results were valid, the approach would be a useful contribution to precision trial design and real-world evidence analysis. The paper also has a sensible motivation: moving from average effects to interpretable subgroups. However, the significance is severely limited by the fact that the headline subgroup effects are not independent tests: the same data are used to learn the HTE threshold and to estimate the within-subgroup Cox HRs. The extreme treatment imbalance (149 versus 8,865) and the absence of propensity-matching details further undermine the reliability of the point estimates. No code, no hold-out validation, and no external replication are provided, so the paper does not currently support its claim of identifying clinically actionable subgroups.

major comments (5)
  1. [§4.4 and §2.2.2/2.2.4] The decision-tree threshold (-0.03501) is learned by training on the transformed outcome Y* = Y × (T/p − (1−T)/(1−p)), which directly contains the mortality outcome Y and treatment assignment T. The same patients' outcomes are then used in §2.2.4 to fit Cox models within the resulting subgroups. The reported HRs (0.203 and 6.680) and p-values are therefore conditional on a data-chosen split and are not valid unconditional tests. A hold-out validation or sample-splitting procedure is required before any inferential claim can be made.
  2. [Table 1 and §4.2] The treatment groups are extremely unbalanced: 149 dapagliflozin patients versus 8,865 placebo patients. The beneficial and harmful subgroups contain only 47 and 102 treated patients, respectively, so the subgroup HRs are driven by a very small number of events. The manuscript does not report the propensity-score model specification, matching ratio, caliper, or balance diagnostics. These omissions are load-bearing because the survival estimates depend entirely on the quality of confounding adjustment in this highly imbalanced setting.
  3. [§1 vs §4.1] The paper contains an internal contradiction about the DAPA-HF trial's all-cause mortality result. §1 states that dapagliflozin was associated with a significant reduction in all-cause mortality (HR 0.83, p=0.0217), whereas §4.1 states that the original trial 'found no statistically significant difference between the two treatments.' This inconsistency affects the interpretation of the emulation benchmark and needs to be resolved.
  4. [§2.2.5 and §4.5] The eligibility-criteria optimization introduces a second, additional layer of selection. The optimized beneficial and harmful cohorts are chosen using HTE-derived labels as reference targets and are then re-evaluated through the full emulation workflow. The resulting HRs (0.175 and 2.125) are selected estimates, subject to the same circularity as the threshold-based subgroups, with no correction for multiple comparisons or independent validation. The claim that these are 'clinically actionable' criteria is not supported by the analysis.
  5. [Discussion (§3)] The paper acknowledges that external validation is necessary, but it does not temper the abstract and results sections, which present the directionally distinct subgroup HRs as established findings. Given the selection-circularity, the extreme imbalance, the internal contradiction about DAPA-HF, and the lack of any hold-out or external validation, the central claim is not supported as stated.
minor comments (4)
  1. [Throughout] Several grammatical issues: 'one independent modelling strategies' in §4.3, 'Meta-S' capitalization inconsistency, and the transformed-outcome formula in §4.4 appears garbled (should read Y* = Y × (T/p − (1−T)/(1−p))).
  2. [§2.2.3] The text mentions 'two distinct HTE estimation strategies,' but only the Meta-S learner is described in the methods. Clarify whether a second learner was used or remove the reference.
  3. [Figure 1] The figure caption says panels (b) and (c) present stratified curves, but the labels in the figure are not described precisely enough to map them to 'beneficial' and 'harmful.'
  4. [Table 1] Several rows report '9,014 (100.00%)' for variables like 'Elevated NT-proBNP,' which is likely an artifact of how the EHR eligibility criteria were operationalized. This should be explained, as it seems implausible that all patients had elevated NT-proBNP.

Circularity Check

2 steps flagged

Subgroup HRs are fit on the same outcome data used to choose the HTE threshold, so the headline effects are selected, not predicted.

specific steps
  1. fitted input called prediction [§2.2.1–2.2.2, §2.2.4, §4.4]
    "we applied a decision tree–based thresholding approach, in which a shallow regression tree was trained on the transformed outcome, incorporating both mortality events and treatment assignment, to learn a data-driven HTE cutoff. ... The final threshold is -0.03501 (HTE value). ... In contrast, the beneficial subgroup ... (HR = 0.203, 95% CI, 0.087–0.476, p = 0.0002), whereas the harmful subgroup ... (HR = 6.680, 95% CI, 2.759–16.171, p < 0.0001)."

    The decision-tree cutoff is fit to a pseudo-outcome that explicitly contains mortality Y and treatment assignment T, so the split at -0.03501 is a function of the same full-cohort outcomes that are later used to compute the subgroup Cox HRs. The 'beneficial' and 'harmful' labels are therefore not independent predictions: they are partitions selected by an outcome-weighted transform, and the reported HRs and p-values are conditional on that data-chosen split. No held-out or external sample is used before the headline p-values are presented.

  2. fitted input called prediction [§4.5; results in §2.2.5]
    "Qualified candidates were then evaluated using the same downstream analysis framework, including propensity-score matching and adjusted Cox proportional-hazards modeling. Final beneficial candidates were required to have HR < 1 and (p < 0.05), whereas final harmful candidates were required to have HR > 1 and (p < 0.05)."

    The optimized eligibility criteria are selected precisely because they produce a significant Cox HR on the same mortality endpoint used to build the HTE threshold. Reporting the selected candidates' HRs (0.175 and 2.125) as validation re-uses the outcome to choose the cohort and then to estimate the effect in that cohort. This is outcome-based selection, not an independent confirmation of the optimized subgroups.

full rationale

Most of the pipeline is standard and the comparison to the DAPA-HF benchmark is a useful external reference, but the two headline subgroup findings are not independent of their fitting inputs. In §4.4 the HTE cutoff is learned from a pseudo-outcome that explicitly embeds the mortality outcome and treatment indicator; the same data are then used in §2.2.4 to compute Cox HRs inside the resulting splits. This makes the reported HRs (0.203 and 6.680) selected statistics: the split has already been optimized against a function of Y and T. The eligibility-criteria optimization in §4.5 is doubly selective, since candidate cohorts are retained only if their Cox HR is significant on the same endpoint. I am not treating the self-citation to [20] as load-bearing for the subgroup claim, and the inconsistent description of DAPA-HF's own mortality result is a correctness/evidence concern rather than circularity. The paper's own limitation statement—'Prospective evaluation and external validation across other data sources are necessary to determine whether these HTE-derived partitions truly define groups'—confirms that the in-sample results are not yet independently validated. The score is 7 rather than 8–10 because the Cox HR magnitudes are not algebraically identical to the fitted threshold; a properly designed out-of-sample or external validation could, in principle, support the claim. As written, however, the central claim reduces to in-sample outcome-based selection.

Axiom & Free-Parameter Ledger

4 free parameters · 5 axioms · 0 invented entities

The central claim rests on a fitted HTE threshold, an unspecified HTE model, an unvalidated causal identification assumption, and a data-dependent rule-selection procedure. No new physical or theoretical entities are introduced; the fragile inputs are statistical and design choices rather than invented constructs.

free parameters (4)
  • HTE threshold = -0.03501
    Learned by a shallow decision tree on the transformed outcome (§4.4) and used to split patients into beneficial and harmful subgroups; the entire subgroup result depends on this fitted cutoff.
  • Meta-S learner hyperparameters = not reported
    The gradient boosting or neural-network implementation, its hyperparameters, and training details are not specified, but the HTE scores driving the threshold depend on them.
  • Propensity score model specification = not reported
    The paper repeatedly invokes propensity-score matching but never reports the model, covariates, matching ratio, or balance diagnostics; these fitted values underlie all Cox estimates.
  • Eligibility rule filtering thresholds = not reported
    Candidate optimized cohorts were ranked by precision, recall, Jaccard index, cohort size, balance, and event availability, but no numeric cutoffs are given; final rules were selected by requiring HR<1 or HR>1 with p<0.05, a data-dependent criterion.
axioms (5)
  • domain assumption Electronic health records can emulate a randomized trial with valid causal interpretation.
    Section 4.1 states the study emulates DAPA-HF using MCC data and defers causal-validity details to reference [20], without presenting them here.
  • domain assumption Unconfoundedness after propensity-score matching: treatment assignment is independent of potential outcomes given measured covariates.
    Section 4.2/4.5 assume the propensity-matched cohort supports causal HR estimates, but no matching diagnostics are shown.
  • standard math The transformed outcome Y* = Y × (T/p_treat - (1-T)/(1-p_treat)) is a valid pseudo-outcome for treatment-effect estimation.
    Athey & Imbens [22,23] provide the IPW-style construction, but the formula in §4.4 is garbled and the derivation is deferred to an absent Supplementary Note 1.
  • ad hoc to paper Subgroup assignment based on HTE estimates learned from the same data can be used to validate treatment effects within those subgroups.
    The paper's central validation step (§2.2.4) estimates Cox HRs in subgroups defined by a threshold fitted on the same cohort, which assumes the threshold is independent of the outcome—an assumption not justified.
  • domain assumption All-cause mortality is captured in the EHR without informative censoring or outcome misclassification.
    Section 4.2 defines all-cause mortality as the outcome but does not discuss data completeness, loss to follow-up, or censoring mechanisms.

pith-pipeline@v1.3.0-alltime-deepseek · 12039 in / 10821 out tokens · 108525 ms · 2026-08-01T19:29:16.418999+00:00 · methodology

0 comments
read the original abstract

Traditional randomized trials often obscure clinically meaningful heterogeneity in treatment response by focusing on average effects. Leveraging real-world data to emulate clinical trials and estimate heterogeneous treatment effects (HTEs) offers a promising path toward more precise and efficient trial design. In this study, we emulate the DAPA-HF trial using electronic health records from the Mayo Clinic Cloud (MCC) to investigate whether HTE-guided stratification can identify patient subgroups with distinct treatment responses to dapagliflozin versus placebo in patients with heart failure with reduced ejection fraction. All-cause mortality was evaluated using Cox proportional hazards models, with HTEs estimated using a Meta-S learner and subgroups defined using a decision tree-based thresholding approach. In the overall cohort of the emulation, no significant treatment difference was observed (HR, 1.681; 95% CI, 0.828-3.413; p = 0.1507). However, compared with the overall emulated cohort, in which dapagliflozin showed no statistically significant survival benefit, HTE-driven stratification identified subgroups with significant and directionally distinct treatment effects. The beneficial (low-HTE) subgroup showed a significant survival benefit from dapagliflozin (HR = 0.203, 95% CI, 0.087-0.476, p = 0.0002), whereas the harmful (high-HTE) subgroup showed a significant harmful association with markedly increased mortality risk (HR = 6.680, 95% CI, 2.759-16.171, p < 0.0001). These findings indicate that HTE-guided stratification can uncover clinically meaningful beneficial and harmful treatment-effect patterns that are masked in the full-cohort emulation.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

21 extracted references

  1. [1]

    Assessing the gold standard—lessons from the history of RCTs

    Bothwell LE, Greene JA, Podolsky SH, Jones DS. Assessing the gold standard—lessons from the history of RCTs. N engl j med. 2016;374(22):2175-81

  2. [2]

    Why do we need some large, simple randomized trials? Statistics in medicine

    Yusuf S, Collins R, Peto R. Why do we need some large, simple randomized trials? Statistics in medicine. 1984;3(4):409-20

  3. [3]

    Challenges and lessons learned from COVID-19 trials: should we be doing clinical trials differently? Canadian Journal of Cardiology

    Janiaud P, Hemkens LG, Ioannidis JP. Challenges and lessons learned from COVID-19 trials: should we be doing clinical trials differently? Canadian Journal of Cardiology. 2021;37(9):1353-64

  4. [4]

    A literature review on the representativeness of randomized controlled trial samples and implications for the external validity of trial results

    Kennedy-Martin T, Curtis S, Faries D, Robinson S, Johnston J. A literature review on the representativeness of randomized controlled trial samples and implications for the external validity of trial results. Trials. 2015;16(1):495

  5. [5]

    to whom do the results of this trial apply?

    Rothwell PM. External validity of randomised controlled trials:“to whom do the results of this trial apply?”. The Lancet. 2005;365(9453):82-93

  6. [6]

    Evaluating eligibility criteria of oncology trials using real-world data and AI

    Liu R, Rizzo S, Whipple S, Pal N, Pineda AL, Lu M, et al. Evaluating eligibility criteria of oncology trials using real-world data and AI. Nature. 2021;592(7855):629-33

  7. [7]

    Broadening eligibility criteria and diversity among patients for cancer clinical trials

    Kaur M, Frahm F, Lu Y, Ascha MS, Guadamuz JS, Dotan E, et al. Broadening eligibility criteria and diversity among patients for cancer clinical trials. NEJM evidence. 2024;3(4):EVIDoa2300236

  8. [8]

    GIST 2.0: A scalable multi-trait metric for quantifying population representativeness of individual clinical studies

    Sen A, Chakrabarti S, Goldstein A, Wang S, Ryan PB, Weng C. GIST 2.0: A scalable multi-trait metric for quantifying population representativeness of individual clinical studies. Journal of biomedical informatics. 2016;63:325-36

  9. [9]

    Using big data to emulate a target trial when a randomized trial is not available

    Hernán MA, Robins JM. Using big data to emulate a target trial when a randomized trial is not available. American journal of epidemiology. 2016;183(8):758-64

  10. [10]

    Emulating randomized clinical trials with nonrandomized real-world evidence studies: first results from the RCT DUPLICATE initiative

    Franklin JM, Patorno E, Desai RJ, Glynn RJ, Martin D, Quinto K, et al. Emulating randomized clinical trials with nonrandomized real-world evidence studies: first results from the RCT DUPLICATE initiative. Circulation. 2021;143(10):1002-13

  11. [11]

    Real-world evidence and real-world data for evaluating drug safety and effectiveness

    Corrigan-Curay J, Sacks L, Woodcock J. Real-world evidence and real-world data for evaluating drug safety and effectiveness. Jama. 2018;320(9):867-8

  12. [12]

    Emulation of randomized clinical trials with nonrandomized database analyses: results of 32 clinical trials

    Wang SV, Schneeweiss S, Franklin JM, Desai RJ, Feldman W, Garry EM, et al. Emulation of randomized clinical trials with nonrandomized database analyses: results of 32 clinical trials. Jama. 2023;329(16):1376-85

  13. [13]

    Emulate randomized clinical trials using heterogeneous treatment effect estimation for personalized treatments: Methodology review and benchmark

    Ling Y, Upadhyaya P, Chen L, Jiang X, Kim Y. Emulate randomized clinical trials using heterogeneous treatment effect estimation for personalized treatments: Methodology review and benchmark. Journal of biomedical informatics. 2023;137:104256

  14. [14]

    Simulating Colorectal Cancer Trials Using Real-World Data

    Chen Z, Zhang H, George TJ, Guo Y, Prosperi M, Guo J, et al. Simulating Colorectal Cancer Trials Using Real-World Data. JCO Clinical Cancer Informatics. 2022;6:e2100195

  15. [15]

    Electronic medical records can be used to emulate target trials of sustained treatment strategies

    Danaei G, Rodríguez LAG, Cantero OF, Logan RW, Hernán MA. Electronic medical records can be used to emulate target trials of sustained treatment strategies. Journal of clinical epidemiology. 2018;96:12-22

  16. [16]

    Estimation and inference of heterogeneous treatment effects using random forests

    Wager S, Athey S. Estimation and inference of heterogeneous treatment effects using random forests. Journal of the American Statistical Association. 2018;113(523):1228-42

  17. [17]

    Metalearners for estimating heterogeneous treatment effects using machine learning

    Künzel SR, Sekhon JS, Bickel PJ, Yu B. Metalearners for estimating heterogeneous treatment effects using machine learning. Proceedings of the national academy of sciences. 2019;116(10):4156-65

  18. [18]

    Learning representations for counterfactual inference

    Johansson F, Shalit U, Sontag D, editors. Learning representations for counterfactual inference. International conference on machine learning; 2016: PMLR

  19. [19]

    Dapagliflozin in patients with heart failure and reduced ejection fraction

    McMurray JJ, Solomon SD, Inzucchi SE, Køber L, Kosiborod MN, Martinez FA, et al. Dapagliflozin in patients with heart failure and reduced ejection fraction. New England Journal of Medicine. 2019;381(21):1995-2008

  20. [21]

    Rapid and intensive guideline-directed medical therapy for heart failure: 5 core principles

    Greene SJ, Butler J, Fonarow GC. Rapid and intensive guideline-directed medical therapy for heart failure: 5 core principles. Circulation. 2024;150(6):422-4

  21. [22]

    Machine learning methods for estimating heterogeneous causal effects

    Athey S, Imbens GW. Machine learning methods for estimating heterogeneous causal effects. stat. 2015;1050(5):1-26. 21