REVIEW 2 major objections 5 minor 48 references
Leveraging External Data for Testing Experimental Therapies with Biomarker Interactions in Randomized Clinical Trials
T0 review · 2 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read A permutation test can borrow external data in randomized trials while keeping false positives at the nominal level regardless of external-data quality.
desk verdict Solid methodological paper: exact type I error control with external data is real, the BEP optimality proof is sound, and the power gains are honestly conditional on ED informativeness—but the abstract oversells the power claim. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the posterior-predictive conditional likelihood $m(D) = \int q_\theta(Y|X,A)\pi(\theta|D_E)d\theta$, used as the permutation test statistic. The permutation distribution over treatment assignments is the device that converts arbitrary Bayesian borrowing from external data into exact frequentist error control, because under the null hypothesis the joint distribution of outcomes, covariates, and treatments is invariant to permutations of treatment labels.
What would settle it
Enumerate all treatment permutations for a small randomized trial (say $n=20$) under a null where external controls come from a shifted population with unmeasured confounding; if the randomized ED-PT rejects in more than $\alpha$ of the permutations, the finite-sample false-positive claim fails.
Extended reading notes
Core claim
The discovery is that permutation invariance of the null hypothesis makes external data essentially free for error control: any test statistic, including one built from a posterior that incorporates external data, yields an exact level-$\alpha$ permutation test. The test statistic is the posterior-predictive conditional likelihood $m(D) = \int q_\theta(Y|X,A)\pi(\theta|D_E)d\theta$, which averages the working-model likelihood over the posterior distribution informed by external data. Under the null, permuting treatment labels leaves the distribution of the trial data unchanged, so permutation p-values are exact in finite samples; a Lehmann–Stein argument then shows the resulting test maximizes Bayesian expected power. Simulations with binary and continuous outcomes, and retrospective glioblastoma in silico trials, confirm that the test controls type I error even when external and internal control populations differ, while naive procedures that merge external data show inflated false-positive rates.
Load-bearing premise
The usefulness of the method, though not its false-positive control, depends on the external data and the working model actually being informative about control-arm outcomes in the trial population; when they are not, the power gains shrink or reverse.
Editorial extensions
If this is right
- A trialist can incorporate historical controls or electronic health record data into a confirmatory subgroup test and still report exact finite-sample type I error control, without needing to argue that the external data are unbiased.
- The randomized ED-PT is optimal in Bayesian expected power among all level-alpha tests, so it provides a principled benchmark for borrowing strength while maintaining a frequentist error guarantee.
- The practical Monte Carlo implementation with a finite number of permutations converges to the exact test as the number of permutations grows, making the method computationally feasible.
- In the glioblastoma in silico trials, ED-PT improved power over internal-data-only tests in heterogeneous-effect scenarios, with gains equivalent to increasing trial size by 10 to 33 percent in some configurations.
- Procedures that naively merge external and internal data, such as Wald tests, likelihood ratio tests, matching, and inverse probability weighting, showed inflated type I error when external and internal control populations differed, highlighting that automatic error control is not achieved by simply pooling data.
Reading between the lines
- The same permutation wrapper should work with any external-data-informed statistic, including a machine-learning prognostic score, a hierarchical meta-analysis, or a causal-model prediction, because exact level control comes from the permutation step rather than from the Bayesian model; only power depends on the model.
- A natural extension, not pursued in the paper, is an adaptive combination of several external sources, with the posterior automatically down-weighting sources that disagree with the trial's internal control arm.
- The one-sided modifications suggest a general decision-rule template: replace the two-sided evidence score with the posterior probability of a positive effect or with expected regret, and the permutation wrapper still controls size; this could extend to go/no-go decisions and registration-trial eligibility.
- The glioblastoma analysis implies that external-data choice matters, so a testable diagnostic comparing internal and external control predictions could help trialists select external sources that actually improve power.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes an external-data-augmented permutation test (ED-PT) for randomized clinical trials with potentially heterogeneous treatment effects. The test statistic is a posterior predictive likelihood m(D) under a Bayesian working model, with the posterior informed by external data; the null hypothesis is the randomization hypothesis that the joint distribution of outcomes, covariates, and treatment labels is invariant under permutations of treatment labels. The authors prove (Proposition 1) that the randomized permutation test has exact level α and maximizes Bayesian expected power among all α-level tests, and they provide asymptotic power approximations (Propositions 2 and 3) for binary and normal outcomes. They also introduce one-sided modifications, study binary and normal working models, and apply the procedure to resampled glioblastoma trial data, comparing power and type I error with several alternative procedures. The formal guarantees of exact type I error control and BEP optimality appear sound and follow from standard permutation theory (Lehmann-Romano and Lehmann-Stein).
Significance. If the results are correct, this is a practically valuable contribution: it provides a principled way to borrow external information while retaining exact finite-sample frequentist type I error control, together with a decision-theoretic optimality property under the stated working model. The supplementary proofs are explicit, the asymptotic power functions are derived in closed or near-closed form, and the simulation and retrospective studies are extensive and include realistic mismatches between external and internal data. The main reservations concern the strength of the power claims and the lack of a formal treatment of the one-sided variants; these do not undermine the core type I error control result.
major comments (2)
- [Abstract; §2.3] The abstract frames the procedure as testing the null that the experimental therapy 'does not improve' outcomes in any subpopulation, which is a one-sided null, but the test in §2.2 and Proposition 1 concerns the two-sided null in Eq. (1). The one-sided modifications in §2.3 are supported only by simulation evidence (e.g., the modified S1 scenario and Figure SM3); no formal claim or proof is given that these versions control the type I error rate under the composite null of no positive effect, which includes negative effects. The permutation argument in SM1 relies on invariance of the full data distribution under the two-sided H0 and does not automatically extend to negative-effect configurations. Please either provide a formal treatment of the one-sided versions (for example, monotonicity conditions on the modified statistics under the one-sided null) or revise the abstract and Section 2.3 to present the two-sided test as the primary contribution and the one-sided versions as empirical extensions.
- [§3, Scenario 2; §4] In the retrospective analysis with DFCI external data, the paper reports that ED-PT is 9.6% less powerful than the ID-only Wald test Test-B in Scenario 2, while no pre-specified criterion or diagnostic is offered for deciding when a given external dataset will improve power. The abstract's statement that the permutation test 'leverages the available external data to increase power' is therefore stronger than what is demonstrated if the comparison is against the best ID-only test. Please qualify the power claim (for example, by specifying that the improvement is relative to the ED-free permutation Test-A, as the simulation results show) and add a practical discussion, suitable for inclusion in a trial protocol, of when borrowing is expected to be beneficial.
minor comments (5)
- [Throughout] The mathematical notation in the provided text is heavily corrupted (for example, α appears as 'Ă', sample sizes as 'Ĥ', and parameters as 'Ă'); the published version must be typeset with a correctly rendering math font.
- [SM1; SM2] There are typos in the supplementary material: 'nomial level' in SM1 and 'randomzied' in SM2 should be corrected to 'nominal' and 'randomized'.
- [Proposition 2] The statement 'pv ˜č ˆpv ˜č → 1' is confusing; it should be written as pv_φ / \hat{pv}_φ → 1 in probability, with the ratio explicitly defined.
- [Figure 3] The caption and legend contain an incomplete label 'ED-PT-' in the top-left panel; the test-statistic subscript should be completed (presumably 'ED-PT-m').
- [Algorithm 1] The Monte Carlo p-value uses (1+J) in the denominator; the authors should state explicitly that this is the standard conservative Monte Carlo p-value, since readers may expect 1/J.
Circularity Check
No significant circularity: the ED-PT's exact type I error control and Proposition 1 optimality are proven from external permutation theorems, not from fitted inputs or self-citations.
full rationale
The paper's central guarantees are self-contained. The type I error proof (SM1) applies Lehmann-Romano Theorem 15.2.1 to the randomization hypothesis H0 in (1): under H0 the RCT data are invariant to permutations of treatment labels, so any measurable statistic, including m(D) built from a misspecified or mismatched external-data posterior, yields an exact alpha-level test. Proposition 1 uses the Lehmann-Stein theorem to show that the permutation test based on the integrated likelihood m(D) maximizes the Bayesian expected power defined in (5) with respect to the stated working model and posterior; this is a mathematical optimality statement relative to the model's own decision-theoretic criterion, not an empirical prediction forced by the data. The external data enter only as fixed weights in the test statistic; no parameter is fitted to the RCT outcome and later renamed as a prediction. The only self-citations (e.g., Ventz et al. 2019, 2022; Rahman et al. 2023) justify the structural prior assumption in the GBM case study; this is a modeling choice for a retrospective illustration and is not load-bearing for the method's guarantees. The power-improvement claim's dependence on ED informativeness is transparently demonstrated by the paper's own results, including Scenario 2 with DFCI electronic health records where ED-PT is 9.6% less powerful than the ID-only Test-B; this is a limitation of the empirical power claim, not a circular step.
Assumptions & free parameters
free parameters (2)
- Prior variance for regression coefficients in the working model =
sigma2 = 10 (S1 continuous simulations), 100 (GBM application)
- GBM prior discrepancy coefficients for EOR and MGMT x KPS =
normal(0, 100) priors
assumptions (5)
- domain assumption The trial treatment assignments are independent of pre-treatment characteristics and are exchangeable under the null hypothesis (Section 2.1, Equation 1).
- standard math Lehmann-Romano Theorem 15.2.1 and Lehmann-Stein Theorem 2 provide exactness and optimality of permutation tests (SM1, SM2).
- domain assumption The external data are independent of the trial data and are used only through the fixed posterior distribution pi(theta|D_E) in the test statistic (Section 2.2).
- domain assumption For the power gains in the application, the working model M and the external data contain useful information about the control outcome distribution in the RCT population (Section 2.2 and Section 3).
- standard math For the asymptotic power results, local alternatives of the form delta = a/n1^(1/2) and beta0 = b/n1^(1/2), plus standard CLT and delta-method regularity conditions, are assumed (Assumptions A1 and A2, Section 2.4).
Cite this review
Pith. "Pith review of Leveraging External Data for Testing Experimental Therapies with Biomarker Interactions in Randomized Clinical Trials." pith.science (2026). https://pith.science/paper/42SAQYVF
@misc{pith2026250604128,
author = {Pith},
title = {Pith review of: Leveraging External Data for Testing Experimental Therapies with Biomarker Interactions in Randomized Clinical Trials},
year = {2026},
howpublished = {\url{https://pith.science/paper/42SAQYVF}},
note = {Machine review of arXiv:2506.04128}
}
abstract
In oncology the efficacy of novel therapeutics often differs across patient subgroups, and these variations are difficult to predict during the initial phases of the drug development process. The relation between the power of randomized clinical trials and heterogeneous treatment effects has been discussed by several authors. In particular, false negative results are likely to occur when the treatment effects concentrate in a subpopulation but the study design did not account for potential heterogeneous treatment effects. The use of external data from completed clinical studies and electronic health records has the potential to improve decision-making throughout the development of new therapeutics, from early-stage trials to registration. Here we discuss the use of external data to evaluate experimental treatments with potential heterogeneous treatment effects. We introduce a permutation procedure to test, at the completion of a randomized clinical trial, the null hypothesis that the experimental therapy does not improve the primary outcomes in any subpopulation. The permutation test leverages the available external data to increase power. Also, the procedure controls the false positive rate at the desired $\alpha$-level without restrictive assumptions on the external data, for example, in scenarios with unmeasured confounders, different pre-treatment patient profiles in the trial population compared to the external data, and other discrepancies between the trial and the external data. We illustrate that the permutation test is optimal according to an interpretable criteria and discuss examples based on asymptotic results and simulations, followed by a retrospective analysis of individual patient-level data from a collection of glioblastoma clinical trials.
Figures
Reference graph
Works this paper leans on
-
[1]
Arel-Bundock, V., N. Greifer, and A. Heiss (2024). How to interpret statistical models using marginaleffects for r and python. Journal of Statistical Software\/ 111 , 1--32
work page 2024
-
[2]
Berger, J. O. (2013). Statistical decision theory and Bayesian analysis . Springer
work page 2013
-
[3]
Berger, J. O., X. Wang, and L. Shen (2014). A bayesian approach to subgroup identification. J Biopharm Stat\/ 24\/ (1), 110--129
work page 2014
-
[4]
Bonetti, M. and R. D. Gelber (2004). Patterns of treatment effects in subsets of patients in clinical trials. Biostatistics\/ 5\/ (3), 465--481
work page 2004
-
[5]
Brown, B. W., J. Herson, E. N. Atkinson, and M. E. Rozell (1987). Projection from previous studies: a bayesian and frequentist compromise. Controlled clinical trials\/ 8\/ (1), 29--44
work page 1987
- [6]
-
[7]
Chinot, O. L., W. Wick, W. Mason, R. Henriksson, F. Saran, et al. (2014). Bevacizumab plus radiotherapy–temozolomide for newly diagnosed glioblastoma. NEJM\/ 370\/ (8), 709--722. PMID: 24552318
work page 2014
-
[8]
Chu, Y. and Y. Yuan (2018). Blast: Bayesian latent subgroup design for basket trials accounting for patient heterogeneity. JRSS-C: Applied Statistics\/ 67\/ (3), 723--740
work page 2018
Show all 48 references
-
[9]
Johansson, A
Dar, H., A. Johansson, A. Nordenskj \"o ld, A. Iftimi, C. Yau, et al. (2021). Assessment of 25-year survival of women with estrogen receptor--positive/erbb2-negative breast cancer treated with and without tamoxifen therapy. JAMA Network Open\/ 4\/ (6), e2114904--e2114904
2021
-
[10]
De Bruijn, N. G. (1981). Asymptotic methods in analysis , Volume 4. Courier Corporation
1981
-
[11]
Feller, and L
Ding, P., A. Feller, and L. Miratrix (2016). Randomization inference for treatment effect variation. Journal of the Royal Statistical Society Series B: Statistical Methodology\/ 78\/ (3), 655--671
2016
-
[12]
Feller, and L
Ding, P., A. Feller, and L. Miratrix (2019). Decomposing treatment effect variation. Journal of the American Statistical Association\/ 114\/ (525), 304--317
2019
-
[13]
Freidlin, B., L. M. McShane, and E. L. Korn (2010). Randomized clinical trials with biomarkers: design issues. JNCI\/ 102\/ (3), 152--160
2010
-
[14]
Freidlin, B., L. M. McShane, M.-Y. C. Polley, and E. L. Korn (2012). Randomized phase ii trial designs with biomarkers. JCO\/ 30\/ (26), 3304
2012
-
[15]
Kim, and V
Haslam, A., M. Kim, and V. Prasad (2021). Updated estimates of eligibility for and response to genome-targeted oncology drugs among us cancer patients, 2006-2020. Annals of Oncology\/ 32\/ (7), 926--932
2021
-
[16]
Ho, D., K. Imai, G. King, and E. A. Stuart (2011). Matchit: nonparametric preprocessing for parametric causal inference. Journal of statistical software\/ 42 , 1--28
2011
-
[17]
Kennedy, A. D., D. J. Torgerson, M. K. Campbell, and A. M. Grant (2017). Subversion of allocation concealment in a randomised controlled trial: a historical case study. Trials\/ 18\/ (1), 1--6
2017
-
[18]
Lauko, A., A. Lo, M. S. Ahluwalia, and J. D. Lathia (2022). Cancer cell heterogeneity & plasticity in glioblastoma and brain tumors. In Seminars in Cancer Biology , Volume 82, pp.\ 162--175. Elsevier
2022
-
[19]
Lehmann, E. and J. Romano (2005). Testing statistical hypotheses , Volume 3. Springer
2005
-
[20]
Li, F., K. L. Morgan, and A. M. Zaslavsky (2018). Balancing covariates via propensity score weighting. JASA\/ 113\/ (521), 390--400
2018
-
[21]
Liau, L. M., K. Ashkan, S. Brem, J. L. Campian, and J. E. a. o. Trusheim (2023). Association of autologous tumor lysate-loaded dendritic cell vaccination with extension of survival among patients with newly diagnosed and recurrent glioblastoma. JAMA Onc\/ 9\/ (1), 112--121
2023
-
[22]
Liu, F. (2018). Assessment of bayesian expected power via bayesian bootstrap. Stat Med\/ 37\/ (24), 3471--3485
2018
-
[23]
Morita, S. and P. M \"u ller (2017). Bayesian population finding with biomarkers in a randomized clinical trial. Biometrics\/ 73\/ (4), 1355--1365
2017
-
[24]
Murphy, S. A. (2003). Optimal dynamic treatment regimes. JRSS-B: Statistical Methodology\/ 65\/ (2), 331--355
2003
-
[25]
Nugent, B. M., R. Madabushi, B. Buch, V. Peiris, V. Crentsil, et al. (2021). Heterogeneity in treatment effects across diverse populations. Pharmaceutical Statistics\/ 20\/ (5), 929--938
2021
-
[26]
Ventz, J
Rahman, R., S. Ventz, J. McDunn, B. Louv, I. Reyes-Rivera, et al. (2021). Leveraging external data in the design and analysis of clinical trials in neuro-oncology. Lancet Onc\/ 22\/ (10), e456--e465
2021
-
[27]
Ventz, R
Rahman, R., S. Ventz, R. Redd, T. Cloughesy, B. M. Alexander, et al. (2023). Accessible Data Collections for Improved Decision Making in Neuro-Oncology Clinical Trials . CCR\/ 29\/ (12), 2194--2198
2023
-
[28]
Baiocchi, and S
Rigdon, J., M. Baiocchi, and S. Basu (2018). Preventing false discovery of heterogeneous treatment effect subgroups in randomized trials. Trials\/ 19\/ (1), 1--15
2018
-
[29]
Ventz, V
Russo, M., S. Ventz, V. Wang, and L. Trippa (2023). Inference in response-adaptive clinical trials when the enrolled population varies over time. Biometrics\/ 79\/ (1), 381--393
2023
-
[30]
Sherman, R. E., S. A. Anderson, G. J. Dal Pan, G. W. Gray, T. Gross, et al. (2016). Real-world evidence—what is it and what can it tell us. NEJM\/ 375\/ (23), 2293--2297
2016
-
[31]
Clark, S
Slevin, M., P. Clark, S. Joel, S. Malik, R. Osborne, et al. (1989). A randomized trial to evaluate the effect of schedule on the activity of etoposide in small-cell lung cancer. JCO\/ 7\/ (9), 1333--1340
1989
-
[32]
Stupp, R., W. P. Mason, M. J. Van Den Bent, M. Weller, B. Fisher, et al. (2005). Radiotherapy plus concomitant and adjuvant temozolomide for glioblastoma. NEJM\/ 352\/ (10), 987--996
2005
-
[33]
Tierney, L. and J. B. Kadane (1986). Accurate approximations for posterior moments and marginal densities. JASA\/ , 82--86
1986
-
[34]
van der Wal, W. M. and R. B. Geskus (2011). ipw: an r package for inverse probability weighting. Journal of Statistical Software\/ 43 , 1--23
2011
-
[35]
Comment, B
Ventz, S., L. Comment, B. Louv, R. Rahman, P. Y. Wen, et al. (2022). The use of external control data for predictions and futility interim analyses in clinical trials. Neuro Onc\/ 24\/ (2), 247--256
2022
-
[36]
Khozin, B
Ventz, S., S. Khozin, B. Louv, J. Sands, P. Y. Wen, R. Rahman, L. Comment, B. M. Alexander, and L. Trippa (2022). The design and evaluation of hybrid controlled trials that leverage external data and randomization. Nat Commun\/ 13\/ (1), 5783
2022
-
[37]
Ventz, S., A. Lai, T. F. Cloughesy, P. Y. Wen, L. Trippa, and B. M. Alexander (2019). Design and evaluation of an external control arm using prior clinical trials and real-world data. CCR\/ 25\/ (16), 4993--5001
2019
-
[38]
Wager, S. and S. Athey (2018). Estimation and inference of heterogeneous treatment effects using random forests. JASA\/ 113\/ (523), 1228--1242
2018
-
[39]
Zhang, and R
Wang, J., H. Zhang, and R. Tiwari (2023). A propensity-score integrated approach to bayesian dynamic power prior borrowing. Stat Biopharm Research\/ , 1--23
2023
-
[40]
Wang, R., D. A. Schoenfeld, B. Hoeppner, and A. E. Evins (2015). Detecting treatment-covariate interactions using permutation methods. Stat Med\/ 34\/ (12), 2035--2047
2015
-
[41]
White, H. (1980). A heteroskedasticity-consistent covariance matrix estimator and a direct test for heteroskedasticity. Econometrica\/ , 817--838
1980
-
[42]
Zhang, H
Xu, J., H. Zhang, H. Zhang, J. Bian, and F. Wang (2023). Machine learning enabled subgroup analysis with real-world data to inform clinical trial eligibility criteria design. Scientific Reports\/ 13\/ (1), 613
2023
-
[43]
Yang, S., F. Li, M. A. Starks, A. F. Hernandez, R. J. Mentz, et al. (2020). Sample size requirements for detecting treatment effect heterogeneity in cluster randomized trials. Stat Med\/ 39\/ (28), 4218--4237
2020
-
[44]
K\"oll, and N
Zeileis, A., S. K\"oll, and N. Graham (2020). Various versatile variances: An object-oriented implementation of clustered covariances in R . Journal of Statistical Software\/ 95\/ (1), 1--36
2020
-
[45]
Ziegler, A., A. Koch, K. Krockenberger, and A. Gro hennig (2012). Personalized medicine using dna biomarkers: a review. Human genetics\/ 131 , 1627--1638
2012
-
[46]
Ding, C. G. (1992). Algorithm as 275: computing the non-central 2 distribution function. JRSS-C: Applied Statistics\/ 41\/ (2), 478--482
1992
-
[47]
Feller, W. (1968). An introduction to probability theory and its applications. volume 1 (3rd edition). pp.\ 525P--525P
1968
-
[48]
Lehmann, E. L. and C. Stein (1949). On the theory of some non-parametric hypotheses. The Annals of Mathematical Statistics\/ 20\/ (1), 28--45
1949
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.