Pith. sign in

REVIEW 2 major objections 5 minor 4 references

Bayesian leveraging of historical control data for a clinical trial with time-to-event endpoint

T0 review · 2 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read A Bayesian meta-analytic prior can borrow historical control data for time-to-event trials, with a robust mixture guarding against prior-data conflict.

desk verdict Useful MAP extension to time-to-event endpoints, but the NSCLC design results rest on an algebraically wrong Kaplan-Meier extraction formula and need correction before the aggregate-data claim can be trusted. read the letter →

arxiv 1908.07265 v2 pith:2REFOTWL submitted 2019-08-20 stat.AP

classification stat.AP MSC 62F1562N0162P10
keywords historicalcontroldatameta-analytic-predictivepriorpiecewiseexponentialmodeltime-to-eventendpointrobustmixtureeffectivenumberofeventsBayesianclinicaltrialdesignsurvivalanalysis
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper claims that the meta-analytic-predictive (MAP) prior, a Bayesian method for borrowing information from historical trials, can be extended from simple one-dimensional parameters to time-to-event endpoints modeled with piecewise exponential hazards. The paper shows that historical control data, even when only Kaplan-Meier curves are available in publications, can be converted into per-interval event counts and exposure times and then combined across trials into a prior for the control hazard in a new trial. To protect against the new trial differing from the historical ones, the prior can include a robust mixture component that automatically down-weights the historical data when they conflict with the new data. If this works, trialists could run smaller, faster randomized trials by using historical controls for the control arm, while still controlling error rates.

What carries the argument

The load-bearing object is the interval-wise Poisson likelihood for piecewise exponential survival, in which the number of events $r_{jk}$ in trial $j$ and interval $k$ follows $\text{Poisson}(\lambda_{jk} E_{jk})$, with exposure times reconstructed from published Kaplan-Meier curves. Around this likelihood, the paper builds a hierarchical model on log-hazards with a first-order dynamic linear model (a random walk on the interval means) and half-normal priors on between-trial standard deviations. The robustness mechanism is a mixture of exchangeability and non-exchangeability per interval, with fixed prior weights that are updated to posterior weights depending on how similar the new data are to the historical data. Finally, the expected local-information-ratio effective sample size converts the MCMC-based MAP prior into an effective number of events for design purposes.

What would settle it

Construct a simulated cohort with known piecewise hazards and a known censoring pattern, draw its Kaplan-Meier curve, apply the paper's Appendix A extraction steps to reconstruct interval event counts and exposure times, and compare them with the true values; substantial bias under realistic late-interval deaths or non-uniform censoring would show that the extracted Poisson data are not sufficient statistics, and the operating characteristics in Section 3.2 would need re-evaluation.

Watch

Extended reading notes

Core claim

The central discovery is a hierarchical Poisson model for piecewise exponential survival data in which, for each time interval, the log-hazards of the historical trials and of the new trial are exchangeable normal draws around a common mean with an across-trial standard deviation. The MAP prior for the new trial's control hazards is the conditional distribution of those hazards given the historical event counts and exposure times; in the meta-analytic-combined version the historical and new data are analyzed jointly, which is equivalent to the two-step MAP approach. A robust extension replaces the exchangeability assumption with a mixture: each interval of the new trial is exchangeable with probability $w$ and comes from a weakly informative non-exchangeable prior with probability $1-w$, and the weights update dynamically in the posterior. The information content of the prior is quantified by the expected local-information-ratio effective number of events, summing the effective sample sizes across intervals. Applications to ovarian carcinoma and non-small-cell lung cancer illustrate both analysis of a new trial and design of a new trial with reduced control-arm enrollment.

Load-bearing premise

The load-bearing premise is that event counts and exposure times reconstructed from published Kaplan-Meier curves, assuming deaths at mid-interval and constant censoring within each interval, are accurate enough that the Poisson likelihood and the resulting MAP prior are not materially biased.

Editorial extensions

If this is right

  • In the lung cancer design example, borrowing historical control data allowed the trial to enroll 130 patients instead of 230, with the final analysis at 110 events rather than 133.
  • When the true control median matches the historical data, borrowing increases power substantially; in the example, power rose from about 85 percent under stratification to over 95 percent under robust borrowing.
  • If the true control median is misspecified, full exchangeability can inflate type-I error above 10 percent, whereas the 50-50 robust mixture keeps it much closer to the nominal 2.5 percent.
  • The effective number of events of the MAP prior gives a single design-time number, here 36 or 58 events in the two applications, that trialists can use to judge how much historical information is being contributed.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A testable extension would be to simulate known survival and censoring processes, generate Kaplan-Meier curves from them, and apply the paper's digitization and exposure-time reconstruction; if the reconstructed Poisson counts are biased when deaths cluster at interval boundaries or censoring is non-uniform, the MAP prior and the reported operating characteristics would shift.
  • The paper's EXNEX weights are fixed per interval, but the discussion hints at trial-specific weights; allowing the weights to depend on trial-level covariates could yield a meta-regression version that explains rather than merely discounts heterogeneity.
  • The type-I error inflation seen under full exchangeability suggests that, in a regulatory setting, the robust mixture weight should be pre-specified and its sensitivity reported, rather than chosen after seeing the new data.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper proposes a Bayesian meta-analytic-predictive (MAP) approach for leveraging historical control data in trials with time-to-event endpoints. The method extends the standard MAP prior for one-dimensional parameters to piecewise exponential (PWE) hazards, adding a robust exchangeability/non-exchangeability mixture to guard against prior-data conflict. The authors describe how to derive MAP priors from historical aggregate or individual data, define an effective number of events (ENE) based on expected local information ratios, and illustrate the method in two applications: a joint analysis of ovarian carcinoma trials and the design of a non-small-cell lung cancer (NSCLC) trial. The design application reports frequentist operating characteristics under several borrowing schemes.

Significance. If the methodology is correct, the paper addresses a practically important problem: using published Kaplan-Meier curves to borrow historical control information in time-to-event trials, which could reduce control-arm sizes while preserving frequentist properties. The hierarchical PWE model and the robust mixture extension are natural and clinically motivated, and the use of operating-characteristic simulations is appropriate. The ENE computation via expected local information ratios is a useful contribution. However, the current manuscript contains load-bearing errors in the Kaplan-Meier data-extraction formula and in the supplied WinBUGS code, so the reported NSCLC design results and the reproducibility of both applications are not yet trustworthy. These issues are correctable, but the affected analyses must be redone.

major comments (2)
  1. [Appendix A, Step 4] The printed equation log(S_hat(t_{i+1})) - log(S_hat(t_i)) = 1 - d_i/n_i is algebraically incorrect. The Kaplan-Meier recursion is S(t_{i+1}) = S(t_i)(1 - d_i/n_i), so the left-hand side equals log(1 - d_i/n_i), which is approximately -d_i/n_i for small hazards, not 1 - d_i/n_i. If this formula was used to digitize the seven NSCLC Kaplan-Meier curves in Section 3.2, the extracted event counts d_i are systematically wrong. Consequently the MAP prior (median 4.88 months, ENE 36) and all operating characteristics in Table 3 and Figure 4 rest on invalid summary statistics. The extraction formula must be corrected and the entire NSCLC design analysis re-run before the central design claims can be accepted.
  2. [Appendix B.1, WinBUGS code] The supplied WinBUGS code contains several errors that affect the reported MAC analyses. In the output loop, the lines 'log(hazard[h,s,t]) <- log.hazard.base[s,t] + inprod(Xout[h,1:Ncov],beta[1:Ncov,1])' and 'log.hazard[h,s,t] <- log(hazard.base[s,t])' appear in sequence, so the second assignment overwrites the first and the covariate/treatment effect is dropped from survival outputs. Also, the expression 'step(Nstudies-1.5)' is identically 1 for Nstudies > 1, so the random effect RE[s,t] is always included, including for the predicted new study, which contradicts the intended prediction structure. Finally, when MAP.prior=TRUE the function increments Nstudies but does not expand p.exch or the NEX mean/sd matrices, so the printed Fiocco MAP-prior call in B.3.1 indexes beyond the supplied matrix dimensions and cannot run as written. Correcting these code defects and re-running the analyses is essential for the manuscript's reproducibility.
minor comments (5)
  1. [Section 3.1, paragraph 5] The text states that the EXNEX analysis has a median survival of 2.59 months, but Table 2 reports 2.62 years; this appears to be a units typo.
  2. [Section 3.1, paragraph 5] The text reports the EX median survival as 2.01 years, while Table 2 gives 2.24 years; the numbers should be reconciled.
  3. [Equation (11)] The displayed formula for ESS_ELIR is incomplete: it ends with a dangling '=' and no closed-form expression, and the definition of the Fisher information term is split awkwardly.
  4. [General] All in-text citations appear as '?' placeholders, making it impossible to verify the reference list and prior literature claims; this should be fixed.
  5. [Figure 3] The figure labels 'Expos/Day' and 'Events' are ambiguous; the caption and text should clarify whether these are exposure per day or total exposure, and event counts per interval.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity found; the MAP prior is derived from historical data only, and new-trial inference is a standard posterior update.

full rationale

The paper's derivation chain is not circular. The MAP prior for the new trial's control log-hazards is defined as the posterior of the new-trial parameter given only the historical data (Eq. (1), Eq. (10)); the new-trial likelihood (Eq. (12)) enters only in the analysis or design stage, and no quantity derived from the new trial or from the targeted operating characteristic is used to construct the prior. The robust mixture weights and the non-exchangeable variances are explicit analyst-specified inputs (Section 2.5), and the reported effective numbers of events are summary measures of the prior under the ESS_ELIR definition, not predictions fitted to the new data. The citations to Neuenschwander et al. and Schmidli et al. provide prior methodology and the MAP/MAC equivalence; these are external results with stated assumptions that do not include the target piecewise-exponential extension, and the extension itself is fully specified in Sections 2.2-2.5 and in the WinBUGS code of Appendix B. The main caveat is Appendix A Step 4's event-count equation, log(S_hat(t_{i+1})) - log(S_hat(t_i)) = 1 - d_i/n_i, which is not the Kaplan-Meier recursion (the left side should be log(1 - d_i/n_i)); this is a validity and correctness risk for the NSCLC aggregate-data extraction and the resulting design operating characteristics, but it is not a circularity. The stated extraction assumptions (constant censoring rates, mid-interval events) are similar limitations. No step in the derivation is equivalent by construction to its own inputs, and no predicted quantity is a renamed fitted parameter.

Assumptions & free parameters 5 free parameters · 4 assumptions · 0 invented entities

The central claim rests on standard Bayesian hierarchical model choices plus domain assumptions about exchangeability and proportional hazards. The free parameters are analyst-chosen hyperparameters and modeling choices, not hidden fits to the new trial outcome. No new physical or statistical entities are postulated beyond the robust mixture component, which is a standard modeling device.

free parameters (5)
  • robust mixture weight w = 0.5 or 0.9 in the applications
    Chosen by analysts to control the degree of borrowing. Higher w means more borrowing and larger type-I error, as shown in the operating characteristics.
  • between-trial heterogeneity prior scale s_tau = 0.5
    Half-normal prior on sigma_k with scale 0.5, chosen to cover small to large heterogeneity (Section 2.2, equation 9).
  • NEX prior variance = 1 on the log-hazard scale
    Variance approximately worth one observation, used for the non-exchangeability component (Section 2.5).
  • unit-information prior mean for mu1 = -1.171 for Fiocco, -5.167 for lung cancer
    Set to the log of the overall estimated hazard from the historical data. This centers the prior for the first interval's mean log-hazard on the historical estimate.
  • interval partition = 12 intervals for Fiocco, 9 for lung cancer
    The piecewise exponential model requires pre-specified cut points; results depend on this choice.
assumptions (4)
  • standard math Poisson likelihood for piecewise exponential survival (equation 4)
    Standard result: with constant hazards per interval, the Poisson likelihood arises from the piecewise exponential model.
  • domain assumption Exchangeability of trial-specific log-hazards (equation 5)
    The MAP prior assumes historical and new trial hazards are exchangeable around the mean, a key modeling assumption that drives borrowing.
  • domain assumption Dynamic linear model for interval means (equations 7-8)
    The temporal structure of the mean log-hazards is modeled as a random walk with discounting, imposing smoothness across time intervals.
  • domain assumption Proportional hazards for treatment effect (equation 12)
    The test treatment hazard is assumed to be a constant multiple of the control hazard; standard in oncology trials.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Bayesian leveraging of historical control data for a clinical trial with time-to-event endpoint." pith.science (2026). https://pith.science/paper/2REFOTWL

@misc{pith2026190807265,
  author       = {Pith},
  title        = {Pith review of: Bayesian leveraging of historical control data for a clinical trial with time-to-event endpoint},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/2REFOTWL}},
  note         = {Machine review of arXiv:1908.07265}
}
read the original abstract

The recent 21st Century Cures Act propagates innovations to accelerate the discovery, development, and delivery of 21st century cures. It includes the broader application of Bayesian statistics and the use of evidence from clinical expertise. An example of the latter is the use of trial-external (or historical) data, which promises more efficient or ethical trial designs. We propose a Bayesian meta-analytic approach to leveraging historical data for time-to-event endpoints, which are common in oncology and cardiovascular diseases. The approach is based on a robust hierarchical model for piecewise exponential data. It allows for various degrees of between trial-heterogeneity and for leveraging individual as well as aggregate data. An ovarian carcinoma trial and a non-small-cell cancer trial illustrate methodological and practical aspects of leveraging historical data for the analysis and design of time-to-event trials.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

4 extracted references · 1 canonical work pages

  1. [1]

    @esa (Ref

    \@ifxundefined[1] #1\@undefined \@firstoftwo \@secondoftwo \@ifnum[1] #1 \@firstoftwo \@secondoftwo \@ifx[1] #1 \@firstoftwo \@secondoftwo [2] @ #1 \@temptokena #2 #1 @ \@temptokena \@ifclassloaded agu2001 natbib The agu2001 class already includes natbib coding, so you should not add it explicitly Type <Return> for now, but then later remove the command n...

  2. [2]

    \@lbibitem[] @bibitem@first@sw\@secondoftwo \@lbibitem[#1]#2 \@extra@b@citeb \@ifundefined br@#2\@extra@b@citeb \@namedef br@#2 \@nameuse br@#2\@extra@b@citeb \@ifundefined b@#2\@extra@b@citeb @num @parse #2 @tmp #1 NAT@b@open@#2 NAT@b@shut@#2 \@ifnum @merge>\@ne @bibitem@first@sw \@firstoftwo \@ifundefined NAT@b*@#2 \@firstoftwo @num @NAT@ctr \@secondoft...

  3. [3]

    QѴM 񼎊8 P. ɷ

    @open @close @open @close and [1] URL: #1 \@ifundefined chapter * \@mkboth \@ifxundefined @sectionbib * \@mkboth * \@mkboth\@gobbletwo \@ifclassloaded amsart * \@ifclassloaded amsbook * \@ifxundefined @heading @heading NAT@ctr thebibliography [1] @ \@biblabel @NAT@ctr \@bibsetup #1 @NAT@ctr @ @openbib .11em \@plus.33em \@minus.07em 4000 4000 `\.\@m @bibit...

  4. [4]

    write newline

    " write newline "" before.all 'output.state := FUNCTION blank.sep after.quote 'output.state := FUNCTION fin.entry doi empty output.state after.quoted.block = 'skip 'add.period if if write newline FUNCTION new.block output.state before.all = 'skip output.state after.quote = after.quoted.block 'output.state := after.block 'output.state := if if FUNCTION new...

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.