REVIEW 2 major objections 5 minor 35 references
A prespecified Bayesian rule can borrow phase II survival data and still pick when the next phase III interim look should happen, while keeping type I error under control.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
B²-FIC calibrates Bayesian phase-II borrowing for type I error, then uses IA1 predictive probability to schedule the earliest admissible IA2, yielding earlier decisions than fixed GSD when evidence is favorable while holding empirical FWER near 0.025.
T0 review reviewed 2026-07-11 challenge →
load-bearing objection Solid, regulatorily-aware design paper that correctly couples type-I-calibrated dynamic borrowing with predictive selection of the next interim time; simulation evidence is extensive under a Weibull DGM. the 2 major comments →
A Bayesian predictive framework for adaptive interim-analysis timing with robust borrowing in confirmatory trials
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
The authors claim that type I error-calibrated Bayesian borrowing of continuing phase II treatment-arm survival data can be coupled with a Bayesian predictive-probability rule that selects the earliest acceptable timing for the second interim analysis, and that the resulting B^{2}-FIC design maintains empirical type I error while improving interim power and shortening calendar time to decision when evidence is favorable.
What carries the argument
B^{2}-FIC: a design-stage-calibrated pair of robust borrowing priors (spike-slab commensurate and elastic) whose posteriors feed a predictive-probability timing rule that chooses the earliest candidate IA2 information fraction whose predicted probability of crossing the O'Brien–Fleming efficacy boundary meets a fixed threshold.
Load-bearing premise
The whole calibration and predictive schedule rest on a shared-shape Weibull proportional-hazards model with known common shape and purely administrative censoring; if true hazards are non-proportional or shapes differ across phases, both the type-I anchors and the predictive probabilities can mislead.
What would settle it
Re-run the simulation battery under delayed-effect or non-proportional hazards (different Weibull shapes or cure-model mixtures) and check whether overall one-sided type I error still stays ≤0.025 at the discrepancy boundaries used for calibration; if it exceeds the bound while power or timing gains disappear, the central operating-characteristic claim fails.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes B²-FIC, a design-stage framework for confirmatory time-to-event trials that couples type-I-error-calibrated Bayesian borrowing of continuing phase-II treatment-arm data with Bayesian predictive probability used to select the earliest admissible IA2 information fraction after IA1. Candidate borrowing procedures (commensurate and elastic priors) are calibrated at boundary discrepancies δ=±δ* so that overall one-sided FWER stays at or below the nominal level, after which power is maximized within the feasible set. At IA1 the calibrated posterior yields predictive probabilities over a candidate set of IA2 fractions; the earliest fraction meeting a prespecified threshold is recommended, and recommendations are summarized in a monotone design-stage decision table. Simulations under a shared-shape Weibull PH model report empirical type-I control near the non-borrowing GSD benchmark for the calibrated procedures, interim-power gains relative to GSD, and earlier IA2 recommendations when phase-II and phase-III IA1 evidence are favorable. Two oncology case studies illustrate the decision tables.
Significance. If the simulation results hold under the stated data-generating assumptions, the paper supplies a concrete, prospectively specified answer to a design problem that is becoming practically relevant for first-in-class confirmatory trials: how to let borrowing-adjusted evidence influence the timing of a future interim look without abandoning type-I control. Strengths include the hard design-stage type-I constraint, the explicit separation of efficacy boundaries from the timing rule, the borrowing-adjusted information-fraction approximation for commensurate-prior predictive probability (with reported validation against nested simulation), and the monotone rule-50 decision table obtained by Dykstra projection. These features make the proposal more operationally usable than pure design-stage timing optimization or fixed-schedule dynamic-borrowing GSDs. The contribution is simulation-based rather than analytic, so its significance is that of a carefully calibrated design framework rather than a general theorem.
major comments (2)
- [§2.2, §3.1, Tables 2–3] §2.2 and §3.1 fix a shared Weibull shape k=0.8, proportional hazards, and purely administrative censoring for both calibration and all operating-characteristic claims. Tables 2–3 and Figures 2–5 therefore establish empirical type-I control and interim-power gains only under this DGM. The Discussion notes non-PH and shape-mismatch extensions as future work, but the central claim in the Abstract and §3 is currently scoped only to the simulated model. A modest sensitivity grid (e.g., exponential, increasing hazard, or mild delayed effect) would substantially strengthen the load-bearing claim that the calibrated procedures remain near the GSD FWER benchmark when the DGM is misspecified.
- [§2.4.2, Appendix S5, Figure 5] §2.4.2 and Appendix S5 replace nested MCMC predictive probability for B²-CP by a normal approximation that uses a borrowing-adjusted effective information fraction I_1,eff. The authors report mean absolute differences of 0.070 at r=0.5 down to 0.005 at r=0.8 against the direct estimator, and use the approximation for design-stage search. Because the earliest candidates are precisely those that drive the reported time savings (Figure 5, Tables 4–5), the manuscript should either (i) recompute the final decision tables with direct nested simulation for the selected rules, or (ii) quantify how often the approximation changes the recommended action relative to the direct estimator. Without that check, the adaptive-timing claim for B²-CP rests partly on an approximation whose largest error coincides with the most consequential candidates.
minor comments (5)
- [§3.1, §3.3.1] Several cross-references are incomplete or inconsistent (e.g., “Section X” in §3.1 Methods; Tables S15/S16 referenced in the main text while the printed tables are numbered 4–5 and S13–S16). A single consistent numbering pass is needed.
- [§3.1] The control-arm prior shape α_C=500 is described as “substantial event-scale prior information” but is not varied. A short sensitivity note (or a single additional column in Table 2) would clarify that treatment-arm operating characteristics are not driven by an unrealistically informative control prior.
- [§2.4.1, Table 6] Notation for the predictive threshold switches between η_PP, η_BPP and γ in different sections and in Table 6. Unify the symbol and state its numerical value once in the main text.
- [Data Availability Statement] Code and simulation seeds are not stated to be publicly available. Given the computational intensity of the nested predictive calculations, a repository or archive link would aid reproducibility.
- [Figure 2] Figure 2 panel labels and the conjugate-prior FWER values that exceed the plotted range would be clearer with an explicit note that some conjugate points are truncated or annotated.
Circularity Check
No significant circularity: operating characteristics and decision tables are Monte Carlo outputs under explicitly stated scenarios, not tautologies of the inputs.
full rationale
The paper proposes a design framework (B^{2}-FIC) whose two Bayesian pieces—type-I-calibrated dynamic borrowing and Bayesian predictive probability for IA2 timing—are fully specified at the design stage. Calibration of the commensurate and elastic priors is performed by constrained Monte Carlo search over discrepancy boundaries δ=±δ* so that empirical overall one-sided FWER ≤ α*; power and ESS are then evaluated only inside that admissible class (Tables 2–3, Figs. 2–4). The adaptive-timing rule evaluates posterior predictive success probabilities over a prespecified candidate set R and records the earliest r that meets η_PP; the resulting decision table is an empirical summary (rule-50 of the integrated action distribution, projected onto a monotone cone by Dykstra). None of these steps is definitional: the type-I control is an empirical operating characteristic estimated from 10 000 null replicates, the power gains are relative to a non-borrowing GSD comparator under the same Weibull DGM, and the earlier IA2 recommendations are observed simulation outcomes, not algebraic identities. The B^{2}-CP predictive approximation is validated against nested simulation (mean difference 0.070→0.005), confirming it is an approximation rather than a tautology. Free design choices (δ*=0.2, spike–slab scales, quantile anchors, α_C=500) are inputs, not quantities derived from the target claim. No self-citation is load-bearing for uniqueness or for the central operating-characteristic claims. The derivation chain is therefore self-contained against the paper’s own simulation benchmarks; score 0 is appropriate.
Axiom & Free-Parameter Ledger
free parameters (6)
- δ* (boundary discrepancy for calibration) =
0.2
- spike-slab half-normal scales for B²-CP =
0.25 / 2.0
- elastic quantile anchors (q0,q1) and resulting (a_el,b_el)
- predictive-probability threshold η_PP / γ =
0.9 (case studies)
- control-arm prior shape α_C =
500
- Weibull shape k =
0.8
axioms (5)
- domain assumption Shared-shape Weibull proportional-hazards model with known common shape k across arms and phases
- domain assumption Censoring is purely administrative; no loss-to-follow-up or competing risks
- ad hoc to paper Empirical type-I control at the two boundary discrepancies δ=±δ* is an adequate surrogate for confirmatory FWER control
- domain assumption O’Brien–Fleming boundaries on the posterior-probability scale remain valid decision thresholds under borrowing
- domain assumption Phase-II treatment-arm cohort remains under follow-up and is exchangeable enough for dynamic borrowing after calibration
invented entities (3)
-
B²-FIC calibration rule (hard type-I constraint + power maximization within feasible set)
no independent evidence
-
Borrowing-adjusted effective information fraction I_1,eff for approximate BPP under commensurate prior
no independent evidence
-
Monotone rule-50 decision table obtained by Dykstra projection of empirical action CDFs
no independent evidence
Cite this review
Pith. "Pith review of A Bayesian predictive framework for adaptive interim-analysis timing with robust borrowing in confirmatory trials." pith.science (2026). https://pith.science/paper/YXCZQPZ5
@misc{pith2026260704205,
author = {Pith},
title = {Pith review of: A Bayesian predictive framework for adaptive interim-analysis timing with robust borrowing in confirmatory trials},
year = {2026},
howpublished = {\url{https://pith.science/paper/YXCZQPZ5}},
note = {Machine review of arXiv:2607.04205}
}
abstract
Confirmatory phase III trials require rigorous evidence, yet for first-in-class (FIC) therapies they must often be designed when same-mechanism evidence is scarce. This uncertainty motivates planned interim analyses and makes phase II data from the same therapy a relevant source of prior evidence. However, both borrowing and repeated interim analyses must be calibrated to control the overall type I error rate. Because borrowing changes the evidence available at interim analyses relative to a non-borrowing group sequential design (GSD), it also raises the question of whether interim analysis timing should be prospectively adapted to the borrowing-adjusted evidence base. We propose a prespecified adaptive interim-timing framework based on Bayesian information borrowing and Bayesian predictive probability $B^2$-FIC. The borrowing model is calibrated against phase II--phase III discrepancy scenarios to control overall type I error rate. At the first interim analysis (IA1), the calibrated model combines phase II information with accumulating phase III data to update the posterior. Bayesian predictive probabilities from this posterior select the earliest information fraction for the second interim analysis (IA2) that meets the efficacy criterion. In simulations, $B^2$-FIC maintained empirical type I error and improved interim power across different scenarios. Predictive probabilities derived from phase II and phase III IA1 data selected earlier IA2 than GSD when evidence was favorable. Two oncology case studies illustrate the framework. Overall, $B^2$-FIC provides a calibrated framework for adapting interim timing to borrowing-adjusted evidence, an emerging design problem in confirmatory trials.
Figures
Reference graph
Works this paper leans on
-
[1]
ICH E20 Guideline on Adaptive Designs for Clinical Trials: Step 2b
International Council for Harmonisation . ICH E20 Guideline on Adaptive Designs for Clinical Trials: Step 2b. 2025. Draft guideline, Step 2b, endorsed June 25, 2025. Accessed June 1, 2026.https://www.ema.europa.eu/en/documents/scientific- guideline/ich-e20-guideline-adaptive-designs-clinical-trials-step- 2b_en.pdf. 23
2025
-
[2]
Bauer P, Bretz F, Dragalin V, K ¨onig F, Wassmer G. Twenty-five years of confirmatory adaptive designs: opportunities and pitfalls.Statistics in Medicine.2016;35(3):325–347. doi:10.1002/sim.6472
-
[3]
Boumendil L, Chevret S, L ´evy V, Biard L. Two-stage randomized clinical trials with a right-censored endpoint: Comparison of frequentist and Bayesian adaptive designs. Statistics in Medicine.2024;43(18):3364–3382. doi:10.1002/sim.10130
-
[4]
Togo K, Iwasaki M. Optimal timing for interim analyses in clinical trials.Journal of Biopharmaceutic Statistics.2013;23(5):1067–1080. doi:10.1080/10543406.2013.813522
-
[5]
Feng B, Zee B. Robust time selection for interim analysis in the Bayesian phase 2 exploratory clinical trial.Journal of Biopharmaceutical Statistics.2024;34(3):413–423. doi:10.1080/10543406.2023.2208665
-
[6]
Optimal scheduling of interim analyses in group sequential trials
He Z, Cro S, Billot L. Optimal scheduling of interim analyses in group sequential trials. arXiv.2025;abs/2509.05537. doi:10.48550/arXiv.2509.05537
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.2509.05537 2025
-
[7]
Wu X, Xu Y, Carlin BP. Optimizing interim analysis timing for Bayesian adaptive commensurate designs.Statistics in Medicine.2020;39(4):424–437. doi:10.1002/sim.8414
-
[8]
Kotalik A, Vock DM, Hobbs BP, Koopmeiners JS. A group-sequential randomized trial design utilizing supplemental trial data.Statistics in Medicine.2022;41(4):698–718. doi:10.1002/sim.9249
-
[9]
Zhang W, Pan Z, Yuan Y. A Bayesian group sequential design for randomized biosimilar clinical trials with adaptive information borrowing from historical data.Journal of Biopharmaceutic Statistics.2022;32(3):359–372. doi:10.1080/10543406.2022.2080700
-
[10]
Chiaruttini MV, Lorenzoni G, Gregori D. Bayesian dynamic borrowing in group-sequential design for medical device studies.BMC Medical Research Methodology.2025;25(1):78. doi:10.1186/s12874-025-02520-6
-
[11]
Dmitrienko A, Wang MD. Bayesian predictive approach to interim monitoring in clinical trials.Statistics in Medicine.2006;25(13):2178–2195. doi:10.1002/sim.2204
-
[12]
Saville BR, Connor JT, Ayers GD, Alvarez J. The utility of Bayesian predictive probabilities for interim monitoring of clinical trials.Clinical Trials.2014;11(4):485–493. doi:10.1177/1740774514531352
-
[13]
Yin G, Chen N, Lee JJ. Phase II trial design with Bayesian adaptive randomization and predictive probability.Journal of the Royal Statistical Society Series C: Applied Statistics. 2012;61(2):219–235. doi:10.1111/j.1467-9876.2011.01006.x. 24
-
[14]
Chen L, Pan J, Wu Y,et al.Bayesian two-stage design for phase II oncology trials with binary endpoint.Statistics in Medicine.2022;41(12):2291–2301. doi:10.1002/sim.9355
-
[15]
Broglio KR, Connor JT, Berry SM. Not too big, not too small: a goldilocks approach to sample size selection.Journal of Biopharmaceutic Statistics.2014;24(3):685–705. doi:10.1080/10543406.2014.888569
-
[16]
Rufibach K, Jordan P, Abt M. Sequentially updating the likelihood of success of a Phase 3 pivotal time-to-event trial based on interim analyses or external information.Journal of Biopharmaceutic Statistics.2016;26(2):191–201. doi:10.1080/10543406.2014.972508
-
[17]
Aubel P, Antigny M, Fougeray R, Dubois F, Saint-Hilary G. A Bayesian approach for event predictions in clinical trials with time-to-event outcomes.Statistics in Medicine. 2021;40(28):6344–6359. doi:10.1002/sim.9186
-
[18]
Fu J, Zhao D, Skanji D, Liu H, Tang RS, Yuan Y. Bayesian Prediction of Event Times Using Mixture Model for Blinded Randomized Controlled Trials.Statistics in Medicine. 2025;44(28-30):e70310. doi:10.1002/sim.70310
-
[19]
Chen MH, Ibrahim JG, Shao QM. Power prior distributions for generalized linear models.Journal of Statistical Planning and Inference.2000;84(1):121–137. doi:10.1016/S0378-3758(99)00140-8
-
[20]
Summarizing historical information on controls in clinical trials.Clinical Trials.2010;7(1):5–18
Neuenschwander B, Capkun-Niggli G, Branson M, Spiegelhalter DJ. Summarizing historical information on controls in clinical trials.Clinical Trials.2010;7(1):5–18. doi:10.1177/1740774509356002
-
[21]
Jin H, Yin G. Unit information prior for adaptive information borrowing from multiple historical datasets.Statistics in Medicine.2021;40(25):5657–5672. doi:10.1002/sim.9146
-
[22]
Hobbs BP, Carlin BP, Mandrekar SJ, Sargent DJ. Hierarchical Commensurate and Power Prior Models for Adaptive Incorporation of Historical Information in Clinical Trials. Biometrics.2011;67(3):1047–1056. doi:10.1111/j.1541-0420.2011.01564.x
-
[23]
Jiang L, Nie L, Yuan Y. Elastic priors to dynamically borrow information from historical data in clinical trials.Biometrics.2023;79(1):49–60. doi:10.1111/biom.13551
-
[24]
Marion J, Lorenzi E, Allen-Savietta C, Berry S, Viele K. Predictive Probabilities Made Simple: A Fast and Accurate Method for Clinical Trial Decision-Making.Statistics in Medicine.2025;44(13–14):e70120. doi:10.1002/sim.70120
-
[25]
An Algorithm for Restricted Least-Squares Regression
Dykstra RL. An Algorithm for Restricted Least-Squares Regression. Journal of the American Statistical Association.1983;78(384):837–842. doi:10.1080/01621459.1983.10477029. 25
-
[26]
New York: John Wiley & Sons 1988
Robertson T, Wright FT, Dykstra RL.Order Restricted Statistical Inference. New York: John Wiley & Sons 1988
1988
-
[27]
Using simulation studies to evaluate statistical methods
Morris TP, White IR, Crowther MJ. Using simulation studies to evaluate statistical methods. Statistics in Medicine.2019;38(11):2074–2102. doi:10.1002/sim.8086
doi:10.1002/sim.8086 2019
-
[28]
Guidance Principle on the Application of Bayesian External Information Borrowing Methods in Drug Clinical Trials (Trial Version)
Center for Drug Evaluation, National Medical Products Administration . Guidance Principle on the Application of Bayesian External Information Borrowing Methods in Drug Clinical Trials (Trial Version). tech. rep.Center for Drug Evaluation, National Medical Products Administration 2026. In Chinese
2026
-
[29]
Food and Drug Administration
U.S. Food and Drug Administration . Use of Bayesian Methodology in Clinical Trials of Drug and Biological Products: Guidance for Industry. tech. rep.U.S. Food and Drug Administration 2026. Draft guidance
2026
-
[30]
Wiesenfarth M, Calderazzo S. Quantification of prior impact in terms of effective current sample size.Biometrics.2020;76(1):326–336. doi:10.1111/biom.13124
-
[31]
Kopp-Schneider A, Calderazzo S, Wiesenfarth M. Power gains by using external information in clinical trials are typically not possible when requiring strict type I error control.Biometrical Journal.2020;62(2):361–374. doi:10.1002/bimj.201800395
-
[32]
Finn RS, Martin M, Rugo HS,et al.Palbociclib and letrozole in advanced breast cancer.N Engl J Med.2016;375(20):1925–1936. doi:10.1056/NEJMoa1607303
-
[33]
doi:10.1016/S1470-2045(14)71159-3
Finn RS, Crown JP, Lang I,et al.The cyclin-dependent kinase 4/6 inhibitor palbociclib in combination with letrozole versus letrozole alone as first-line treatment of oestrogen receptor-positive, HER2-negative, advanced breast cancer (PALOMA-1/TRIO-18): a randomised phase 2 study.Lancet Oncol.2015;16(1):25–35. doi:10.1016/S1470-2045(14)71159-3
-
[34]
Li J, Qin S, Xu J,et al.Randomized, Double-Blind, Placebo-Controlled Phase III Trial of Apatinib in Patients With Chemotherapy-Refractory Advanced or Metastatic Adenocarcinoma of the Stomach or Gastroesophageal Junction.Journal of Clinical Oncology.2016;34(13):1448–1454. doi:10.1200/JCO.2015.63.5995
-
[35]
Journal of Clinical Oncology.2013;31(26):3219–3225
Li J, Qin S, Xu Je,et al.Apatinib for Chemotherapy-Refractory Advanced Metastatic Gastric Cancer: Results From a Randomized, Placebo-Controlled, Parallel-Arm, Phase II Trial. Journal of Clinical Oncology.2013;31(26):3219–3225. doi:10.1200/jco.2013.48.8585. 26 Table 1.Scenario-varying parameters in the simulation study. Notation Parameter Values used 𝑚 (𝐼 ...
This paper was first reviewed by grok-4.5 on July 11, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.