REVIEW 2 major objections 5 minor 37 references
Impact of interim analyses on Bayesian and frequentist operating characteristics of Bayesian clinical trials
T0 review · 2 major / 5 minor · reviewed 2026-07-11 · grok-4.5
Pith's one-line read Bayesian clinical trials are no less harmed by repeated interim analyses than frequentist ones; multiplicity still matters.
desk verdict Solid, fully reproducible simulation paper that settles a live practical confusion: Bayesian sequential trials are not multiplicity-immune once you look at either Type I or Bayesian discovery metrics under prior disagreement. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Positive true discovery odds (pTDO = Pr(μ>0|V)/Pr(μ≤0|V)), equal to the Bayes factor of an efficacy conclusion when the prior is symmetric about zero, evaluated by Monte Carlo under a reviewer’s prior that may differ from the investigator’s analysis prior, together with the analogous Bayesian risk A and power 1−B.
What would settle it
Repeat the simulation suite with non-conjugate logistic or survival models, unknown variance, and decisions based on expected utility or a minimal clinically important difference; if Type I error and pTDO no longer degrade with the number of looks, the central claim is false.
Extended reading notes
Core claim
Repeated interim analyses alter both frequentist and Bayesian operating characteristics of group-sequential trials, regardless of the framework used. Without multiplicity adjustment the Type I error rate rises with the number of analyses. Bayesian alternatives—risk of erroneous efficacy conclusions and the informative value of an efficacy conclusion measured by positive true discovery odds—are meaningful but degrade with more analyses and with prior divergence between investigator and reviewer; even under full prior agreement that informative value falls. Frequentist operating characteristics of Bayesian trials can be controlled by known methods, but Bayesian operating characteristics cannot
Load-bearing premise
Binary stop-or-continue decisions under a normal–normal conjugate model with known unit variance and equally spaced looks adequately represent how interim analyses affect real trial conclusions.
Editorial extensions
If this is right
- Multiplicity must be designed for in every trial with interim analyses, Bayesian or frequentist.
- Reviewers of positive Bayesian sequential trials should prefer pTDR or pTDO under their own prior rather than the authors’ reported posterior probability.
- Perpetual Bayesian designs that analyse continuously cannot keep erroneous-efficacy risk controlled indefinitely; the informative value of an efficacy conclusion approaches zero.
- The classical Type I error under the null remains a useful upper-bound metric for Bayesian designs with many looks because it is the most sceptical prior.
- Public, reproducible simulation code is required to assess operating characteristics of Bayesian adaptive designs.
Reading between the lines
- Platform and master protocols that drop or add arms under Bayesian rules will inherit the same degradation of decision value unless they hard-code multiplicity-adjusted thresholds calibrated across a range of reviewer priors.
- Even decision-theoretic designs that maximise expected utility rather than threshold a posterior probability can still suffer sequential-selection bias in the reported posterior once early stopping is allowed.
- Guidance that treats Bayesian designs as free of multiplicity penalties risks systematically overstating the reliability of early-stopping claims in adaptive trials.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript evaluates how repeated interim analyses affect both frequentist and Bayesian operating characteristics of group-sequential clinical trials under a normal–normal conjugate model. Through large Monte Carlo simulations (100 000 trials; m up to 500 equally spaced looks), the authors show that unadjusted efficacy stopping inflates Type I error for Bayesian designs much as for frequentist ones (especially under weakly informative analysis priors), that Bayesian analogues A and 1−B and the proposed positive true discovery odds (pTDO) degrade with more looks and with investigator–reviewer prior divergence, and that even under exact prior agreement pTDO falls as the number of analyses rises. Frequentist Type I error of Bayesian designs can be controlled by boundary recalibration (Pocock-style), whereas Bayesian OCs cannot be guaranteed across divergent reviewer priors. The central claim is that Bayesian trials are not immune to multiplicity, contrary to some statements in the literature.
Significance. The paper addresses a persistent and practically important confusion in adaptive and Bayesian trial design. Its contribution is clear: it separates the likelihood-principle invariance of the posterior from the operating characteristics of binary stopping rules, defines transparent Bayesian decision metrics (including pTDO), and demonstrates with reproducible code and extensive simulations that both frequentist and Bayesian OCs are altered by repeated looks. The public GitHub repository and conjugate closed forms strengthen credibility. If the qualitative conclusions hold beyond the normal model—as the authors argue in §6.6—the work should influence how Bayesian group-sequential and “perpetual” designs are planned, reported, and reviewed, and it usefully reframes Type I error as a consensus upper bound under the most sceptical prior.
major comments (2)
- [§5.2, Figure 4, §6.5] §5.2 and Figure 4: After calibrating e_i and n so that A = 0.05 and 1−B = 0.80 under π_investigator = π_reviewer = N(0, 0.25²), the authors correctly show that more sceptical reviewer priors inflate A and collapse pTDO. The practical implication is left somewhat open. A short discussion of whether designers should (a) control frequentist Type I as a consensus bound (§6.5(iv)), (b) pre-specify a range of reviewer priors and control worst-case A, or (c) report pTDO under a grid of σ_reviewer would make the recommendation actionable without changing the central claim.
- [§6.4, Supplementary Figures S74–S78] §6.4 and Supplementary Figures S74–S78: The observation that, with many looks and a fixed posterior-probability threshold, the posterior probability at stopping concentrates at the threshold itself is load-bearing for the claim that the posterior loses quantitative interpretability. Elevating a concise version of this result (and one main-text figure) would strengthen the argument that binary stopping, not conjugacy per se, drives the loss of informativeness, and would better support the advice that external reviewers should rely on pTDO rather than the reported posterior.
minor comments (5)
- [Table 1, §2.2] Table 1 and §2.2: Clarify that pTDO equals the Bayes factor of V for μ > 0 only when the reviewer prior is symmetric about zero (as assumed throughout the simulations); the general odds form is already stated but easy to miss.
- [Figure 1] Figure 1B and analogous power-vs-α plots: The left-to-right ordering of points by number of analyses is described in the caption but hard to read; a small annotation or colour gradient for m would help.
- [§2.4] §2.4 / Table 3: State explicitly that the unit-variance assumption is without loss of generality (by rescaling) so that readers do not misread the design as restricted to σ = 1 endpoints.
- [References] References: The FDA Bayesian guidance link is dated “accessed 14 June 2026”; confirm the access date and citation format for the journal.
- [Supplementary tables] Supplementary tables S1–S6 are dense; consider moving a single representative pTDO table (e.g., under prior agreement) into the main text to illustrate the decline with m without forcing readers into the appendix.
Circularity Check
No circularity: operating characteristics are Monte-Carlo estimates under independently generated data and priors, not algebraic restatements of the decision rules or fitted inputs.
full rationale
The paper's central claims rest on simulation studies that generate data under fixed nulls (μ=0), alternatives (μ=0.1112), or reviewer priors π_reviewer=N(0,σ_reviewer²), apply investigator decision rules (posterior probability thresholds or p-value boundaries), and estimate Type I error, power, A, 1-B, pTDR, pFDR and pTDO as empirical frequencies (Methods §§2.3–2.4; Tables 4–7; Figs. 1–4). These quantities are not defined in terms of one another, nor are parameters fitted to a subset of the same outcomes and then re-presented as predictions. pTDO is obtained from Bayes' theorem applied to the simulated joint distribution of (μ,V) and equals (1-B)/A under the paper's symmetric priors; it is not forced by construction of the stopping boundary. Boundary recalibration (Pocock-style or numerical optimisation of e_i and n) is performed to hit pre-specified targets and then re-evaluated under mismatched priors; the degradation is an observed simulation result, not an identity. Self-citations (e.g., the authors' platform-trial review) supply only methodological background and are not load-bearing premises. The normal–normal conjugate model is an explicit modelling choice, not a uniqueness theorem imported from prior work by the same authors. Consequently the derivation chain contains no self-definitional steps, no fitted-input-as-prediction, and no circular self-citation.
Assumptions & free parameters
free parameters (5)
- Maximum sample size n =
500 (unadjusted); re-derived for controlled scenarios
- Alternative effect μ1 =
0.1112
- Investigator prior SD σ_investigator =
grid in {0.001, 0.01, 0.05, 0.1, 0.25, 1, 10, 100}
- Reviewer prior SD σ_reviewer =
same grid as investigator
- Efficacy/futility thresholds e_i, f_i =
0.05 unadjusted; scenario-specific when controlled
assumptions (5)
- domain assumption Observations X_i are i.i.d. N(μ,1) with known variance (w.l.o.g. by rescaling).
- domain assumption Investigator prior is normal N(0, σ_investigator²), conjugate to the likelihood.
- domain assumption Trial decisions are binary efficacy/futility conclusions from comparing posterior Pr(μ>0|data) or one-sided p-values to fixed thresholds at pre-specified looks.
- ad hoc to paper Bayesian operating characteristics are defined by averaging over a reviewer prior π_reviewer that may differ from the analysis prior.
- standard math Standard group-sequential multiplicity methods (e.g., Pocock boundaries) control frequentist Type I error under the stated correlation structure of sequential z-statistics.
invented entities (1)
-
Positive true discovery odds (pTDO)
Cite this review
Pith. "Pith review of Impact of interim analyses on Bayesian and frequentist operating characteristics of Bayesian clinical trials." pith.science (2026). https://pith.science/paper/NSDQ4CR7
@misc{pith2026260704976,
author = {Pith},
title = {Pith review of: Impact of interim analyses on Bayesian and frequentist operating characteristics of Bayesian clinical trials},
year = {2026},
howpublished = {\url{https://pith.science/paper/NSDQ4CR7}},
note = {Machine review of arXiv:2607.04976}
}
read the original abstract
While Bayesian methods are increasingly used in clinical research, confusion persists as to whether Bayesian designs are affected by repeated interim analyses, and how such effects should be evaluated. We aimed to clarify this question by evaluating both frequentist and Bayesian operating characteristics of group-sequential trials. We conducted simulation studies with normally distributed outcomes, examining designs with repeated analyses, with and without futility stopping rules and multiplicity adjustments. We show that repeated interim analyses alter both frequentist and Bayesian operating characteristics in group-sequential trials, regardless of the inferential framework adopted. Without proper adjustment for multiplicity, the Type I error rate increases with the number of analyses. Bayesian operating characteristics such as the risk of erroneous conclusions and the informative value of an efficacy conclusion are meaningful alternatives to classical frequentist metrics, but they are sensitive to prior divergence between stakeholders, and this sensitivity grows with the number of analyses. Even under full prior agreement, the informative value of an efficacy conclusion is reduced. While it is possible to control frequentist operating characteristics of Bayesian trials with appropriate methods, it is not possible to guarantee such control for Bayesian operating characteristics because they are prior-dependent and different stakeholders may adopt different priors. Contrary to claims that Bayesian inference is immune to multiplicity, our results show that Bayesian clinical trials are no less affected by repeated analyses than frequentist ones, regardless of the framework used to evaluate them.
Figures
Reference graph
Works this paper leans on
-
[1]
𝟎𝟓 and 𝑷𝒓𝛍∼𝑵(𝟎,𝟎.𝟐𝟓𝟐)( 𝑽 ∣∣ 𝛍 > 𝟎 ) = 𝟎. 𝟖𝟎
-
[2]
non-informative
Discussion In this work, we show that repeated interim analyses of accumulating data in clinical trials, without adjustment for multiple testing, fundamentally alter their operating characteristics, regardless of whether a frequentist or Bayesian framework is adopted for the design of the interim analyses, the interpretation of the operating characteristi...
-
[3]
Conclusion Consistently with the current literature, we confirmed that group-sequential Bayesian trials are not immune to Type I error rate inflation. Moreover, we showed that Bayesian operating characteristics such as the risk of erroneous conclusions and the informative value of an efficacy conclusion degrade as the number of interim analyses increases ...
-
[4]
The testing of statistical hypotheses in relation to probabilities a priori
Neyman J, Pearson ES. The testing of statistical hypotheses in relation to probabilities a priori. Math Proc Camb Phil Soc 1933; 29: 492–510
1933
-
[5]
Controlling the False Discovery Rate: A Practical and Powerful Approach to Multiple Testing
Benjamini Y, Hochberg Y. Controlling the False Discovery Rate: A Practical and Powerful Approach to Multiple Testing. Journal of the Royal Statistical Society Series B (Methodological) 1995; 57: 289–300
1995
-
[6]
A Sharper Bonferroni Procedure for Multiple Tests of Significance
Hochberg Y. A Sharper Bonferroni Procedure for Multiple Tests of Significance. Biometrika 1988; 75: 800–802
1988
-
[7]
Multiple Comparison Procedures
Hochberg Y, Tamhane AC. Multiple Comparison Procedures. 1st edn. Wiley. Epub ahead of print 21 September 1987. DOI: 10.1002/9780470316672
-
[8]
A Simple Sequentially Rejective Multiple Test Procedure
Holm S. A Simple Sequentially Rejective Multiple Test Procedure. Scandinavian Journal of Statistics 1979; 6: 65–70
1979
Show all 37 references
-
[9]
The problem of multiple comparisons
Tukey JW. The problem of multiple comparisons. Princeton University
-
[10]
Multiple testing in clinical trials
Bauer P. Multiple testing in clinical trials. Statist Med 1991; 10: 871–890
1991
-
[11]
Repeated Significance Tests on Accumulating Data
Armitage P, McPherson CK, Rowe BC. Repeated Significance Tests on Accumulating Data. Journal of the Royal Statistical Society Series A (General) 1969; 132: 235
1969
-
[12]
A case for bayesianism in clinical trials
Berry DA. A case for bayesianism in clinical trials. Statistics in Medicine 1993; 12: 1377–1393
1993
-
[13]
Bayesian clinical trials
Berry DA. Bayesian clinical trials. Nat Rev Drug Discov 2006; 5: 27–36
2006
-
[14]
Bayesian adaptive trials offer advantages in comparative effectiveness trials: an example in status epilepticus
Connor JT, Elm JJ, Broglio KR. Bayesian adaptive trials offer advantages in comparative effectiveness trials: an example in status epilepticus. Journal of Clinical Epidemiology 2013; 66: S130–S137
2013
-
[15]
A practical guide to adopting Bayesian analyses in clinical research
Gunn-Sandell LB, Bedrick EJ, Hutchins JL, et al. A practical guide to adopting Bayesian analyses in clinical research. Journal of Clinical and Translational Science 2024; 8: e3
2024
-
[16]
Characteristics, Progression, and Output of Randomized Platform Trials: A Systematic Review
Griessbach A, Schönenberger CM, Taji Heravi A, et al. Characteristics, Progression, and Output of Randomized Platform Trials: A Systematic Review. JAMA Netw Open 2024; 7: e243109
2024
-
[17]
Characteristics, design, and statistical methods in platform trials: a systematic review
Massonnaud CR, Schönenberger CM, Chiaborelli M, et al. Characteristics, design, and statistical methods in platform trials: a systematic review. Journal of Clinical Epidemiology 2025; 184: 111827
2025
-
[18]
Application of Bayesian approaches in drug development: starting a virtuous cycle
Ruberg SJ, Beckers F, Hemmings R, et al. Application of Bayesian approaches in drug development: starting a virtuous cycle. Nat Rev Drug Discov 2023; 22: 235–250
2023
-
[19]
Bayesian analysis in confirmatory clinical trials: A narrative review and discussion of current practice
Turner RM, Tweed CD, Duong T, et al. Bayesian analysis in confirmatory clinical trials: A narrative review and discussion of current practice. Clinical Trials 2026; 17407745261437669
2026
-
[20]
Interim Analysis in Clinical Trials: The Role of the Likelihood Principle
Berry DA. Interim Analysis in Clinical Trials: The Role of the Likelihood Principle. The American Statistician 1987; 41: 117–122
1987
-
[21]
A Bayesian group sequential design for a multiple arm randomized clinical trial
Rosner GL, Berry DA. A Bayesian group sequential design for a multiple arm randomized clinical trial. Statistics in Medicine 1995; 14: 381–394
1995
-
[22]
Do we need to adjust for interim analyses in a Bayesian adaptive trial design? BMC Med Res Methodol 2020; 20: 150
Ryan EG, Brock K, Gates S, et al. Do we need to adjust for interim analyses in a Bayesian adaptive trial design? BMC Med Res Methodol 2020; 20: 150
2020
-
[23]
Comparison of Bayesian with group sequential methods for monitoring clinical trials
Freedman LS, Spiegelhalter DJ. Comparison of Bayesian with group sequential methods for monitoring clinical trials. Controlled Clinical Trials 1989; 10: 357–367
1989
-
[24]
Problems of multiplicity in clinical trials
Simon R. Problems of multiplicity in clinical trials. Journal of Statistical Planning and Inference 1994; 42: 209–221
1994
-
[25]
Comparison of Bayesian and frequentist group-sequential clinical trial designs
Stallard N, Todd S, Ryan EG, et al. Comparison of Bayesian and frequentist group-sequential clinical trial designs. BMC Med Res Methodol 2020; 20: 4
2020
-
[26]
Bayesian Analytical Methods in Cardiovascular Clinical Trials: Why, When, and How
Heuts S, Kawczynski MJ, Sayed A, et al. Bayesian Analytical Methods in Cardiovascular Clinical Trials: Why, When, and How. Canadian Journal of Cardiology 2025; 41: 30–44
2025
-
[27]
FDA. Use of Bayesian Methodology in Clinical Trials of Drug and Biological Products, https://www.fda.gov/regulatory-information/search-fda-guidance-documents/use-bayesian- methodology-clinical-trials-drug-and-biological-products (accessed 14 June 2026)
2026
-
[28]
The Positive False Discovery Rate: A Bayesian Interpretation and the q-Value
Storey JD. The Positive False Discovery Rate: A Bayesian Interpretation and the q-Value. The Annals of Statistics 2003; 31: 2013–2035
2003
-
[29]
R: A Language and Environment for Statistical Computing
R Core Team. R: A Language and Environment for Statistical Computing. Vienna, Austria: R Foundation for Statistical Computing, https://www.R-project.org/ (2023)
2023
-
[30]
Group sequential methods in the design and analysis of clinical trials
Pocock SJ. Group sequential methods in the design and analysis of clinical trials. Biometrika 1977; 64: 191–199
1977
-
[31]
The choice of sequential boundaries based on the concept of power spending
Bauer P. The choice of sequential boundaries based on the concept of power spending. Controlled Clinical Trials 1991; 12: 637
1991
-
[32]
A Multiple Testing Procedure for Clinical Trials
O’Brien PC, Fleming TR. A Multiple Testing Procedure for Clinical Trials. Biometrics 1979; 35: 549–556
1979
-
[33]
Interim analysis: The alpha spending function approach
Demets DL, Lan KKG. Interim analysis: The alpha spending function approach. Statist Med 1994; 13: 1341–1352
1994
-
[34]
Group Sequential Designs: A Tutorial
Lakens D, Pahlke F, Wassmer G. Group Sequential Designs: A Tutorial. Preprint, PsyArXiv. Epub ahead of print 27 January 2021. DOI: 10.31234/osf.io/x4azm
2021 doi
-
[35]
Control of Type I Error Rates in Bayesian Sequential Designs
Shi H, Yin G. Control of Type I Error Rates in Bayesian Sequential Designs. Bayesian Anal; 14. Epub ahead of print 1 June 2019. DOI: 10.1214/18-BA1109
2019 doi
-
[36]
A Bayesian sequential design using alpha spending function to control type I error
Zhu H, Yu Q. A Bayesian sequential design using alpha spending function to control type I error. Stat Methods Med Res 2017; 26: 2184–2196
2017
-
[37]
analysis.idx
Amrhein V, Greenland S, McShane B. Scientists rise up against statistical significance. Nature 2019; 567: 305–307. Supplementary material for the article Impact of interim analyses on Bayesian and frequentist operating characteristics of Bayesian clinical trials Clément R. Mas...
2019
Reviewed July 11, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.