Pith. sign in

REVIEW 2 major objections 4 minor 26 references

Under an asymptotic normal likelihood, posterior odds increase with the Wald statistic, so a Bayesian rejection rule can be calibrated to coincide exactly with the frequentist UMP test; the paper builds this into a group sequential design w

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-03 13:32 UTC pith:WRGPUWDK

load-bearing objection The fixed-prior calibration is clean and useful; the dynamic borrowing step in Remark (i) has a real double-use-of-data problem that the paper should fix before publication. the 2 major comments →

arxiv 2607.29108 v1 pith:WRGPUWDK submitted 2026-07-31 stat.ME

Frequentist-calibrated Bayesian group sequential design with dynamic borrowing

classification stat.ME MSC 62F1562L0562P10
keywords group sequential designdynamic borrowingposterior oddsfrequentist calibrationuniformly most powerful testmeta-analytic predictive priorHellinger distancetuberculosis prevention trial
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper aims to show that Bayesian and frequentist hypothesis testing in clinical trials need not be opposed: under an asymptotically normal estimator, posterior odds are a monotone function of the usual test statistic, which makes it possible to choose an evidential threshold that reproduces the frequentist uniformly most powerful test exactly. The authors embed this correspondence in a group sequential design and attach a second, non-informative threshold that lets the trial borrow historical information when the data and history agree. Dynamic borrowing is implemented through a meta-analytic-predictive prior whose historical weights are discounted by Hellinger-distance similarity, with the between-study variance chosen so historical evidence never outweighs current data. Numerical studies and a phase III tuberculosis-prevention example show the trade-off: borrowing raises power but inflates type I error when historical data conflict with the current null.

Core claim

The central discovery is the exact mapping between posterior-odds decision rules and the frequentist UMP test: with prior variance sigma^2/n_a, the posterior odds psi_pi(t,n) is increasing in t, so setting e_f = psi_pi(c_f,n) makes the rule "reject if psi_pi(t,n) > e_f" identical to "reject if t > c_f". A second threshold e_b = Phi(c_f)/(1-Phi(c_f)) is fixed under a non-informative prior and kept fixed as the analysis prior becomes informative, so all deviation from the frequentist decision is attributable to prior information. In the group sequential design, e_b,l is fixed in advance from the frequentist boundary c_f,l, while e_f,l is recomputed stage by stage from the current dynamic MAP p

What carries the argument

The load-bearing object is the posterior odds psi_pi(t,n) = Phi[(sqrt(n_a)t_a + sqrt(n)t)/sqrt(n_a+n)] / (1 - Phi[...]), an increasing function of the Wald statistic t. Its monotonicity converts the Bayesian rejection rule into a threshold on t, so the frequentist UMP threshold c_f maps to a Bayesian threshold e_f = psi_pi(c_f,n), and the non-informative threshold e_b = Phi(c_f)/(1-Phi(c_f)) serves as a fixed anchor. The other machinery is the dynamic MAP prior: historical studies are weighted by discounted sample sizes s_h * nu_h, with similarity s_h = 1 minus the Hellinger distance between current and historical estimator distributions; between-study variance tau^2 is updated at each inter

Load-bearing premise

The whole calibration rests on the claim in Remark (i) that data accrued up to stage l-1 enter the decision exactly once, through the prior, even though the same data are also part of the stage-l likelihood; if that double use is real, the reported power and type I error curves are not those of a coherent Bayesian update.

What would settle it

Run a two-stage simulation with known theta and the paper's MAP prior. At stage 2, compute the design's psi_pi(t_2,n_2) from prior weights based on D_1; separately compute a coherent posterior odds that updates a historical-only prior with the incremental data D_2 \ D_1, using the same Hellinger weights but never re-entering D_1 in the likelihood. If the distribution of the two odds differs systematically across repeated datasets, the 'no double use' claim fails; if they match, the concern is resolved.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • At every interim, e_f gives a Bayesian procedure with exactly the frequentist UMP test's rejection boundary, so sponsors can report both a Bayesian and a regulatory-calibrated result from the same posterior odds.
  • Because e_b is fixed before the trial from a non-informative prior, the gap between e_b and e_f is a transparent, interpretable measure of how much the analysis prior or borrowed historical data shifts the decision.
  • The Hellinger-distance weighting automatically drives similarity to near zero for historical sources that violate exchangeability; in the TB example, adding two mismatched rifapentine regimens leaves the operating characteristics essentially unchanged.
  • Choosing tau^2 = sigma^2/n_l ensures historical information never dominates: the prior effective sample size is no larger than the cumulative current sample size, moderating type I error inflation while preserving power gains.
  • The calibration is generic in the choice of prior, heterogeneity model, and similarity measure, so the same two-threshold logic can be applied with other dynamic borrowing schemes.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The paper's 'no double use' assertion is not demonstrated: D_{l-1} is a subset of D_l, so stage l-1 data influence both the prior weights and the stage-l likelihood. A direct check would compare the design's posterior odds with a genuinely coherent update that conditions only on the increment D_l \ D_{l-1} given a prior fixed from D_{l-1}; if the two diverge, the reported error and power curves ar
  • The exact UMP equivalence is an asymptotic, known-variance result. For small samples, binary or time-to-event endpoints, the normal approximation may fail and e_f no longer reproduces the frequentist test exactly; a simulation-based recalibration would be needed.
  • The ratio e_f/e_b at each stage could serve as a pre-specified, real-time measure of borrowing influence, analogous to a prior-data conflict diagnostic; trial monitors could stop borrowing if the ratio moves outside a pre-set range.
  • If the double-use concern is resolved, the same two-threshold design could be extended to adaptive sample size re-estimation: use e_b to decide whether to borrow and e_f to determine whether the frequentist operating characteristics are preserved.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. The paper proposes a group sequential Bayesian design with dynamic borrowing based on posterior odds. Section 3.1 establishes a monotone mapping between the Bayesian rejection rule ψ_π(t,n)>e and the frequentist UMP rejection t>c_f; setting e=e_f=ψ_π(c_f,n) retrieves the frequentist decision exactly, while e_b=Φ(c_f)/(1-Φ(c_f)) anchors the threshold to a flat prior. Section 3.2 extends the idea to L-stage designs with a dynamic MAP prior whose similarity weights are updated at stage ℓ using data accrued up to stage ℓ−1. Numerical studies (Tables 1–3) and a tuberculosis trial application (Tables 4–5, Figure 1) illustrate operating characteristics. The authors claim in Remark (i) of Section 3.2.3 that this dynamic construction avoids double use of the data.

Significance. The core calibration in Section 3.1 is correct and genuinely useful: for a normal likelihood with known variance and a normal analysis prior, the posterior odds is monotone in the test statistic, so the e_f rule is exactly the frequentist UMP test. This provides a transparent bridge between Bayesian and frequentist decision rules. The e_b threshold is a simple closed-form anchor to a non-informative prior. The simulation studies are extensive, and the provision of R code and a ShinyApp is a strength for reproducibility. The main weakness is in the dynamic borrowing extension: at stage ℓ, data up to stage ℓ−1 are used both to construct the analysis prior and as part of the likelihood, so the resulting e_b rule is not a coherent Bayesian posterior update and its operating characteristics cannot be interpreted as purely reflecting external/historical information.

major comments (2)
  1. [Section 3.2.3, Remark (i) and Eq. (16)] The claim that data accrued up to stage ℓ−1 contribute to the decision only once, 'avoiding a double use of the data', is incorrect. At stage ℓ the dynamic MAP prior in Eq. (15) is built from S_{ℓ−1}, which depends on D_{ℓ−1} through the similarity measures s_{h,ℓ}; the posterior odds in Eq. (16) then uses the likelihood of all cumulative data D_ℓ, which includes D_{ℓ−1}. Thus D_{ℓ−1} appears in both the prior and the likelihood. The quantity in Eq. (16) is therefore not a standard Bayesian posterior; it is a pseudo-posterior that over-weights early trial data. Consequently, Tables 4–5 and Figure 1 describe the operating characteristics of an ad hoc decision rule, not of a coherent Bayesian update, and the deviation of the e_b rule from the frequentist test cannot be attributed solely to external/historical prior information. The fix is straightforward: if the stage-ℓ prior uses D_{ℓ−1},
  2. [Section 3.1 vs. Section 3.2.3] The statement that any deviation from the frequentist UMP test is 'entirely due to the additional information encoded in the prior' is valid for the fixed-prior setting of Section 3.1 but not for the dynamic borrowing setting. In Section 3.2.3 the analysis prior itself is informed by the trial's own accumulating data (through the similarity measures), so the e_b rule's deviation is partly driven by internal trial data, not only by external information. This distinction should be made explicit, as it affects how readers interpret the power gains and type I error inflation reported in Section 5.
minor comments (4)
  1. [Eq. (16)] The displayed formula has an extra/mismatched parenthesis in the numerator: Φ( (√n_{a,ℓ} t_{a,ℓ} + √n_ℓ t_ℓ) / √(n_{a,ℓ}+n_ℓ) ). Please correct the typography.
  2. [Table 3 and caption] The caption says the table reports both γ̃_b and γ̃_f, but the tabulated entries appear to be only γ̃_b; γ̃_f is implicitly equal to the bold n_a=0 row. For non-flat priors, γ̃_f values are not displayed, although they are claimed to be the same across priors. Please make the table explicit or revise the caption.
  3. [Section 3.2.2] After the reparameterization of ω_h, the phrase 'historical sample sizes now replaced by their discounted counterparts s_hν_h' is exact only when τ²=0. For τ²>0, the effective sample size is further attenuated by τ²; consider clarifying this to avoid confusion.
  4. [Section 5] The construction of the combined normal analysis prior for the log-odds ratio from separate arm-specific MAP priors is deferred to Web Appendix D. A brief summary of this combination in the main text would improve readability and reproducibility.

Circularity Check

1 steps flagged

Central e_f–UMP equivalence is a sound reparameterization; dynamic borrowing Remark (i) double-counts D_{ℓ−1}.

specific steps
  1. other [Section 3.2.3, Remark (i); Eqs. (15)–(16)]
    "At each stage ℓ≥2, step (2) guarantees that the data accrued up to stage ℓ−1 contribute to the decision only once and exclusively through the posterior odds, thus avoiding a double use of the data."

    Step (2) constructs the stage-ℓ dynamic MAP prior from similarity measures s_{h,ℓ} based on θ̂(x_{ℓ−1}), hence on D_{ℓ−1}, while the posterior is then obtained by updating that prior with 'all current data up to stage ℓ', i.e. D_ℓ, which contains D_{ℓ−1}. Thus D_{ℓ−1} enters both the prior (through S_{ℓ−1} and n_{a,ℓ}) and the likelihood (through t_ℓ and n_ℓ) in Eq. (16). The stage-ℓ posterior odds is therefore not a standard Bayesian update; it is a pseudo-posterior that double-uses early trial data. Consequently, the claim that the e_b rule's deviation from the frequentist test is 'entirely due to the additional information encoded in the prior' is not supportable when the prior itself is informed by the trial's own accumulating data.

full rationale

The main calibration identity in Section 3.1 is not circular: e_f is defined as ψ_π(c_f,n), and because ψ_π(t,n) is monotonically increasing in t, the rule ψ_π(t,n) > e_f is algebraically equivalent to t > c_f for any fixed prior. This is a change of variable, not a fitted prediction, and the same reasoning applies stagewise for e_f,ℓ once the dynamic prior is conditionally fixed using only data up to ℓ−1. The numerical matching of frequentist rejection rates under e_f is a logical consequence of the identity, not an independent empirical confirmation. The self-citation to De Santis et al. (2026) is not load-bearing because the derivation is reproduced in the text. The genuine problem is in the dynamic borrowing extension: Remark (i)'s assertion that D_{ℓ−1} contributes only once is contradicted by the equations, since D_{ℓ−1} is a subset of D_ℓ and appears both in the similarity-weighted prior and in the likelihood of the posterior odds. This double use is a coherence flaw that weakens the interpretation of the dynamic-borrowing operating characteristics and of e_b as isolating purely external prior information, but it does not invalidate the central analytic e_f/e_b calibration. Hence the score is moderate rather than high.

Axiom & Free-Parameter Ledger

2 free parameters · 6 axioms · 0 invented entities

The paper does not introduce new physical or conceptual entities; the dynamic MAP prior and the e_b threshold are methods. The main free inputs are the borrowing hyperparameter tau^2 and the ad hoc cap on prior effective sample size. The central derivation rests on the standard normal-approximation and group-sequential distributional assumptions, plus the exchangeability and Hellinger-distance modeling choices.

free parameters (2)
  • tau^2 (between-study variance) = 0 or sigma^2/n_l (chosen by hand in Section 5)
    Controls the amount of historical borrowing; in the application the authors compare tau^2=0 with tau^2=sigma^2/n_l, which strongly affects operating characteristics.
  • Prior effective sample size cap n_max = n_l at each interim via tau^2=sigma^2/n_l
    An ad hoc design choice in Section 3.2.2 to prevent historical data from dominating the current trial; it is not derived from first principles.
axioms (6)
  • domain assumption Asymptotic normality of the MLE with known variance sigma^2
    Equation (1) and (3); the entire calibration and UMP equivalence require this approximation. In practice sigma^2 is estimated, so exactness is only asymptotic.
  • standard math Joint canonical distribution of sequential Wald statistics is multivariate normal with covariance sqrt(n_l/n_j)
    Equation (8), standard group sequential theory from Jennison and Turnbull (1997).
  • domain assumption Exchangeability of historical and current studies in the MAP hierarchical model
    Section 3.2.2; the dynamic MAP prior is only valid if the historical and current study parameters are exchangeable after accounting for between-study heterogeneity.
  • domain assumption Hellinger distance is an appropriate similarity measure with s_h = 1 - d
    Section 3.2.2 and Eq. (14); a modeling choice that determines how much historical information is retained.
  • ad hoc to paper tau^2 = sigma^2/n_l caps the prior effective sample size at the current trial sample size
    Introduced in Section 3.2.2 as a compromise; it is not derived from a formal optimality criterion.
  • ad hoc to paper Data D_{l-1} can be used to set prior weights without double-counting
    Asserted in Section 3.2.3 Remark (i); the paper claims no double use, but D_{l-1} enters both the prior via similarity and the likelihood via D_l.

pith-pipeline@v1.3.0-daily-deepseek · 15957 in / 20181 out tokens · 176708 ms · 2026-08-03T13:32:15.908394+00:00 · methodology

0 comments
read the original abstract

Bayesian analysis is increasingly used in clinical trials. However, assessment of the design with respect to the frequentist operating characteristics, such as type I error and power, remains a regulatory requirement in many cases. It is well established that, when information is borrowed from external sources to the trial, imposing strict frequentist type I error rate control is equivalent to offsetting the borrowing, which results in no power gains. We propose a Bayesian group sequential design with dynamic borrowing that exploits an explicit correspondence between Bayesian decision criteria based on posterior odds and frequentist uniformly most powerful (UMP) tests. At each interim analysis, two evidential thresholds are made available: the one that exactly retrieves the frequentist UMP decision; the other, that allows the investigator to incorporate historical information when appropriate. We assess the performance of the proposed approach in numerical studies, and apply the framework to the design of a phase III tuberculosis prevention trial, incorporating historical adult and pediatric trial data.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

26 extracted references · 3 linked inside Pith

  1. [1]

    Calderazzo, S., Wiesenfarth, M., Weru, V., and Kopp-Schneider, A. (2026). Principled type I error rate inflation in two-arm clinical trial designs with external control information borrowing.arXiv preprint arXiv:2508.16348. De Santis, F., Gubbiotti, S., and Mariani, F. (2026). Operating characteristics of Bayes factors.Journal of the Royal Statistical Soc...

  2. [2]

    Menzies, D. (2018). Safety and side effects of rifampin versus isoniazid in children.New England Journal of Medicine379,454–463

  3. [3]

    Dye, C., Glaziou, P., Floyd, K., and Raviglione, M. (2013). Prospects for tuberculosis elimination.Annual review of public health34,271–286. FDA (January 2026). Use of Bayesian methodology in clinical trials of drug and biological products. Frequentist-calibrated Bayesian group sequential design with dynamic borrowing25

  4. [4]

    Hagar, L., L., M., Golchi, S., and Menzies, D. (2025). An efficient approach to design Bayesian platform trials.arXiv preprint arXiv: 2507.12647

  5. [5]

    Harun, N., Liu, C., and Kim, M.-O. (2020). Critical appraisal of Bayesian dynamic borrowing from an imperfectly commensurate historical control.Pharmaceutical Statistics 19,613–625

  6. [6]

    Higgins, J., Thompson, S., and Spiegelhalter, D. (2009). A re-evaluation of random-effects meta-analysis.Journal of the Royal Statistical Society: Series A172,137–159

  7. [7]

    and Turnbull, B

    Jennison, C. and Turnbull, B. W. (1997). Group-sequential analysis incorporating covariate information.Journal of the American Statistical Association92,1330–1341

  8. [8]

    Kass, R. E. and Raftery, A. E. (1995). Bayes factors.Journal of the American Statistical Association90,773–795

  9. [9]

    Kopp-Schneider, A., Calderazzo, S., and Wiesenfarth, M. (2020). Power gains by using external information in clinical trials are typically not possible when requiring strict type I error control.Biometrical Journal62,361–374

  10. [10]

    Lesaffre, E., Qi, H., Banbeta, A., and van Rosmalen, J. (2024). A review of dynamic borrowing methods with applications in pharmaceutical research.Brazilian Journal of Probability and Statistics38,1–31

  11. [11]

    Mariani, F., De Santis, F., and Gubbiotti, S. (2024). A dynamic power prior approach to non-inferiority trials for normal means.Pharmaceutical Statistics23,242–256

  12. [12]

    Menzies, D. (2024). An adaptive trial to find the safest and shortest TB preventive regimens. ClinicalTrials.gov, Identifier No. NCT06498414.https://clinicaltrials.gov/study/ NCT06498414

  13. [13]

    Goldberg, H., Valiquette, C., Hornby, K., Dion, M.-J., Li, P.-Z., Hill, P., Schwartzman, K., and Benedetti, A. (2018). Four months of rifampin or nine months of isoniazid for latent tuberculosis in adults.New England Journal of Medicine379,440–453

  14. [14]

    Neuenschwander, B., Capkun-Niggli, G., Branson, M., and Spiegelhalter, D. J. (2010). Summarizing historical information on controls in clinical trials.Clinical Trials7,5–18

  15. [15]

    Ollier, A., Morita, S., Ursino, M., and Zohar, S. (2020). An adaptive power prior for sequential clinical trials - application to bridging studies.Statistical Methods in Medical Research29,2282–2294

  16. [16]

    and Held, L

    Pawel, S. and Held, L. (2026). Bayes factor group sequential designs.arXiv preprint arXiv: 2601.02851

  17. [17]

    Pocock, S. J. (1976). The combination of randomized and historical controls in clinical trials.Journal of Chronic Diseases29,175–188

  18. [18]

    Psioda, M. A. and Ibrahim, J. G. (2019). Bayesian clinical trial design using historical data that inform the treatment effect.Biostatistics20,400–415

  19. [19]

    Schmidli, H., Gsteiger, S., Roychoudhury, S., O’Hagan, A., Spiegelhalter, D., and Neuen- schwander, B. (2014). Robust meta-analytic-predictive priors in clinical trials with historical control information.Biometrics70,1023–1032

  20. [20]

    and Walker, S

    Shively, T. and Walker, S. (2018). On Bayes factors for the linear model.Biometrika105, 739–744

  21. [21]

    J., Abrams, K

    Spiegelhalter, D. J., Abrams, K. R., and Myles, J. P. (2004).Bayesian Approaches to Clinical Trials and Health-Care Evaluation. John Wiley & Sons

  22. [22]

    Conde, M., Bozeman, L., Horsburgh, C.R., J., and Chaisson, R. (2011). Three months of rifapentine and isoniazid for latent tuberculosis infection.New England Journal of Frequentist-calibrated Bayesian group sequential design with dynamic borrowing27 Medicine365,2155–2166

  23. [23]

    Mohapi, L., da Silva Escada, R., Mawlana, S., Banda, P., Severe, P., Hakim, J., Kanyama, C., Langat, D., Moran, L., Andersen, J., Fletcher, C., Nuermberger, E., and Chaisson, R. (2019). One month of rifapentine plus isoniazid to prevent HIV-related tuberculosis.New England Journal of Medicine380,1001–1011

  24. [24]

    Nakatani, H., and Raviglione, M. (2015). WHO’s new end TB strategy.Lancet (London, England)385,1799–1801

  25. [25]

    Ibrahim, J., Kinnersley, N., Lindborg, S., Micallef, S., Roychoudhury, S., and Thompson, L. (2014). Use of historical control data for assessing treatment effects in clinical trials. Pharmaceutical Statistics13,41–54

  26. [26]

    V., and Jaki, T

    Wadsworth, I., Hampson, L. V., and Jaki, T. (2018). Extrapolation of efficacy and other data to support the development of new medicines for children: A systematic review of methods.Statistical Methods in Medical Research27,398–413. World Health Organization (2025). Global tuberculosis report. Received July2026