Pith. sign in

REVIEW 1 major objections 4 minor 1 cited by

Sequential Bayes factor designs can be planned without Monte Carlo simulation: by writing Bayes factors as functions of the z-statistic, all design characteristics reduce to multivariate normal integrals.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-03 12:26 UTC pith:JE2IAB2N

load-bearing objection Worth a serious referee, but Table 1's directional-null vs. directional-alternative Bayes factor has an inverted prior-odds term that must be fixed. the 1 major comments →

arxiv 2601.02851 v2 pith:JE2IAB2N submitted 2026-01-06 stat.ME

Bayes Factor Group Sequential Designs

classification stat.ME MSC 62F1562L1062P10
keywords Bayes factorsequential designsgroup sequentialz-statisticmultivariate normal integrationdesign priorsample size determinationstopping rules
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper establishes that sequential Bayes factor designs can be planned without Monte Carlo simulation. The insight is to write the Bayes factor as a function of the z-statistic, which turns each stopping rule into a region — typically a rectangle — in the space of cumulative z-statistics. Under the canonical multivariate normal distribution of those statistics, the probability of stopping at each analysis, the expected sample size, and its variability all become multivariate normal integrals. This makes exploring designs with many interim looks fast and exact up to numerical integration error, as demonstrated on a clinical trial, a rat experiment, and a 61-look psychological design. If correct, it removes the main computational obstacle to broader use of sequential Bayes factor designs.

Core claim

The central claim is that the design characteristics of a sequential Bayes factor design — the per-analysis stopping probabilities, the probability of conclusive or misleading evidence, and the expected sample size and its standard deviation — can be computed by multivariate normal integration rather than simulation. This is achieved by expressing the Bayes factor at each analysis as a function of the z-statistic, so that the conditions 'stop for H0' and 'stop for H1' become intervals or unions of intervals in z-space. Combined with the canonical result that cumulative z-statistics follow a multivariate normal distribution (or a slightly richer normal when a normal design prior is used), eac

What carries the argument

The key machinery is a pair of translations. First, each Bayes factor threshold is converted into critical z-value(s) at each analysis, making the event 'stop at analysis i' a rectangle (or union of rectangles) in the space of cumulative z-statistics. Second, the joint distribution of those cumulative z-statistics is taken to be the canonical multivariate normal with mean θ times the square-root information vector and covariance equal to the square-rooted information ratio (Equation 1); a normal design prior adds a rank-one term to the covariance (Equation 2). Multivariate normal integration over the stopping rectangles then yields all design characteristics. For Bayes factors with two criti

Load-bearing premise

The whole calculation rests on the assumption that the accumulating z-statistics follow the canonical multivariate normal distribution given in Equation (1), which is exact when estimates are normally distributed with known variance and only approximate for binary outcomes, t-tests with small samples, or other non-normal settings.

What would settle it

Take a sequential Bayes factor design for a binary outcome with small per-group sample sizes where the normal approximation is doubtful, compute the stopping probabilities via the paper's multivariate-normal integration, and compare them to a very large Monte Carlo simulation based on the exact binomial likelihood; any discrepancy substantially larger than Monte Carlo error would refute the claim that the method is generally fast and accurate.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Sequential Bayes factor designs can be explored interactively: changing thresholds, analysis times, or maximum sample size yields updated stopping probabilities in seconds, not hours of simulation.
  • The framework makes classical frequentist calibrations (type-I error under a point null) a special case of the same computation, letting Bayesian and frequentist operating characteristics be reported side by side.
  • Designs with very many interim analyses — the paper demonstrates 61 looks in a t-test setting — become computationally feasible, matching or reproducing simulation-based results.
  • Expected sample size and its variability are obtained directly, supporting the ethical goal of stopping early in clinical trials and animal experiments.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Inference: the same z-statistic rectangle representation could be turned around to calibrate Bayes factor thresholds automatically, e.g., choosing k0 and k1 to hit a target expected sample size while bounding misleading evidence, which the paper does not explicitly pursue.
  • Inference: the approach's reliance on the canonical normal approximation suggests a natural test bed — small samples with binary or time-to-event outcomes, where the authors themselves note simulation may still be needed; one could check how far the approximation holds.
  • Inference: a possible extension is to non-normal design priors: the marginal in Equation (2) is derived under a normal design prior; other priors would require a mixture of normal integrals, which the framework could in principle accommodate.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

1 major / 4 minor

Summary. This paper presents a framework for the design and analysis of sequential Bayes factor designs. The key idea is to express Bayes factor stopping rules as functions of the z-statistic and to exploit the canonical multivariate normal distribution of cumulative z-statistics under a normal design prior. The authors show that the stopping regions form hyper-rectangles (for one-sided BFs) or unions of hyper-rectangles (for two-sided BFs), and that the probabilities of stopping for H0 or H1, the expected sample size, and its variability can be computed by multivariate normal integration without simulation. The approach is illustrated with a clinical trial, a rat experiment, and a sequential t-test example, and is implemented in the R package bfpwr. The paper also derives the marginal distribution of z-statistics under a normal design prior (Appendix B) and validates the approximation against existing simulation-based results.

Significance. If correct, the paper provides a valuable and practical advance: it makes the computation of sequential Bayes factor design characteristics as fast and convenient as classical group sequential design, while allowing for design priors that reflect parameter uncertainty. Strengths of the manuscript include a clear derivation of the marginal distribution of the z-statistics, a detailed description of the stopping-region geometry, reproducible code and data, an open-source R package, and an explicit validation against BFDA simulation results. The examples are well chosen and demonstrate the method's speed and scalability. However, the general correctness of the framework rests on the Bayes factor formulas in Table 1.

major comments (1)
  1. [Table 1, first row (Directional null vs. directional alternative)] The printed Bayes factor is BF01 = [1−Φ(μ*/τ*)]/[Φ(μ*/τ*)] × [1−Φ(μ/τ)]/[Φ(μ/τ)]. The first factor is the posterior odds for H0 vs H1, the second is the prior odds. Since BF01 is defined as the ratio of posterior to prior odds (p. 3), the correct formula is the first factor divided by the second. As σ→∞, μ*→μ and τ*→τ, so the correct formula converges to 1, while the printed formula converges to ([1−Φ(μ/τ)]/Φ(μ/τ))², which is 1 only when μ=0. For informative directional priors (μ≠0), the critical z-values derived in the same row are therefore wrong, and all design characteristics computed from them are biased. The paper's examples using this row (e.g., Figure 2) set μ=0, so the error is latent, but the table and the accompanying R package claim general applicability. This undermines the central claim of 'fast, accurate, and simulation-free' design evaluation for the Bayes factors in Tabl
minor comments (4)
  1. [Section 3.1, p. 11] For the point null vs. two-sided alternative Bayes factor, the text states that the H0 stopping condition is 'the z-statistic being in the interval around zero'. This is only correct when μ=0; in general the interval is [z_crit−(k0), z_crit+(k0)], centered at M = −μσ/τ² (as given in Table 1). The wording should be adjusted to avoid confusion for nonzero prior means.
  2. [Section 6, p. 25] Typo: 'by expressing Bayes factors as functions of z-statistics and and extending results' — remove the duplicate 'and'.
  3. [Section 4.1, p. 18] Typo: 'number of analysises' should be 'number of analyses'.
  4. [Abstract and Section 1] The abstract states that 'no closed-form or efficient numerical methods exist' for computing design characteristics. Since the proposed method is itself a numerical method, this phrasing is somewhat misleading; suggest rewording to 'no efficient numerical methods have been previously available' or similar.

Circularity Check

0 steps flagged

No significant circularity: design characteristics are computed from explicit BF-z formulas and the known canonical normal distribution; BFDA comparisons are independent benchmarks, not fitted inputs.

full rationale

Walking the derivation chain: the paper's design-characteristic formulas (Section 2) decompose stopping probabilities into sums of probabilities that the cumulative z-vector lies in BF-derived hyper-rectangular regions (Section 3.1). The z-distribution (Equations 1 and 2) is taken from classical group sequential theory (Jennison and Turnbull, external) or derived by a direct covariance calculation in Appendix B; it is neither fitted nor derived from the target characteristics. Table 1 provides explicit closed-form BF-z links, or defines numerical root-finding; these formulas are stated in full, not imported as an unverifiable black box. The 'predictions' (stopping probabilities, expected sample size) are mathematical consequences of those stated inputs. Validation against BFDA simulation in Section 5 is an independent benchmark, not a fit: no parameter is tuned to reproduce the reported 70.6%, 1.6%, and 69 values. Self-citations (Held and Ott 2018; Pawel and Held 2025; Pawel et al. 2023) support auxiliary statements (test-statistic BF class; t-test numerical critical value; replication design priors) and are accompanied either by explicit formulas or by external citations (e.g., Wong and Tendeiro 2025), so none is load-bearing. Section 6's limitations concern approximation validity, not circularity. The directional-null/directional-alternative BF formula flagged by a reviewer may be incorrect, but that is a correctness/accuracy issue, not an identity between input and output; it does not change the circularity verdict.

Axiom & Free-Parameter Ledger

0 free parameters · 4 axioms · 0 invented entities

The central claim relies on standard distributional assumptions from group sequential theory and the specific forms of Bayes factors for normal models. There are no fitted free parameters or invented entities; the user-specified thresholds, priors, and sample size schedule are part of the design input, not tuned to force the results.

axioms (4)
  • domain assumption Canonical distribution of cumulative z-statistics (Equation 1): Z|θ ~ N_m(θI, Σ) with Cov(Z_i,Z_j) = sqrt(I_i/I_j).
    Section 3.2 assumes this distribution, following Jennison and Turnbull (1999). It is exact for normal data with known variance; for other settings it is an approximation.
  • domain assumption Design prior θ ~ N(µ_d, τ_d²) leads to marginal distribution (2).
    Equation (2) and Appendix B derive the marginal using a normal design prior; non-normal priors would not yield the same form.
  • domain assumption Bayes factors in Table 1 are expressible as functions of the z-statistic for the listed hypothesis/prior combinations.
    The method requires this property (Section 3). It holds for the listed normal-theory models and approximately for the t-test with n≥30.
  • domain assumption For the sequential t-test, t|θ ≈ N(θ√n, 1) for n≥30.
    Section 5 invokes this approximation to embed the BF t-test in the canonical z-statistic framework.

pith-pipeline@v1.3.0-alltime-deepseek · 18869 in / 11446 out tokens · 98021 ms · 2026-08-03T12:26:57.046605+00:00 · methodology

0 comments
read the original abstract

The Bayes factor, the data-based updating factor from prior to posterior odds, is a principled measure of relative evidence for two competing hypotheses. It is naturally suited to sequential data analysis in settings such as clinical trials and animal experiments, where early stopping for efficacy or futility is desirable. However, designing such studies is challenging because computing design characteristics, such as the probability of obtaining conclusive evidence or the expected sample size, typically requires computationally intensive Monte Carlo simulations, as no closed-form or efficient numerical methods exist. To address this issue, we extend results from classical group sequential design theory to sequential Bayes factor designs. The key idea is to derive Bayes factor stopping regions in terms of the z-statistic and use the known distribution of the cumulative z-statistics to compute stopping probabilities through multivariate normal integration. The resulting method is fast, accurate, and simulation-free. We illustrate it with examples from clinical trials, animal experiments, and psychological studies. We also provide an open-source implementation in the bfpwr R package. Our method makes exploring sequential Bayes factor designs as straightforward as classical group sequential designs, enabling experiments to rapidly design informative and efficient experiments.

Figures

Figures reproduced from arXiv: 2601.02851 by Leonhard Held, Samuel Pawel.

Figure 1
Figure 1. Figure 1: Illustration of a sequential Bayes factor analysis. The Bayes factor quantifying the evidence for a point null hypothesis H0 : θ = 0 against an alternative hypothesis H1 : θ ̸= 0 is monitored as data accumulate (top). The bottom plot shows the corresponding z-statistic along with the Bayes factor stopping boundaries. of evidence for either hypothesis. After around 20 observations, the Bayes factor sur￾pass… view at source ↗
Figure 2
Figure 2. Figure 2: Critical z-values where BF01 = 1/10 for different types of Bayes factors (BF) from [PITH_FULL_IMAGE:figures/full_fig_p010_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Illustration of critical z-values such that BF01 ≥ k0 (dotted lines) and BF01 ≤ k1 (dashed lines) for Bayes factors with one critical value (left) and two critical values (right) such that BF01 = k. Integrating the colored regions produces the indicated analysis-wise stopping probabilitites. For example, integrating the green region gives the probability to not stop in the first analysis but stop for H1 in… view at source ↗
Figure 4
Figure 4. Figure 4: Sequential Bayes factor probabilities for the Low-PV trial (Barbui et al., 2021) [PITH_FULL_IMAGE:figures/full_fig_p016_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Sequential Bayes factor design characteristics for the Low-PV trial (Barbui et al., 2021) for differing maximum sample sizes and number of analyses. The dashed lines show the maximum sample sizes corresponding to 90% probability of correct evidence for 3 analyses. size for designs with many analyses (not shown). Increasing the number of analyses reduces the expected sample size (third-row plots) due to the… view at source ↗
Figure 6
Figure 6. Figure 6: Weight loss in rats assigned to different levels of a candidate drug (Kang et al., 2025) (top plot). The bottom plot shows a sequential Bayes factor analysis contrasting the null hy￾pothesis of no mean difference (H0 : θi = 0) to a mean difference of 5 grams (H0 : θi = 5) for each treatment group i ∈ {low, medium, high} indicated by the color. At each step an observation from control and treatment group is… view at source ↗
Figure 7
Figure 7. Figure 7: Sequential Bayes factor probabilities for replications of weight loss rat experiments (Kang et al., 2025). Probabilities are computed assuming a design prior corresponding to pos￾terior distributions for the mean difference θ based on data from the original experiments (in￾dicated in the plot panels). Analyses are assumed to be conducted after each pair of rats from control and treatment groups. Suppose th… view at source ↗
Figure 8
Figure 8. Figure 8: Sequential Bayes factor design probabilities based on t-test Bayes factor for design from Schonbrodt and Wagenmakers ¨ (2018) [PITH_FULL_IMAGE:figures/full_fig_p024_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: and [PITH_FULL_IMAGE:figures/full_fig_p035_9.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Frequentist-calibrated Bayesian group sequential design with dynamic borrowing

    stat.ME 2026-07 conditional novelty 5.0

    A Bayesian group sequential design provides, at each interim, an evidential threshold exactly matching the frequentist UMP test and a second threshold for dynamic borrowing of historical data.

Reference graph

Works this paper leans on

70 extracted references · 30 canonical work pages · cited by 1 Pith paper

  1. [1]

    Allen, D

    M. Allen, D. Poggiali, K. Whitaker, T. R. Marshall, J. van Langen , and R. A. Kievit. Raincloud plots: a multi-platform tool for robust data visualization. Wellcome Open Research, 4 0 (63), 2021. doi:10.12688/wellcomeopenres.15191.2

  2. [2]

    S. F. Anderson and K. Kelley. Sample size planning for replication studies: The devil is in the design. Psychological Methods, 2022. doi:10.1037/met0000520

  3. [3]

    Armitage, C

    P. Armitage, C. K. McPherson, and B. C. Rowe. Repeated significance tests on accumulating data. Journal of the Royal Statistical Society. Series A (General), 132 0 (2): 0 235, 1969. doi:10.2307/2343787

  4. [4]

    Barbui, A

    T. Barbui, A. M. Vannucchi, V. De Stefano, A. Masciulli, A. Carobbio, A. Ferrari, A. Ghirardi, E. Rossi, F. Ciceri, M. Bonifacio, A. Iurlo, F. Palandri, G. Benevolo, F. Pane, A. Ricco, G. Carli, M. Caramella, D. Rapezzi, C. Musolino, S. Siragusa, E. Rumi, A. Patriarca, N. Cascavilla, B. Mora, E. Cacciola, C. Mannarelli, G. G. Loscocco, P. Guglielmelli, S....

  5. [5]

    S. M. Berry, B. P. Carlin, J. J. Lee, and P. Muller. Bayesian Adaptive Methods for Clinical Trials . Chapman & Hall/CRC, 2010

  6. [6]

    J. M. Bland. Statistics notes: The odds ratio. BMJ, 320 0 (7247): 0 1468--1468, 2000. doi:10.1136/bmj.320.7247.1468

  7. [7]

    Campbell

    G. Campbell. FDA regulatory acceptance of Bayesian statistics. In Bayesian Methods in Pharmaceutical Research, pages 41--51. Chapman and Hall/CRC, 2020. ISBN 9781315180212. doi:10.1201/9781315180212-2

  8. [8]

    Cornfield

    J. Cornfield. Sequential Trials , Sequential Analysis and the Likelihood Principle . The American Statistician, 20: 0 18--23, 1966. doi:10.1080/00031305.1966.10479786

  9. [9]

    Cornfield

    J. Cornfield. Recent methodological contributions to clinical trials. American Journal of Epidemiology, 104 0 (4): 0 408--421, 1976. doi:10.1093/oxfordjournals.aje.a112313

  10. [10]

    D. B. Dahl, D. Scott, C. Roosen, A. Magnusson, and J. Swinton. xtable: Export Tables to LaTeX or HTML, 2019. URL https://CRAN.R-project.org/package=xtable. R package version 1.8-4

  11. [11]

    A. Deng, J. Lu, and S. Chen. Continuous monitoring of A/B tests without pain: Optional stopping in Bayesian testing. In 2016 IEEE International Conference on Data Science and Advanced Analytics (DSAA), pages 243--252, 2016. doi:10.1109/dsaa.2016.33

  12. [12]

    N. I. Drude, L. Martinez Gamboa, M. Danziger, U. Dirnagl, and U. Toelch. Improving preclinical studies through replications. eLife, 10, 2021. doi:10.7554/elife.62101

  13. [13]

    S. S. Ellenberg, T. R. Fleming, and D. L. DeMets. Data Monitoring Committees in Clinical Trials: A Practical Perspective. John Wiley & Sons, 2nd edition, 2019. doi:10.1002/9781119512684

  14. [14]

    Genz and F

    A. Genz and F. Bretz. Computation of Multivariate Normal and t Probabilities . Lecture Notes in Statistics. Springer-Verlag, Heidelberg, 2009. ISBN 978-3-642-01688-2

  15. [15]

    S. N. Goodman. Introduction to Bayesian methods I : measuring the strength of evidence. Clinical Trials, 2 0 (4): 0 282--290, 2005. doi:10.1191/1740774505cn098oa

  16. [16]

    Q. F. Gronau, A. Ly, and E.-J. Wagenmakers. Informed Bayesian t-tests. The American Statistician, 74 0 (2): 0 137--143, 2020. doi:10.1080/00031305.2018.1562983

  17. [17]

    Gsponer, F

    T. Gsponer, F. Gerber, B. Bornkamp, D. Ohlssen, M. Vandemeulebroecke, and H. Schmidli. A practical guide to Bayesian group sequential designs. Pharmaceutical Statistics, 13 0 (1): 0 71--80, 2013. doi:10.1002/pst.1593

  18. [18]

    Held and M

    L. Held and M. Ott. On p -values and B ayes factors. Annual Review of Statistics and Its Application, 5 0 (1), 2018. doi:10.1146/annurev-statistics-031017-100307

  19. [19]

    Jack Lee and C

    J. Jack Lee and C. T. Chu. Bayesian clinical trials in action. Statistics in Medicine, 31 0 (25): 0 2955--2972, 2012. doi:10.1002/sim.5404

  20. [20]

    Jeffreys

    H. Jeffreys. Theory of Probability. Oxford University Press, Oxford, 1961. 3rd edition

  21. [21]

    Jennison and B

    C. Jennison and B. W. Turnbull. Group Sequential Methods with Applications to Clinical Trials. Chapman & Hall, 1999

  22. [22]

    V. E. Johnson. Bayes factors based on test statistics. Journal of the Royal Statistical Society Series B: Statistical Methodology, 67 0 (5): 0 689--701, 2005. doi:10.1111/j.1467-9868.2005.00521.x

  23. [23]

    V. E. Johnson and J. D. Cook. Bayesian design of single-arm phase II clinical trials with continuous monitoring. Clinical Trials, 6 0 (3): 0 217--226, 2009. doi:10.1177/1740774509105221

  24. [24]

    J. A. Kairalla, C. S. Coffey, M. A. Thomann, and K. E. Muller. Adaptive trial designs: a review of barriers and opportunities. Trials, 13 0 (1), 2012. doi:10.1186/1745-6215-13-145

  25. [25]

    J. Kang, T. Koulis, and T. Pourmohamad. Sample size reduction in preclinical experiments: A Bayesian sequential decision-making framework. Journal of Biopharmaceutical Statistics, pages 1--16, 2025. doi:10.1080/10543406.2025.2556680

  26. [26]

    Kass and A

    R. Kass and A. Raftery. Bayes factors. Journal of the American Statistical Association, 90 0 (430): 0 773--795, June 1995. doi:10.1080/01621459.1995.10476572

  27. [27]

    Kassambara

    A. Kassambara. ggpubr: 'ggplot2' Based Publication Ready Plots, 2023. URL https://CRAN.R-project.org/package=ggpubr. R package version 0.6.0

  28. [28]

    Li, M.-H

    W. Li, M.-H. Chen, X. Wang, and D. K. Dey. Bayesian design of non-inferiority clinical trials via the Bayes factor. Statistics in Biosciences, 10 0 (2): 0 439--459, 2017. doi:10.1007/s12561-017-9200-5

  29. [29]

    Linde and D

    M. Linde and D. van Ravenzwaaij. baymedr : an R package and web application for the calculation of Bayes factors for superiority, equivalence, and non-inferiority designs. BMC Medical Research Methodology, 23 0 (1), 2023. doi:10.1186/s12874-023-02097-y

  30. [30]

    Lindon and A

    M. Lindon and A. Malek. Anytime-valid inference for multinomial count data. In S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh, editors, Advances in Neural Information Processing Systems, volume 35, pages 2817--2831, 2022. URL https://proceedings.neurips.cc/paper_files/paper/2022/file/12f3bd5d2b7d93eadc1bf508a0872dc2-Paper-Conference.pdf

  31. [31]

    N. Mani, M. S. Schreiner, J. Brase, K. K\" o hler, K. Strassen, D. Postin, and T. Schultze. Sequential Bayes factor designs in developmental research: Studies on early word learning. Developmental Science, 24 0 (4), 2021. doi:10.1111/desc.13097

  32. [32]

    J. N. Matthews. I ntroduction to R andomized C ontrolled C linical T rials . Chapman & Hall/CRC, second edition, 2006

  33. [33]

    Micheloud and L

    C. Micheloud and L. Held. Power calculations for replication studies. Statistical Science, 37 0 (3): 0 369--379, 2022. doi:10.1214/21-sts828

  34. [34]

    Moerbeek

    M. Moerbeek. Bayesian updating: increasing sample size during the course of a study. BMC Medical Research Methodology, 21 0 (1), 2021. doi:10.1186/s12874-021-01334-6

  35. [35]

    O Hagan and J

    A. O Hagan and J. Stevens. Bayesian assessment of sample size for clinical trials of cost-effectiveness. Medical Decision Making, 21 0 (3): 0 219--230, 2001. doi:10.1177/02729890122062514

  36. [36]

    Pawel and L

    S. Pawel and L. Held. The sceptical Bayes factor for the assessment of replication success. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 84 0 (3): 0 879--911, 2022. doi:10.1111/rssb.12491

  37. [37]

    Pawel and L

    S. Pawel and L. Held. Closed-form power and sample size calculations for Bayes factors. The American Statistician, 79 0 (3): 0 330--344, 2025. doi:10.1080/00031305.2025.2467919

  38. [38]

    Pawel, G

    S. Pawel, G. Consonni, and L. Held. Bayesian approaches to designing replication studies. Psychological Methods, 2023. doi:10.1037/met0000604

  39. [39]

    S. K. Piper, U. Grittner, A. Rex, N. Riedel, F. Fischer, R. Nadon, B. Siegerink, and U. Dirnagl. Exact replication: Foundation of science or game of chance? PLOS Biology, 17 0 (4): 0 e3000188, 2019. doi:10.1371/journal.pbio.3000188

  40. [40]

    Pourmohamad and C

    T. Pourmohamad and C. Wang. Sequential Bayes factors for sample size reduction in preclinical experiments with binary outcomes. Statistics in Biopharmaceutical Research, 15 0 (4): 0 706--715, 2022. doi:10.1080/19466315.2022.2123386

  41. [41]

    M. A. Psioda and J. G. Ibrahim. Bayesian clinical trial design using historical data that inform the treatment effect. Biostatistics, 20 0 (3): 0 400--415, 2018. doi:10.1093/biostatistics/kxy009

  42. [42]

    R: A Language and Environment for Statistical Computing

    R Core Team . R: A Language and Environment for Statistical Computing. R Foundation for Statistical Computing, Vienna, Austria, 2025. URL https://www.R-project.org/

  43. [43]

    D. S. Robertson, B. Choodari‐Oskooei, M. Dimairo, L. Flight, P. Pallmann, and T. Jaki. Point estimation for adaptive trial designs I : A methodological review. Statistics in Medicine, 42 0 (2): 0 122--145, 2022. doi:10.1002/sim.9605

  44. [44]

    G. L. Rosner. Bayesian adaptive designs in drug development. In Bayesian Methods in Pharmaceutical Research, pages 161--184. Chapman and Hall/CRC, 2020. doi:10.1201/9781315180212-8

  45. [45]

    G. L. Rosner, P. W. Laud, and W. O. Johnson. Bayesian Thinking in Biostatistics. Chapman and Hall/CRC, 2021. ISBN 9781439800102. doi:10.1201/9781439800102

  46. [46]

    J. N. Rouder, P. L. Speckman, D. Sun, R. D. Morey, and G. Iverson. Bayesian t tests for accepting and rejecting the null hypothesis. Psychonomic Bulletin & Review, 16 0 (2): 0 225--237, 2009. doi:10.3758/pbr.16.2.225

  47. [47]

    R. Royall. Statistical Evidence: A Likelihood Paradigm . Chapman & Hall, London New York, 1997. ISBN 9780412044113

  48. [48]

    W. M. S. Russell and R. L. Burch. The Principles of Humane Experimental Technique. Methuen, London, U.K., 1959

  49. [49]

    E. G. Ryan, K. Brock, S. Gates, and D. Slade. Do we need to adjust for interim analyses in a Bayesian adaptive trial design? BMC Medical Research Methodology, 20 0 (1), 2020. doi:10.1186/s12874-020-01042-7

  50. [50]

    F. D. Sch\" o nbrodt and E.-J. Wagenmakers. Bayes factor design analysis: Planning for compelling evidence. Psychonomic Bulletin & Review , 25 0 (1): 0 128--142, mar 2018. doi:10.3758/s13423-017-1230-y

  51. [51]

    F. D. Sch\" o nbrodt, E.-J. Wagenmakers, M. Zehetleitner, and M. Perugini. Sequential hypothesis testing with Bayes factors: Efficiently testing mean differences. Psychological Methods, 22 0 (2): 0 322--339, 2017. doi:10.1037/met0000061

  52. [52]

    F. D. Schönbrodt and A. M. Stefan. BFDA : An R package for Bayes factor design analysis (version 0.5.0) , 2019. URL https://github.com/nicebread/BFDA

  53. [53]

    D. J. Spiegelhalter, R. Abrams, and J. P. Myles. Bayesian Approaches to Clinical Trials and Health-Care Evaluation . New York: Wiley, 2004

  54. [54]

    A. M. Stefan, Q. F. Gronau, F. D. Sch\" o nbrodt, and E.-J. Wagenmakers. A tutorial on Bayes factor design analysis using an informed prior. Behavior Research Methods, 51 0 (3): 0 1042–1058, 2019. doi:10.3758/s13428-018-01189-8

  55. [55]

    A. M. Stefan, F. D. Sch\" o nbrodt, N. J. Evans, and E.-J. Wagenmakers. Efficiency in sequential testing: Comparing the sequential probability ratio test and the sequential Bayes factor test. Behavior Research Methods, 54 0 (6): 0 3100--3117, 2022. doi:10.3758/s13428-021-01754-8

  56. [56]

    A. M. Stefan, Q. F. Gronau, and E.-J. Wagenmakers. Interim design analysis using Bayes factor forecasts. Psychological Methods, 2024. doi:10.1037/met0000641

  57. [57]

    Stevely, M

    A. Stevely, M. Dimairo, S. Todd, S. A. Julious, J. Nicholl, D. Hind, and C. L. Cooper. An investigation of the shortcomings of the CONSORT 2010 statement for the reporting of group sequential randomised controlled trials: A methodological systematic review. PLOS ONE, 10 0 (11): 0 e0141104, 2015. doi:10.1371/journal.pone.0141104

  58. [58]

    Food and Drug Administration

    U.S. Food and Drug Administration . Guidance for the use of Bayesian statistics in medical device clinical trials, 2010. URL https://www.fda.gov/regulatory-information/search-fda-guidance-documents/guidance-use-bayesian-statistics-medical-device-clinical-trials

  59. [59]

    Verhagen and E.-J

    J. Verhagen and E.-J. Wagenmakers. Bayesian tests to quantify the result of a replication attempt. Journal of Experimental Psychology, 143 : 0 1457--1475, 2014. doi:10.1037/a0036731

  60. [60]

    A. Wald. Sequential Analysis. Wiley, New York, 1947

  61. [61]

    Wang and A

    F. Wang and A. E. Gelfand. A simulation-based approach to Bayesian sample size determination for performance under a given model and for separating models. Statistical Science, 17 0 (2): 0 193--208, 2002. doi:10.1214/ss/1030550861

  62. [62]

    Wassmer and W

    G. Wassmer and W. Brannath. Group Sequential and Confirmatory Adaptive Designs in Clinical Trials . Springer, New York, 2016

  63. [63]

    H. Wickham. ggplot2: Elegant Graphics for Data Analysis. Springer-Verlag New York, 2016. ISBN 978-3-319-24277-4. URL https://ggplot2.tidyverse.org

  64. [64]

    Wickham, R

    H. Wickham, R. François, L. Henry, K. Müller, and D. Vaughan. dplyr: A Grammar of Data Manipulation, 2023. URL https://CRAN.R-project.org/package=dplyr. R package version 1.1.4

  65. [65]

    T. K. Wong and J. N. Tendeiro. On a generalizable approach for sample size determination in Bayesian t tests. Behavior Research Methods, 57 0 (5), 2025. doi:10.3758/s13428-025-02654-x

  66. [66]

    Y. Xie. Dynamic Documents with R and knitr . Chapman and Hall/CRC, Boca Raton, Florida, 2nd edition, 2015. URL https://yihui.org/knitr/. ISBN 978-1498716963

  67. [67]

    Zellner and A

    A. Zellner and A. Siow. Posterior odds ratios for selected regression hypotheses. Trabajos de Estadistica Y de Investigacion Operativa, 31 0 (1): 0 585--603, 1980. doi:10.1007/bf02888369

  68. [68]

    Zhou and Y

    T. Zhou and Y. Ji. On Bayesian sequential clinical trial designs. The New England Journal of Statistics in Data Science, pages 136--151, 2023. doi:10.51387/23-nejsds24

  69. [69]

    Y. Zhou, R. Lin, and J. J. Lee. The use of local and nonlocal priors in Bayesian test-based monitoring for single-arm phase II clinical trials. Pharmaceutical Statistics, 20 0 (6): 0 1183--1199, 2021. doi:10.1002/pst.2139

  70. [70]

    L. Zhu, Q. Yu, and D. E. Mercante. A Bayesian sequential design for clinical trials with time-to-event outcomes. Statistics in Biopharmaceutical Research, 11 0 (4): 0 387--397, 2019. doi:10.1080/19466315.2019.1629996