Pith. sign in

REVIEW 26 references

An Economical Approach to Design Posterior Analyses

T0 review · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read Two sample sizes suffice to design Bayesian posterior analyses across the whole sample-size space.

desk verdict A useful, well-tested shortcut for Bayesian sample size design whose main working assumption (linear logit quantiles) is a heuristic with asymptotic support but no diagnostic. read the letter →

arxiv 2411.13748 v2 pith:O7LPRPPN submitted 2024-11-20 stat.ME

classification stat.ME
keywords samplesizedeterminationposteriorprobabilitiesoperatingcharacteristicspowertypeIerrorbootstrapconfidenceintervalsBernstein-vonMisesBayesiandesign
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper claims that the operating characteristics of a Bayesian posterior analysis—power and type I error rate—can be accurately estimated at every candidate sample size using simulations at only two sample sizes. The key is that the logits of the posterior probabilities under the null and alternative hypotheses change approximately linearly with the sample size, so quantiles of their sampling distributions can be linearly interpolated or extrapolated. If true, this reduces the computational burden of Bayesian sample-size determination from exploring many sample sizes to evaluating just two, and it also yields bootstrap confidence intervals for the recommended sample size and decision threshold. The method is illustrated on two clinical trial design examples, including one where the recommended design is verified by intensive confirmatory simulation.

What carries the argument

The central object is the logit of a Bernstein-von Mises proxy for the posterior probability of an interval hypothesis, defined as p(n)_delta,j,r = Phi((delta_U - theta_j,r)/$\sqrt$(I(theta_j,r)^{-1}) $\sqrt$(n) - $Phi^{{-1}}$(u_j,r)) minus the analogous term at delta_L; Theorem 1 gives the limiting slope of its logit as a linear function of n. That limiting slope is what licenses the two-sample-size interpolation in Algorithm 2, which sorts logits of estimated posterior probabilities at two sample sizes and connects matching order statistics by straight lines to approximate quantiles at all other sample sizes.

What would settle it

Run Algorithm 2 on a model where the asymptotic MLE normality conditions hold but the sample sizes of interest are small, then simulate the true sampling distributions at several intermediate sizes; if the power and type I error estimated from the two-point linear approximations differ materially from the directly simulated values at those intermediate sizes, the linear-quantile assumption fails.

Watch

Extended reading notes

Core claim

The paper establishes that the derivative with respect to the sample size of the logit of a large-sample proxy for the posterior probability of an interval hypothesis converges to a value determined by the distance of the true parameter from the hypothesis boundaries: it tends to (0.5 minus an indicator that the parameter lies outside the interval) times the smaller squared standardized distance to an endpoint. This limiting slope is used to justify modeling the quantiles of the logits of true posterior probabilities as linear functions of n, so that power and type I error throughout the sample size space can be estimated from simulations at just two sample sizes. The resulting algorithm returns the smallest sample size and critical value gamma satisfying both operating-characteristic criteria, and a bootstrap procedure quantifies simulation variability. In the paper's second example, the recommended design (35, 0.9564) is confirmed by intensive simulation to have power 0.8029 and type I error 0.0500.

Load-bearing premise

The quantiles of the logits of the true posterior probabilities are approximately linear functions of the sample size between the two simulated sizes, and the ordering of simulation repetitions (or of theta subgroups when the design prior is nondegenerate) is roughly preserved as n changes.

Editorial extensions

If this is right

  • Bayesian study designs can be evaluated for power and type I error across all sample sizes after simulating at only two sample sizes, drastically reducing computation relative to binary search or grid exploration.
  • Bootstrap confidence intervals for the optimal sample size and critical value can be computed from the two simulated sampling distributions, giving practitioners a practical way to assess simulation variability.
  • The same two-sample-size simulations can be repurposed to draw contour plots of power and type I error over the (n, gamma) space, enabling exploration of near-optimal designs at almost no extra cost.
  • The approach extends from posterior probabilities to Bayes factors, since Bayes factor decision rules can be viewed as a special case of posterior probability rules.
  • The recommended critical value need not equal 1 - alpha, and the paper shows that fixing gamma = 1 - alpha can fail to control type I error in finite samples or with informative priors.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A testable extension is to check whether the linear-in-n logit quantile approximation remains accurate when the design prior Psi_1 is nondegenerate without subgrouping by order statistics of theta, since the paper only groups by theta order statistics when Psi_1 varies.
  • The linear-quantile heuristic suggests a broader principle: for any posterior summary whose logit has a stable limiting slope, sample-size design could be done from two evaluation points; this invites analogous theorems for other decision summaries such as credible interval coverage or expected loss.
  • The method's reliance on matching order statistics across two sample sizes implicitly assumes that the ranks of simulation repetitions are preserved as n grows; when the true sampling distribution is a mixture under a nondegenerate Psi_1, that assumption only holds within theta subgroups, which is why the subgroup modification is needed.
  • If the linear approximation is exact, contour plots built from one Algorithm 2 run should coincide with brute-force simulation plots; the paper's visual agreement in both examples gives a cheap falsification check that could be automated for new models.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Circularity Check

1 steps flagged · score 2.0 of 10

Minor self-citation in the proxy motivation, but the two-sample-size extrapolation is an independent empirical approximation and is not circular.

  1. other [Section 3, paragraph following Eq. (3)]
    "Theorem 1 of Hagar and Stevens (2024) proved that under the conditions in Appendix A, the total variation distance between the sampling distribution of p(n)δ,j,r and that of Pr (H1|w(n)j,r ) converges in probability to 0 as n→∞. The estimates of power and the type I error rate based on the (proxy) sampling distribution of p(n)δ,j,r are therefore consistent as n→∞. This consistency is adequate for theoretical consideration of the proxy sampling distribution, but we do not use it in Section 4."

    This is a self-citation by the same authors used to bridge the proxy sampling distribution to the true posterior-probability sampling distribution. If Algorithm 2's extrapolation depended on this theorem alone, the argument would reduce to an unverified self-citation; however, the paper explicitly states the consistency result is not used in Section 4, and the final n2 recommendation is computed from empirical logit order statistics at n0 and n1. The self-citation is therefore not load-bearing and does not make the central prediction equivalent to its inputs.

full rationale

The paper's derivation chain is: Theorem 1 computes the limiting n-derivative of logit[p(n)δ,j,r] for the BvM-based proxy, fixing u and η. Algorithm 2 uses that slope only in Lines 2 and 6 to select a starting n1; the final recommendation in Lines 10-13 fits empirical logit order statistics at n0 and n1 and extrapolates. The estimated operating characteristics at n2 are the solution of this linear extrapolation, and the paper is transparent that suitability 'only relies on the accuracy of the empirically estimated slopes.' This is an approximation with an explicit validation step — Section 5.2 confirmatory simulation gives power 0.8029 and type I error 0.0500 at the recommended design — not a fitted parameter renamed as a prediction. The only quasi-circular element is the self-citation of Theorem 1 of Hagar and Stevens (2024) for the TV-distance consistency of the proxy, but the paper states this result is not used in Section 4, and the central algorithm stands on empirical quantiles from two independent simulation batches. The finite-sample linearity of true logit quantiles is a heuristic concern about scope and validity, not circularity.

Assumptions & free parameters 1 free parameters · 4 assumptions · 0 invented entities

The central claim rests on standard asymptotic theory, a self-cited consistency result, and two ad hoc assumptions about linearity and order preservation of logit quantiles. No invented entities are introduced.

free parameters (1)
  • Number of subgroups for nondegenerate Psi_1 = 10
    Example 2 splits logits into 10 subgroups by theta order statistics to handle mixture sampling distributions; the count is ad hoc and no sensitivity analysis is provided.
assumptions (4)
  • domain assumption The model f(w; eta) satisfies the Bernstein-von Mises regularity conditions (Appendix A.1) and the MLE asymptotic normality conditions (Appendix A.2).
    Theorem 1 and the motivating proxy require these standard conditions; they are stated in the supplement but not verified for the examples beyond plausibility.
  • domain assumption The proxy sampling distribution based on BvM approximates the true sampling distribution of posterior probabilities.
    The paper relies on Theorem 1 of Hagar and Stevens (2024), a self-cited published result, for consistency of the proxy in total variation distance.
  • ad hoc to paper The quantiles of the logits of the true posterior probabilities are approximately linear functions of n.
    This is the key assumption behind Algorithm 2, Line 11; it is motivated by Theorem 1 for the proxy but not proven for the true sampling distribution.
  • ad hoc to paper The ordering of logits across simulation repetitions is preserved as n changes, so the same order statistic can be paired at n0 and n1.
    The line construction in Algorithm 2, Line 11 pairs order statistics; the paper acknowledges this requires similar theta values or subgrouping for nondegenerate Psi_1.

how reviews work

0 comments
Cite this review

Pith. "Pith review of An Economical Approach to Design Posterior Analyses." pith.science (2026). https://pith.science/paper/O7LPRPPN

@misc{pith2026241113748,
  author       = {Pith},
  title        = {Pith review of: An Economical Approach to Design Posterior Analyses},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/O7LPRPPN}},
  note         = {Machine review of arXiv:2411.13748}
}
read the original abstract

To design Bayesian studies, criteria for the operating characteristics of posterior analyses - such as power and the type I error rate - are often assessed by estimating sampling distributions of posterior probabilities via simulation. In this paper, we propose an economical method to determine optimal sample sizes and decision criteria for such studies. Using our theoretical results that model posterior probabilities as a function of the sample size, we assess operating characteristics throughout the sample size space given simulations conducted at only two sample sizes. These theoretical results are used to construct bootstrap confidence intervals for the optimal sample sizes and decision criteria that reflect the stochastic nature of simulation-based design. We also repurpose the simulations conducted in our approach to efficiently investigate various sample sizes and decision criteria using contour plots. The broad applicability and wide impact of our methodology is illustrated using two clinical examples.

Figures

Figures reproduced from arXiv: 2411.13748 by the authors.

Figure 1
Figure 1. Left: Contour plots for the type I error rate and power from one sample size calculation for [PITH_FULL_IMAGE:figures/full_fig_p014_1.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

26 extracted references · 18 canonical work pages

  1. [1]

    Bernardo, J. M. and A. F. Smith (2009). Bayesian Theory , Volume 405. John Wiley & Sons

  2. [2]

    Berry, S. M., B. P. Carlin, J. J. Lee, and P. Muller (2010). Bayesian adaptive methods for clinical trials . CRC press

  3. [3]

    De Santis, and S

    Brutti, P., F. De Santis, and S. Gubbiotti (2014). Bayesian-frequentist sample size determination: a game of two priors. Metron\/ 72\/ (2), 133--151

  4. [4]

    De Santis, F. (2007). Using historical data for B ayesian sample size determination. Journal of the Royal Statistical Society: Series A (Statistics in Society)\/ 170\/ (1), 95--113

  5. [5]

    Deng, A. (2015). Objective B ayesian two sample hypothesis testing for online controlled experiments. In Proceedings of the 24th International Conference on World Wide Web , pp.\ 923--928

  6. [6]

    Efron, B. (1982). The Jackknife, the Bootstrap and Other Resampling Plans . SIAM

  7. [7]

    Adaptive designs for clinical trials of drugs and biologics — G uidance for industry

    FDA (2019). Adaptive designs for clinical trials of drugs and biologics — G uidance for industry. Center for D rug E valuation and R esearch, U.S. Food and Drug Administration, Rockville, MD

  8. [8]

    Golchi, S. (2022). Estimating design operating characteristics in B ayesian adaptive clinical trials. Canadian Journal of Statistics\/ 50\/ (2), 417--436

Show all 26 references
  1. [9]

    Golchi, S. and J. J. Willard (2024). Estimating the sampling distribution of posterior decision summaries in bayesian clinical trials. Biometrical Journal\/ 66\/ (8), e70002

  2. [10]

    Gubbiotti, S. and F. De Santis (2011). A B ayesian method for the choice of the sample size in equivalence trials. Australian & New Zealand Journal of Statistics\/ 53\/ (4), 443--460

  3. [11]

    Hagar, L. and N. T. Stevens (2024). Fast power curve approximation for posterior analyses . Bayesian Analysis\/ , 1 -- 26 doi.org/10.1214/24--BA1469

  4. [12]

    Jeffreys, H. (1935). Some tests of significance, treated by the theory of probability. In Mathematical Proceedings of the Cambridge Philosophical Society , Volume 31, No. 2, pp.\ 203--222. Cambridge University Press

  5. [13]

    Kass, R. E. and A. E. Raftery (1995). Bayes factors. Journal of the American Statistical Association\/ 90\/ (430), 773--795

  6. [14]

    Stallrich, S

    Larsen, N., J. Stallrich, S. Sengupta, A. Deng, R. Kohavi, and N. T. Stevens (2024). Statistical challenges in online controlled experiments: A review of A / B testing methodology. The American Statistician\/ 78\/ (2), 135--149

  7. [15]

    Lehmann, E. L. and G. Casella (1998). Theory of point estimation . Springer Science & Business Media

  8. [16]

    Morey, R. D. and J. N. Rouder (2011). Bayes factor approaches for testing interval null hypotheses. Psychological methods\/ 16\/ (4), 406--419

  9. [17]

    O’Hagan, A. and J. W. Stevens (2001). Bayesian assessment of sample size for clinical trials of cost-effectiveness. Medical Decision Making\/ 21\/ (3), 219--230

  10. [18]

    Shi, H. and G. Yin (2019). Control of type I error rates in B ayesian sequential designs. Bayesian Analysis\/ 14\/ (2), 399--425

  11. [19]

    Spiegelhalter, D. J., K. R. Abrams, and J. P. Myles (2004). Bayesian approaches to clinical trials and health-care evaluation , Volume 13. John Wiley & Sons

  12. [20]

    Spiegelhalter, D. J., L. S. Freedman, and M. K. Parmar (1994). Bayesian approaches to randomized trials. Journal of the Royal Statistical Society: Series A (Statistics in Society)\/ 157\/ (3), 357--387

  13. [21]

    Stevens, N. T. and L. Hagar (2022). Comparative probability metrics: Using posterior probabilities to account for practical equivalence in A / B tests. The American Statistician\/ 76\/ (3), 224--237

  14. [22]

    van der Vaart, A. W. (1998). Asymptotic Statistics . Cambridge Series in Statistical and Probabilistic Mathematics. Cambridge University Press

  15. [23]

    Wang, F. and A. E. Gelfand (2002). A simulation-based approach to bayesian sample size determination for performance under a given model and for separating models. Statistical Science\/ 17\/ (2), 193--208

  16. [24]

    Wilding, J. P., R. L. Batterham, S. Calanna, M. Davies, L. F. Van Gaal, I. Lingvay, B. M. McGowan, J. Rosenstock, M. T. Tran, T. A. Wadden, et al. (2021). Once-weekly semaglutide in adults with overweight or obesity. New England Journal of Medicine\/ 384\/ (11), 989--1002

  17. [25]

    Wilson, D. T., R. Hooper, J. Brown, A. J. Farrin, and R. E. Walwyn (2021). Efficient and flexible simulation-based sample size determination for clinical trials with multiple design parameters. Statistical methods in medical research\/ 30\/ (3), 799--815

  18. [26]

    Lehmann, E. L. and G. Casella (1998). Theory of Point Estimation . Springer Science & Business Media

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.