{"id":"b84219ea-abab-4711-b6d7-abfac9e77a2b","arxiv_id":"2411.13748","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A Bayesian study design method estimates power and type I error across all sample sizes from simulations at just two sample sizes, using an asymptotic linearity result for the logit of posterior probabilities.","lead":"This statistics paper proposes a faster way to design Bayesian clinical studies, using simulations at only two sample sizes to find the right study size and decision threshold. It derives a theoretical result about how posterior probabilities change with sample size, and tests the method on two semaglutide trial examples.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The two-sample-size method rests on empirical logit quantiles being linear in n; Theorem 1 only proves the limiting proxy slope, so finite-n linearity is an unvalidated heuristic.","rationale":"The reader's verdict is CONDITIONAL, and my stress-test agrees. The central method is elegant, and the examples include a confirmatory check at (35,0.9564) plus full-grid contour comparisons, which is real supporting evidence. However, the passage that carries the method — Algorithm 2, Line 11 — is an extrapolation of empirical logit quantiles between two sample sizes. The proof of Theorem 1 concerns the proxy p^{(n)} for a fixed pseudorandom point u; it gives the limiting slope and does not bound the error of interpolating true finite-sample quantiles. In fact, for a normal location model the exact quantile curves contain a √n term, so any claim of linearity is only asymptotic. The paper does not provide a diagnostic to tell when the two-point chord is accurate; it only says suitability relies on slope accuracy. A single closed-form normal test with small effect sizes would settle whether this is a benign heuristic or a correctness risk. Given the existing numerical validation, I would not reject, but the missing guarantee justifies keeping CONDITIONAL. No change to the reader's verdict is needed.","tokens_in":19146,"tokens_out":18740,"duration_ms":182907,"concrete_test":"Use a normal model with known variance σ², one-sided H1: θ=δ+dσ for d∈{0.25,0.5,1}, degenerate Ψ0 at θ=δ, α=0.05, β=0.2. Exact logit quantiles are q_n(τ)=logit(Φ(d√n+z_τ)) under H1 and q_n(τ)=logit(τ) under H0. Compute the exact optimal (n*,γ*) and run Algorithm 2 with m=10^5; compare n2 and γ to n*,γ*. If at any d Algorithm 2's recommendation does not satisfy the exact power≥0.8 and type I≤0.05 criteria, or n2 differs from n* by more than one unit, the two-point linearity assumption is not generally reliable and the paper's broad claim needs qualification.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Algorithm 2's final recommendation (Line 11) linearly interpolates each empirical logit order statistic between n0 and n1 and extrapolates to n2. The exact finite-n logit quantiles for even the simplest normal model are q_n(τ)=logit(Φ(θ√n/σ+z_τ)); these are not linear in n — asymptotically they have slope θ²/(2σ²) but also a √n term with coefficient proportional to z_τ. Theorem 1 only establishes the limiting slope of the proxy logit for a fixed u, not the finite-n linearity of the true sampling-distribution quantiles that Line 11 requires. The paper explicitly states suitability of n2 relies only on the accuracy of empirically estimated slopes (Section 4), but provides no diagnostic for when that accuracy fails. The two worked examples are reassuring, yet the central claim — accurate operating characteristics throughout the sample-size space from two simulated sample sizes — is broader than the validated cases. In particular, no example has a nondegenerate Ψ0, and the mixture-quantile case is exactly where linearity is least supported. A concrete failure of the linearity assumption in a simple closed-form model would directly undermine the claimed generality.","agreement_with_reader":"agree"},"referee_report":null,"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a genuinely useful paper with one real, fixable soft spot. The new Theorem 1 (limiting slope of the logit of a posterior probability as a function of n) is correct and original, and the two-sample-size algorithm is a practical way to cut simulation cost in Bayesian sample size determination. The bootstrap intervals and the repurposed contour plots are nice touches, and the numerical work is unusually careful: the authors rerun the whole procedure 1000 times in Example 2 and check coverage, which is the right thing to do.\n\nThe weak point is the step from the theorem to Algorithm 2. Theorem 1 is about a proxy sampling distribution, holding the simulation point u fixed. The algorithm instead assumes that the order statistics of the true posterior probabilities are linear in n, then extrapolates from n0 and n1 to n2. The finite-sample truth for even a one-parameter normal model is logit(Phi(a sqrt(n)+z_tau)); the derivative is a^2/2 + a z_tau/(2 sqrt(n)) + ... . So the line through two points is only asymptotically right, and there is no diagnostic for when the finite-n correction matters. The paper says the recommendation's validity 'only relies on the accuracy of the empirically estimated slopes' and then leaves it there. That's a real gap, especially since both examples use degenerate Psi0; the nondegenerate Psi0 case is mentioned but not tested.\n\nThat said, the gap is addressable. The two examples are encouraging, including a nondegenerate Psi1 with assurance, and the method clearly works in those settings. What's missing is broader simulation support across models and a rule of thumb for when the linear approximation is safe. The self-citation to Hagar and Stevens (2024) is legitimate: Theorem 1 depends on their earlier consistency result, and they say so.\n\nWho should read this: statisticians who design Bayesian trials or A/B tests and want to explore the (n, gamma) space cheaply. The paper deserves a serious referee. A good referee should ask for a simulation study that varies the information in the prior, the distance of the null from the true effect, and the tail quantile, plus some kind of check on the slope estimate. With that, the method would be much more convincing.\n\nVerdict: send to peer review. The central idea is solid, the execution is careful, and the main weakness is a heuristic that can be either justified or carefully scoped in revision.","headline":"A useful, well-tested shortcut for Bayesian sample size design whose main working assumption (linear logit quantiles) is a heuristic with asymptotic support but no diagnostic.","tokens_in":19885,"tokens_out":5117,"would_cite":true,"duration_ms":43381,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Two sample sizes suffice to design Bayesian posterior analyses across the whole sample-size space.","keywords":["sample size determination","posterior probabilities","operating characteristics","power","type I error","bootstrap confidence intervals","Bernstein-von Mises","Bayesian design"],"falsifier":"Run Algorithm 2 on a model where the asymptotic MLE normality conditions hold but the sample sizes of interest are small, then simulate the true sampling distributions at several intermediate sizes; if the power and type I error estimated from the two-point linear approximations differ materially from the directly simulated values at those intermediate sizes, the linear-quantile assumption fails.","tokens_in":18928,"feed_emoji":"📊","tokens_out":1500,"duration_ms":16478,"temperature":0.7,"pith_summary":"This paper claims that the operating characteristics of a Bayesian posterior analysis—power and type I error rate—can be accurately estimated at every candidate sample size using simulations at only two sample sizes. The key is that the logits of the posterior probabilities under the null and alternative hypotheses change approximately linearly with the sample size, so quantiles of their sampling distributions can be linearly interpolated or extrapolated. If true, this reduces the computational burden of Bayesian sample-size determination from exploring many sample sizes to evaluating just two, and it also yields bootstrap confidence intervals for the recommended sample size and decision threshold. The method is illustrated on two clinical trial design examples, including one where the recommended design is verified by intensive confirmatory simulation.","feed_headline":"Two simulations design a Bayesian study at any sample size","feed_subtitle":"Logits of posterior probabilities grow nearly linearly in n, so power and error can be checked without rerunning for every size.","key_machinery":"The central object is the logit of a Bernstein-von Mises proxy for the posterior probability of an interval hypothesis, defined as p(n)_delta,j,r = Phi((delta_U - theta_j,r)/$\\sqrt$(I(theta_j,r)^{-1}) $\\sqrt$(n) - $Phi^{{-1}}$(u_j,r)) minus the analogous term at delta_L; Theorem 1 gives the limiting slope of its logit as a linear function of n. That limiting slope is what licenses the two-sample-size interpolation in Algorithm 2, which sorts logits of estimated posterior probabilities at two sample sizes and connects matching order statistics by straight lines to approximate quantiles at all other sample sizes.","core_discovery":"The paper establishes that the derivative with respect to the sample size of the logit of a large-sample proxy for the posterior probability of an interval hypothesis converges to a value determined by the distance of the true parameter from the hypothesis boundaries: it tends to (0.5 minus an indicator that the parameter lies outside the interval) times the smaller squared standardized distance to an endpoint. This limiting slope is used to justify modeling the quantiles of the logits of true posterior probabilities as linear functions of n, so that power and type I error throughout the sample size space can be estimated from simulations at just two sample sizes. The resulting algorithm returns the smallest sample size and critical value gamma satisfying both operating-characteristic criteria, and a bootstrap procedure quantifies simulation variability. In the paper's second example, the recommended design (35, 0.9564) is confirmed by intensive simulation to have power 0.8029 and type I error 0.0500.","pith_inferences":["A testable extension is to check whether the linear-in-n logit quantile approximation remains accurate when the design prior Psi_1 is nondegenerate without subgrouping by order statistics of theta, since the paper only groups by theta order statistics when Psi_1 varies.","The linear-quantile heuristic suggests a broader principle: for any posterior summary whose logit has a stable limiting slope, sample-size design could be done from two evaluation points; this invites analogous theorems for other decision summaries such as credible interval coverage or expected loss.","The method's reliance on matching order statistics across two sample sizes implicitly assumes that the ranks of simulation repetitions are preserved as n grows; when the true sampling distribution is a mixture under a nondegenerate Psi_1, that assumption only holds within theta subgroups, which is why the subgroup modification is needed.","If the linear approximation is exact, contour plots built from one Algorithm 2 run should coincide with brute-force simulation plots; the paper's visual agreement in both examples gives a cheap falsification check that could be automated for new models."],"forward_implications":["Bayesian study designs can be evaluated for power and type I error across all sample sizes after simulating at only two sample sizes, drastically reducing computation relative to binary search or grid exploration.","Bootstrap confidence intervals for the optimal sample size and critical value can be computed from the two simulated sampling distributions, giving practitioners a practical way to assess simulation variability.","The same two-sample-size simulations can be repurposed to draw contour plots of power and type I error over the (n, gamma) space, enabling exploration of near-optimal designs at almost no extra cost.","The approach extends from posterior probabilities to Bayes factors, since Bayes factor decision rules can be viewed as a special case of posterior probability rules.","The recommended critical value need not equal 1 - alpha, and the paper shows that fixing gamma = 1 - alpha can fail to control type I error in finite samples or with informative priors."],"supporting_citations":[{"why":"Supplies the Bernstein-von Mises theorem and its regularity conditions used to construct the proxy posterior probability in Theorem 1.","marker":"van der Vaart (1998)"},{"why":"Supplies the regularity conditions for asymptotic normality of the MLE that Theorem 1 also requires for the data-generating model.","marker":"Lehmann and Casella (1998)"},{"why":"Provides the result that the total variation distance between the proxy and true sampling distributions of posterior probabilities converges to zero, justifying the proxy.","marker":"Hagar and Stevens (2024)"},{"why":"Provides the percentile bootstrap method used to construct confidence intervals for the recommended sample size and critical value.","marker":"Efron (1982)"},{"why":"Sets the regulatory context requiring frequentist operating characteristics for Bayesian trial designs and the minimum of 10^4 simulation repetitions.","marker":"FDA (2019)"},{"why":"Supplies the two semaglutide clinical trial examples and their summary statistics used to specify the design priors in the numerical studies.","marker":"Wilding et al. (2021)"},{"why":"Supplies the standard uniform limiting distribution of posterior probabilities under the null hypothesis when the true parameter is at an interval endpoint.","marker":"Bernardo and Smith (2009)"}],"fun_headline_variants":["Two simulations size Bayesian trials across all sample sizes","Design Bayesian studies with just two simulation runs","Bayesian power and error from only two sample-size points","One line fits posterior logits, so two simulations suffice","Economical Bayesian design: simulate at two n, infer all n"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The quantiles of the logits of the true posterior probabilities are approximately linear functions of the sample size between the two simulated sizes, and the ordering of simulation repetitions (or of theta subgroups when the design prior is nondegenerate) is roughly preserved as n changes.","fun_headline_variants_meta":{"raw":{"variants":["Two simulations size Bayesian trials across all sample sizes","Design Bayesian studies with just two simulation runs","Bayesian power and error from only two sample-size points","One line fits posterior logits, so two simulations suffice","Economical Bayesian design: simulate at two n, infer all n"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000181,"raw_usage":{"total_tokens":1265,"prompt_tokens":858,"completion_tokens":407,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":474,"completion_tokens_details":{"reasoning_tokens":329}},"tokens_in":474,"tokens_out":407,"duration_ms":4575,"temperature":1.0,"reasoning_tokens":329,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T15:56:46.658489+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run Algorithm 2 on a model where the asymptotic MLE normality conditions hold but the sample sizes of interest are small, then simulate the true sampling distributions at several intermediate sizes; if the power and type I error estimated from the two-point linear approximations differ materially from the directly simulated values at those intermediate sizes, the linear-quantile assumption fails.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the percentile bootstrap method used to construct confidence intervals for the recommended sample size and critical value."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the two semaglutide clinical trial examples and their summary statistics used to specify the design priors in the numerical studies."}],"review_version":1}