{"id":"49cc6199-25e5-4e42-8b3d-9544fd6ea35e","arxiv_id":"2607.22900","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"low","formal_verification":"none","parameter_count":3,"one_line_summary":"Bayesian group sequential designs can now be calibrated on commodity hardware using conjugate-mixture semi-simulation and a cached-per-look posterior tail probability sweep.","lead":"A new computation framework makes Bayesian group sequential trial calibration fast enough for routine use by replacing per-look Markov-chain Monte Carlo with closed-form conjugate-mixture updates and a reusable lookup cache. On a re-design of a 3,800-patient trial, it matches two existing simulators within Monte Carlo error while running far faster.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Prior-approximation error in posterior tails is self-admittedly unquantified; one mild logit-normal example does not secure the broad-flexibility claim.","rationale":"The paper's central speed claim is well supported: Lemma 1 is correct for fixed priors and fixed analysis models, and the empirical benchmarks show agreement within Monte Carlo error for both conjugate and non-conjugate priors. The reader's weakest_assumption identifies exactly the same area as I do: the finite Beta-mixture approximation of an arbitrary user prior. The paper itself flags that the effect of the prior approximation on calibrated thresholds has not been quantified, and Appendix A.1 notes that prior-level diagnostics are necessary but not sufficient. The single non-conjugate example uses a mild prior with small effective sample size, so it does not establish the broad flexibility claim made in the introduction. This is not an internal inconsistency or a flawed proof; it is a missing support for a claimed scope. A conditional acceptance is appropriate: the core method is sound, but the paper should either bound the approximation error in posterior tail probabilities or demonstrate robustness across a wider, more challenging set of priors before the broad flexibility claim is accepted as is.","tokens_in":30719,"tokens_out":16768,"duration_ms":183336,"concrete_test":"Run the Section 3.3.2 comparison again with (a) a logit-normal prior with σ=1.5 (heavy tails) and (b) a two-peaked Beta-mixture prior, at R=10^6 for the five-look ADRENAL design, comparing type I error and power against adaptr's exact inverse-CDF posterior. If the largest absolute difference exceeds about 0.1 percentage points (roughly 5× the binomial MC SE at R=10^6), the finite-mixture approximation is not generically negligible, and the paper should condition its flexibility claim on the Appendix A.1 diagnostics rather than presenting the approximation as broadly validated.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central cache-invariance claim (Lemma 1) is mathematically sound and the benchmarks reproduce adaptr/BATSS within Monte Carlo error. The load-bearing soft spot is the prior-approximation step. Section 2.3 replaces an arbitrary user prior with a finite Beta-mixture and selects L by forward KL. Section 4 explicitly states: 'The prior approximation introduces additional error, whose effect on the calibrated thresholds has not yet been quantified.' Appendix A.1 concedes that prior-level agreement 'need not contract uniformly to the calibrated operating characteristics.' The only non-conjugate validation (Section 3.3.2) uses a mild logit-normal prior with effective sample size roughly 19 per arm, far smaller than the 3,800-patient trial; no stress test varies prior shape or informativeness near the decision boundary. Because the cache stores posterior tail probabilities computed under the approximate prior, any approximation error propagates to every design evaluated from that cache. Thus the framework's broad-flexibility claim rests on an unquantified approximation, even though the demonstrated ADRENAL results themselves are credible. This is the weakest load-bearing link in the generality claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces a semi-simulation framework for the evaluation and calibration of Bayesian group sequential designs with posterior-probability decision rules. Two devices are proposed: (i) a finite Beta-mixture approximation of arbitrary priors, which preserves conjugate per-look updates and reduces decision-rule tail probabilities to one-dimensional numerical integrals; and (ii) a caching scheme (Lemma 1) whereby a single Monte Carlo pass at the union of all candidate look times stores per-look posterior tail probabilities, so that every design in a calibration grid—varying in number/timing of looks, thresholds, and binding/non-binding futility—can be evaluated by a cache sweep without further simulation. The framework is applied to a Bayesian re-design of the ADRENAL trial with up to nine analyses, benchmarked against adaptr and BATSS under both a conjugate Beta(1,1) prior and an informative logit-normal prior, and supplemented by a simulation-budget study and approximation diagnostics.","tokens_in":31026,"tokens_out":5475,"duration_ms":61990,"significance":"If the result holds, the caching strategy is a substantive and practically important contribution: it decouples trial simulation from design evaluation, making joint calibration of decision thresholds and design skeletons feasible on commodity hardware. The paper ships reproducible code and cached outputs, and the central cache-invariance lemma is mathematically sound. The head-to-head validations are credible: at R=10^6, the proposed framework and adaptr agree to within 0.05 percentage points on type I error under both conjugate and logit-normal priors, and the speedups over adaptr and BATSS are large. The main weakness is that the broad flexibility claim rests on an unquantified prior-approximation error, although the authors are transparent about this limitation and provide useful diagnostics for one non-conjugate example.","major_comments":[{"comment":"The prior-approximation error in posterior tail probabilities is not quantified, and this is load-bearing for the framework's generality claim. Eq. (10) bounds prior-level total variation only; Appendix A.1 explicitly says that prior-level agreement 'need not contract uniformly to the calibrated operating characteristics,' and Section 4 states that the prior approximation's 'effect on the calibrated thresholds has not yet been quantified.' Because Lemma 1 caches tail probabilities computed under the approximate prior, this error propagates to every design evaluated from the cache. The single non-conjugate validation (logit-normal prior with ESS about 19 per arm, Section 3.3.2) is a mild prior and does not stress-test prior shape or informativeness near the decision boundary. The operating-characteristic sensitivity in Table A.2 is reassuring for that one design but does not quantify the","section":"Section 2.3 and Section 4; Eq. (10); Tables A.1–A.2"},{"comment":"The effect-and-control prior specification discards between-arm dependence by fitting separate marginal Beta mixtures, justified only by a Bernstein–von Mises heuristic ('this has little effect on the operating characteristics... as information accrues'). No quantified bound or simulation evidence is given for this step. Early interim looks, small samples, or strongly dependent priors could make the discarded dependence non-negligible. At minimum, provide a numerical experiment with a correlated prior on (control rate, effect) comparing the separate-marginal approximation against a joint fit or the exact posterior, especially at the earliest analysis times.","section":"Section 2.3 (effect-and-control specification)"}],"minor_comments":[{"comment":"The caption states 'All three methods use independent Beta(1,1) priors on the per-arm rates, except that BATSS uses...' — this is internally contradictory. Please rephrase to describe the BATSS prior separately from the proposed method and adaptr.","section":"Tables 1–2 captions"},{"comment":"The text in Section 3.3 describes the one-time simulation pass as 'about one minute,' while Figure 1's caption says the dotted intercept is 'about 30 seconds.' Please reconcile these timings.","section":"Section 3.4 and Figure 1"},{"comment":"The sign convention for Δ is reversed relative to Section 2.2 (efficacy is Δ<0 for ADRENAL). This is explained, but a one-sentence pointer at the first use in Section 2.2 would prevent confusion.","section":"Section 3.2"},{"comment":"Algorithm 1 reports E(N) but not E(T), despite Eq. (16) defining expected study duration. Clarify that E(T) is computed separately, or include it in the algorithm output.","section":"Section 2.4 / Algorithm 1"},{"comment":"The construction of the 'union of all candidate look times' for equally spaced schedules with K=1,...,10 is not specified at the level of integer sample sizes. Information fractions such as 1/3 or 2/5 do not map exactly to integer m_k/n_k unless rounding is defined. Please state how cumulative sample sizes at look times are discretized, since cache validity requires exact agreement.","section":"Section 3.4 / Lemma 1"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things you should know. The cache-invariance lemma (Lemma 1) is the real contribution: simulate once at the union of candidate look times, cache per-look posterior tails, then evaluate any design in the grid with sub-second sweeps. That is a genuine architectural advance over BATSS and adaptr, and the ADRENAL benchmark supports it — agreement within Monte Carlo error at both conjugate and non-conjugate priors, with the expected large speedups. The paper is also unusually honest about its own limits: Section 4 explicitly states that the prior-approximation error has not been quantified, and Appendix A.1 concedes that prior-level agreement need not contract uniformly to calibrated operating characteristics. So the stress-test concern is fair but it names something the authors already flag, not a hidden flaw.\n\nWhat is actually new is the precomputation strategy itself, plus the closed-form conjugate-mixture machinery that makes it exact. The quadrature-convergence diagnostics, the simulation-budget analysis, and the plausibility sweep over control rates are all done properly. Code and data are provided. The paper does not oversell: it positions itself as complementary to existing software and notes the like-for-like BATSS comparison is only approximate due to INLA's prior parameterization.\n\nThe soft spots are real but proportionate. The broad-flexibility claim rests on a single mild logit-normal prior (ESS around 19 per arm), and the effect-and-control specification discards between-arm dependence with only a Bernstein–von Mises heuristic rather than a bound. Table A.2 shows insensitivity to mixture order for that one example, but that does not establish the approximation error is small for, say, a sharply informative prior or one with mass near the decision boundary. Minor issue: the per-design speedup over adaptr at matched R=5,000 is 7–16x, not the headline several-hundredfold, which only appears at R=10^6 — the paper is clear about this, but readers could overgeneralize.\n\nBottom line: the central contribution holds. The caching argument is mathematically sound and the empirical work is reproducible and honest. This paper deserves a serious referee and I would bring it to our reading group. The right refereeing outcome is likely acceptance with a request for a broader prior-approximation sensitivity analysis, which the authors already suggest as standard practice.","headline":"The cache-invariance precomputation is a genuine contribution and the validation is careful; the prior-approximation caveat the stress-test flags is real but already disclosed by the authors and does not block the paper.","tokens_in":31458,"tokens_out":1895,"would_cite":true,"duration_ms":23478,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62L10","62F15","62P10"],"pacs":[],"model":"deepseek-v4-flash","headline":"A single Monte Carlo pass can evaluate every candidate design in a Bayesian group sequential calibration grid, because per-look posterior tail probabilities are invariant to stopping thresholds and look schedules.","keywords":["Bayesian group sequential design","posterior probability decision rule","conjugate mixture prior","decision-threshold calibration","operating-characteristic evaluation","precomputed cache","interim analysis"],"falsifier":"For a stress-test prior (e.g., sharply bimodal or with substantial mass near 0 or 1), run the proposed Beta-mixture pipeline at R=10^6 on the five look schedules and compare type I error and power against an exact posterior sampler that does not approximate the prior; if any difference exceeds the binomial Monte Carlo SE (about 0.02 percentage points), the claim that finite-mixture approximation preserves decision-rule accuracy fails. A second, sharper check: add a look time outside the cached union and verify that operating characteristics change only after a fresh simulation pass—if the cach","tokens_in":30645,"feed_emoji":"⚡","tokens_out":8726,"duration_ms":75538,"temperature":0.7,"pith_summary":"Bayesian group sequential designs offer interpretable stopping rules and can borrow external evidence, yet they are rarely used in confirmatory trials, largely because evaluating their operating characteristics has meant re-simulating virtual trials for every candidate design. This paper argues that the bottleneck is removable: if the prior is approximated by a finite mixture of conjugate Beta priors, each look's posterior tail probability can be computed in closed form, and that tail probability depends only on the cumulative event counts and the effect threshold—not on the stopping thresholds, the futility convention, or the rest of the look schedule. A single Monte Carlo pass at the union of all candidate look times can therefore cache these probabilities, and every design in a calibration grid becomes a sub-second sweep over the same cache. Re-designing the ADRENAL trial with up to nine analyses, the framework reproduces the operating characteristics of existing Bayesian trial simulators within Monte Carlo error while cutting per-design wall-clock time by roughly an order of magnitude or more, and it calibrates a 438-design grid in about eight minutes. If it holds up, this makes joint calibration of decision thresholds and the design skeleton a routine workstation task.","feed_headline":"One simulation pass calibrates 438 Bayesian trial designs in minutes","feed_subtitle":"Cached per-look posterior probabilities let every candidate stopping rule be evaluated by a sub-second sweep.","key_machinery":"The engine is a cached per-look posterior tail probability built on finite Beta-mixture conjugate priors. Lemma 1 (cache invariance) states that P(Δ>δ | i, j) at a given look is a function only of the cumulative event counts (i, j) and the effect threshold δ, not of any stopping thresholds, the binding/non-binding futility flag, or the number or timing of looks. The framework computes these tail probabilities by closed-form conjugate updates plus a one-dimensional Gauss–Legendre integral, caches them from one Monte Carlo pass at the union of all candidate look times, and then evaluates each candidate design by comparing its own thresholds against the cache—reducing calibration to a sweep.","core_discovery":"The central claim is the cache-invariance property: for a fixed prior and analysis model, the per-look posterior probability that the treatment effect exceeds a decision threshold depends only on the cumulative event counts at that look and the threshold itself. Because the framework approximates arbitrary priors by finite mixtures of conjugate Beta priors, these probabilities are available in closed form through one-dimensional numerical integration, and the same cached values serve every candidate stopping rule. On the ADRENAL re-design (up to nine analyses, 438 designs), the estimated type I error, power, and expected sample size agree with two established simulation implementations withi","pith_inferences":["The cache architecture suggests that design-search objectives—power, expected sample size, expected duration—can be optimised directly over the cached surface with off-the-shelf optimisers, since each query costs milliseconds; the paper stops at grid sweeps and manual selection.","A parallel cache might be built for predictive-probability decision rules by caching posterior predictive densities rather than tail probabilities; the paper explicitly leaves predictive-probability monitoring out of scope, but the same invariance argument would apply.","In the effect-and-control prior specification, the framework fits the two arms' marginals separately and discards the induced between-arm dependence, justified only by a large-sample posterior-concentration heuristic; a small-trial or strong-prior regime is a natural stress test of that assumption.","If the same decoupling works for non-conjugate settings through fast deterministic posterior approximations, the idea could extend beyond fixed-allocation posterior-probability GSDs to response-adaptive or covariate-adjusted designs, which the paper lists as future work."],"forward_implications":["Joint calibration of decision thresholds and the design skeleton—look count, look timing, futility binding convention—becomes a fast cache sweep rather than one fresh simulation per candidate, so large design grids become explorable interactively.","Because 10^5 virtual trials already put Monte Carlo SE for the type I error below 0.1 percentage point and 10^6 below 0.02, high-precision calibration against a 2.5% one-sided target is affordable on a workstation.","The same cached pass evaluates designs with smaller maximum sample sizes as long as their analysis times lie within the simulated union, so maximum sample size can be optimised jointly with thresholds without re-simulation.","The operating-characteristic surfaces expose a systematic pattern: fixed posterior-probability thresholds inflate type I error as the number of looks grows (roughly 1.0% at one look to 4.1% at nine looks in the studied setting), so calibration is not optional.","Closed-form per-look updates are derived for binary, continuous, count, and time-to-event endpoints through their natural conjugate pairs, though only the binary endpoint is benchmarked here."],"fun_headline_variants":["Cached posteriors make Bayesian trial design calibration sub-second","Semi-simulation framework speeds Bayesian group sequential design","Cache-invariance: one MC pass then each design evaluated in <1s","Bayesian GSD calibration: 7-16x faster via conjugate mixtures"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The broad flexibility claim rests on the premise that an arbitrary user prior can be replaced by a small finite Beta-mixture without moving the posterior tail probabilities that drive stopping decisions by more than Monte Carlo error—demonstrated for one logit-normal prior, not guaranteed for all priors.","fun_headline_variants_meta":{"raw":{"variants":["Cached posteriors make Bayesian trial design calibration sub-second","Semi-simulation framework speeds Bayesian group sequential design","Cache-invariance: one MC pass then each design evaluated in <1s","Bayesian GSD calibration: 7-16x faster via conjugate mixtures"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000211,"raw_usage":{"total_tokens":1295,"prompt_tokens":832,"completion_tokens":463,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":576,"completion_tokens_details":{"reasoning_tokens":388}},"tokens_in":576,"tokens_out":463,"duration_ms":5274,"temperature":1.0,"reasoning_tokens":388,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T04:10:56.870371+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"For a stress-test prior (e.g., sharply bimodal or with substantial mass near 0 or 1), run the proposed Beta-mixture pipeline at R=10^6 on the five look schedules and compare type I error and power against an exact posterior sampler that does not approximate the prior; if any difference exceeds the binomial Monte Carlo SE (about 0.02 percentage points), the claim that finite-mixture approximation preserves decision-rule accuracy fails. A second, sharper check: add a look time outside the cached union and verify that operating characteristics change only after a fresh simulation pass—if the cach","supporting_citations":[],"review_version":1}