{"id":"6444143c-53a0-4748-8a6c-36500c57e00b","arxiv_id":"2507.10269","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Using pilot data in a robust MAP prior reduces the confirmatory trial sample size in simulations, but the reported time savings omit the pilot phase and type I error control is not verified.","lead":"This paper simulates using data from a small pilot study as a Bayesian prior to shrink the sample size of a later confirmatory rare disease trial. It claims this saves time and makes recruitment targets easier to hit, but the time savings are calculated only for the confirmatory phase, not the pilot phase.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The borrowing-efficiency claims rest on an untested fixed-low-heterogeneity assumption and no type I error verification; a between-study heterogeneity sensitivity analysis is needed before the sample-size and duration reductions can be accepted.","rationale":"The paper's central claim is conditional on pilot and definitive trials being exchangeable or at least on between-study heterogeneity being small. The authors acknowledge in the Discussion that their single-pilot robust MAP construction implicitly assumes a fixed, low level of heterogeneity because that parameter cannot be estimated from one historical study. This is exactly the assumption that the reader's weakest-assumption analysis identifies, and it is the least secure point in the argument: the sample-size, duration, and recruitment-probability results in Figures 1-3 are all generated under the no-prior-data-conflict setting, while Figure 4 only perturbs the pilot risk ratio deterministically rather than simulating random heterogeneity. The paper also repeatedly claims type I error control but reports no type I error simulation, leaving the statistical-validity component of the central claim unverified. The duration and recruitment-probability results are deterministic consequences of the sample-size reduction, so they inherit the same fragility. The issue of excluding pilot-phase time and patients is real and the reader correctly notes it, but the heterogeneity and type I error gap is the more fundamental threat to the statistical claim. Because the reader's CONDITIONAL verdict already reflects the need for additional verification of these points, the stress-test does not change the verdict.","tokens_in":9981,"tokens_out":18731,"duration_ms":226348,"concrete_test":"Re-run the Section 4 simulation under a hierarchical data-generating process in which pilot and definitive arm probabilities are drawn from a common logit-normal model with between-study standard deviation tau in {0, 0.1, 0.25, 0.5}, keeping the marginal means at the paper's pC and pT values. For each tau, estimate the minimum definitive sample size achieving 80% power and the type I error at the 0.975 posterior-probability threshold. If the type I error exceeds the nominal 2.5% one-sided level, or if the sample-size reduction versus the no-pilot design disappears at tau = 0.25, then the headline borrowing-efficiency claim fails outside the no-conflict setting.","verdict_should_be":"UNCHANGED","load_bearing_attack":"In Section 3.3, because only one pilot study is available, the robust MAP prior is reduced to a two-component mixture with a fixed initial weight w = 0.5 and no between-study heterogeneity parameter. The paper explicitly concedes in the Discussion that this 'implicitly assumes a fixed, low level of heterogeneity' and that the heterogeneity parameter 'cannot be estimated with only one source of historical data.' The primary simulations in Section 4 fix pC and pT to the same values in the pilot and definitive trials, i.e., no prior-data conflict by construction. The conflict simulations behind Figure 4 only apply a deterministic attenuation factor to the pilot risk ratio; they do not generate pilot and definitive data from a common hierarchical model with random between-study variability, and they do not report type I error at the 0.975 threshold. If true between-study heterogeneity is larger than the implicit low level, the informative component is overtrusted, the data-dependent weight may not downweight enough, and the required sample-size reductions in Figure 1 as well as the duration and recruitment gains in Figures 2-3 could shrink or reverse. Equally, the claim that the robust MAP prior 'maintains type I error control even when disagreement exists' is asserted but never demonstrated under conflict. Since the central claim is precisely that pilot data make confirmatory trials smaller and faster without losing statistical validity, this untested assumption is load-bearing.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes using robust meta-analytic-predictive (MAP) priors to incorporate data from a single protocol-aligned pilot feasibility study into the design and analysis of a definitive confirmatory trial for a binary efficacy outcome in rare diseases. The authors specify a two-component mixture prior for each arm, with an informative Beta component updated from pilot data and a vague Beta(1,1) component, and an initial weight w=0.5. Through simulations, they claim that including pilot data at 10-40% of the definitive sample size reduces the required confirmatory sample size, shortens expected trial duration, and increases the probability of meeting recruitment targets. A secondary simulation applies deterministic attenuation factors to the pilot risk ratio to assess prior-data conflict. The paper is motivated by an IVIg de-escalation feasibility trial in autoimmune inflammatory myopathies.","tokens_in":10311,"tokens_out":4528,"duration_ms":47123,"significance":"If the central claims are fully supported, the paper would offer practical guidance for a common and important problem: leveraging pilot data to make rare disease confirmatory trials smaller and faster without compromising validity. The methodological core is standard—the posterior mixture formulas follow Schmidli et al. (2014) and are correctly derived—and the simulation code is provided in the supplementary materials, which is a strength. However, the operational conclusions (sample size, duration, recruitment probability) rest on assumptions that are not adequately tested, particularly the fixed low between-study heterogeneity and the absence of type I error verification. The paper is therefore a useful illustration of an existing method rather than a new methodological contribution, and its practical recommendations require additional sensitivity analyses before they can be accepted.","major_comments":[{"comment":"The expected trial duration and recruitment-probability analyses use Duration = n/λ and P(N ≥ n) with n equal to the definitive trial sample size only. The enrollment time of the pilot study itself (at least n_pilot/λ months if run sequentially) is omitted from Figures 2 and 3 and from the headline time savings in Section 4.1. Since the total time from pilot start to definitive completion is the quantity of operational interest in rare disease trials, the reported duration reductions are overstated unless the pilot enrollment period is explicitly included or the analysis is clearly labeled as confirmatory-phase-only.","section":"Section 4 (Duration and recruitment calculations)"},{"comment":"The robustness property that anchors the paper's validity claims is not evaluated by the simulations. The main simulation fixes pC and pT at identical values in the pilot and definitive trials (Section 4), so there is no prior-data conflict by construction. The conflict scenarios in Figure 4 only multiply the pilot risk ratio by a deterministic factor (0.80–0.95 × RR); they do not generate pilot and definitive data from a common hierarchical model with random between-study variability, and they do not report type I error at the ϕ = 0.975 threshold. Consequently, the Discussion's claim that the robust MAP prior 'maintains type I error control even when disagreement exists' is unsupported. The authors should add a sensitivity analysis that varies the implicit low heterogeneity (e.g., a logit-normal random effect between pilot and definitive log-odds, or an explicit prior on the mixture weight w) and report both power and type I error under conflict.","section":"Sections 3.3, 4, and 5"}],"minor_comments":[{"comment":"The pilot study proportion is defined inconsistently: at the start of Section 4 it is 'of the definitive sample size', while in Section 4.1 and Figure 1 it is 'of the total required sample size'. This ambiguity affects the interpretation of the sample-size reductions and should be harmonized.","section":"Section 4"},{"comment":"Typo: 'arising form heterogeneity' should be 'arising from heterogeneity'.","section":"Section 1"},{"comment":"The phrase 'non-informative informative component' in the Discussion is contradictory; it should read 'non-informative component' or 'vague component'.","section":"Section 5"},{"comment":"The Gamma prior λ | λ0 ~ Gamma(2λ0, 2) fixes the coefficient of variation at 1/sqrt(2λ0); the authors do not justify this choice or test sensitivity to it, although the recruitment probability results in Figure 3 depend on it.","section":"Section 4"},{"comment":"The figures report point estimates only; adding variability across simulation replicates would help readers gauge the precision of the duration and recruitment-probability estimates.","section":"Figures 2 and 3"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope for stat.AP and addresses a practically relevant problem. The main issues are not with the derivations (which are standard) but with the gap between the strength of the operational claims and the evidence provided. The requested sensitivity analyses—random between-study heterogeneity, type I error reporting, and inclusion of pilot enrollment time—are feasible within the manuscript's scope and would substantially strengthen the paper. No concerns about novelty disclosure or citation practices beyond the normal expectation that the relationship to Schmidli et al. (2014) and Qi et al. (2022) remains clearly stated."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a competent applied simulation paper, not a new method. The conditional claim—that a protocol-aligned pilot's data, fed through a robust MAP prior, can cut the confirmatory sample size when pilot and definitive are compatible—holds up. The operational claims about time savings and type I error control do not yet hold up. The paper deserves refereeing, but it needs real work before the headline results can be trusted.\n\nWhat is genuinely useful: the framing is practical, the simulation grid is sensible, and the code is promised in supplementary materials. I also appreciate that the authors state openly that the single-pilot setting forces a fixed low heterogeneity assumption that cannot be estimated. That is honest.\n\nWhere it is soft, in order of seriousness:\n1. The duration numbers are computed as n/lambda for the confirmatory phase alone. The pilot phase's own enrollment time is ignored, so the 'shorter expected trial duration' claim is overstated. If the pilot takes 12–24 months, the total timeline may not shrink at all. This needs a proper total-time accounting.\n2. The paper asserts that robust MAP 'maintains type I error control even when disagreement exists,' but no type I error simulation is reported. The conflict scenarios in Figure 4 use a deterministic attenuation of the pilot RR and only report power/sample size. That is not evidence on error control. A random-effects heterogeneity simulation with type I error as the endpoint is needed before the validity claim is credible.\n3. The conflict simulations are not from a hierarchical model; they are fixed shifts. So the 'robustness' of the weight updating is not really stress-tested.\n4. A smaller point: the motivating example seems to flip the direction of the hypothesis—'flare rate' vs 'success probability'—which could confuse readers.\n5. The total patient count (pilot plus definitive) increases relative to no pilot, so the 'efficiency' story is only about new enrollees in the definitive trial. The paper should say this explicitly.\n\nWho should read it: applied statisticians and trial planners working on rare disease feasibility studies. The conditional sample size reduction is a useful, reproducible demonstration. It is not a methodological advance, but it does not need to be. The paper is serious, not sloppy. I would send it to referees, with a request for major revision: add type I error simulations, redo duration with pilot time included, and clarify the endpoint direction.","headline":"A solid conditional result on sample size reduction from pilot data via robust MAP priors, but the duration-savings and type I error claims outrun the simulations.","tokens_in":10776,"tokens_out":5451,"would_cite":false,"duration_ms":61389,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F15","62P10"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that reusing outcome data from a single protocol-aligned pilot feasibility study as a robust Bayesian prior can make confirmatory rare-disease trials smaller, faster, and more likely to complete recruitment without…","keywords":["Bayesian clinical trials","pilot feasibility study","rare diseases","robust meta-analytic-predictive prior","sample size determination","trial duration","recruitment probability","binary outcome"],"falsifier":"Run the same sample-size simulation with the pilot drawn under a true risk ratio 20% lower than the definitive trial's and with the implicit between-study heterogeneity set to a moderate value rather than a low one; if the robust-MAP design then requires as many new enrollees as the no-pilot design (or more) under H1, or rejects H0 at more than the nominal rate under H0, the central efficiency-and-control claim is refuted.","tokens_in":1788,"feed_emoji":"🧪","tokens_out":2856,"duration_ms":98173,"temperature":0.7,"pith_summary":"Pilot feasibility studies in rare diseases are usually treated as operational checkpoints, and their outcome data are discarded at the analysis stage. This paper argues that when the pilot is deliberately designed to mirror the planned confirmatory trial, its binary outcome data can be carried forward as an informative prior using a robust meta-analytic-predictive (MAP) prior. In simulations, incorporating pilot data equal to 10-40% of the definitive sample size reduced the required number of new enrollees by roughly 10-23%, shortened expected trial duration at realistic recruitment rates, and increased the probability of meeting recruitment targets within a fixed window. The mechanism is a mixture prior that borrows from the pilot when the data agree and down-weights it when they conflict, so the efficiency gains are claimed to come without inflating the type I error rate. The practical upshot for rare-disease research is that a small, harmonized pilot can do double duty: feasibility assessment and prior evidence.","feed_headline":"Pilot data can cut rare-disease trial sample sizes by up to 23%","feed_subtitle":"Reusing a protocol-aligned pilot as a Bayesian prior cuts timelines and boosts recruitment-target odds.","key_machinery":"The central mechanism is the robust meta-analytic-predictive (MAP) prior, approximated here as a two-component mixture for each arm's success probability: a vague component $\\operatorname{Beta}(1,1)$ and an informative component $\\operatorname{Beta}(a_0+y_g^{(1)}, b_0+n_g^{(1)}-y_g^{(1)})$ built from the pilot counts, with initial mixture weight $w=0.5$. After the definitive trial observes $y_g^{(2)}$ successes in $n_g^{(2)}$ patients, the posterior for $p_g$ is again a Beta mixture whose weight $\\tilde{w}_g$ is updated by the marginal likelihood under each component, so the pilot's influence shrinks automatically when pilot and definitive data conflict. All updates are closed-form because Beta is conjugate to the binomial likelihood, which makes the design easy to simulate and compute sample sizes for.","core_discovery":"The central claim is that a single pilot study, when protocol-aligned with a confirmatory trial, can be formally incorporated into that trial's analysis through a robust MAP prior, converting early-phase data into sample-size savings and operational feasibility gains. Concretely, the simulations show that for a control event rate of $p_C=0.06$ and risk ratio $1.9$, pilot data equal to 20% of the definitive sample size reduced required new enrollees from 846 to 736 (13%), and 40% pilot data reduced it by 23%; at $p_C=0.25$ the reductions were 10% and 17%, and at $p_C=0.6$ they were 11% and 17%. Expected durations shrank by several months to over a year depending on recruitment rate, and the probability of completing recruitment within a fixed period rose by roughly 0.06 to 0.17. The paper frames this as an ethical as well as operational gain: pilot participants' data are not wasted, and the need to recruit fewer patients in a small population is itself valuable.","pith_inferences":["A testable extension would randomize the between-study heterogeneity parameter in the robust MAP prior and re-run the operating-characteristic simulations; the paper does not do this because a single pilot cannot estimate that parameter, but doing so would show how much the sample-size savings degrade as heterogeneity grows.","The same design logic could be applied to the IVIg taper example with a two-sided or non-inferiority hypothesis, since the motivating trial actually expects a reduction in flare rates rather than the increase simulated here; the paper notes the framework is reparameterizable but does not simulate that orientation.","A more ambitious implication, left implicit, is that funders and ethics committees could condition approval of a feasibility study on protocol alignment with the future confirmatory trial, turning a common operational step into a formal evidence-generating component.","One could build an adaptive version where the mixture weight is re-estimated at an interim analysis and borrowing is dynamically adjusted; the paper mentions this as future work, so it is a natural next step rather than a demonstrated result."],"forward_implications":["A pilot study containing 10-20% of the definitive sample size is enough to produce meaningful reductions in required enrollment, so the added cost of a slightly larger, protocol-aligned pilot can pay for itself in a smaller confirmatory trial.","The time savings are largest in slow-recruiting, low-event-rate settings, which are precisely the rare-disease scenarios where trial viability is most at risk.","The recruitment-probability model, a Gamma-Poisson process with negative binomial marginal, lets planners translate sample-size reductions into concrete probabilities of finishing on time.","Pessimistic pilot results do not erase the benefit in most settings: at $p_C=0.06$ and $0.25$, even a pilot risk ratio of $0.8$ times the true $RR$ still yields sample sizes below the no-pilot benchmark, though at $p_C=0.6$ a pessimistic pilot can push sample size slightly higher.","The robust MAP framework extends beyond binary outcomes to any conjugate exponential-family setting and to time-to-event endpoints via MCMC, so the same pilot-reuse logic applies to other rare-disease designs."],"supporting_citations":[{"why":"Supplies the robust MAP prior construction: the mixture of an informative historical-data component and a vague component with marginal-likelihood weight updating that the paper adapts to a single pilot study.","marker":"Schmidli et al. (2014)"},{"why":"Establishes that MAP priors can substantially reduce required sample size without inflating type I error when historical and current data are commensurate, the result this paper extends to pilot data and operational endpoints.","marker":"Qi et al. (2022)"},{"why":"Provides the caution against pooling pilot and main-trial data directly, motivating the use of calibrated Bayesian borrowing instead of simple combination.","marker":"Leon et al. (2011)"},{"why":"Original source of the recruitment prediction model, a Poisson process with Gamma-distributed rate, used to compute the probability of meeting recruitment targets.","marker":"Anisimov and Fedorov (2007)"},{"why":"Recent Bayesian adaptive feasibility design whose recruitment-target probability approach the paper follows for estimating the chance of completing enrollment on time.","marker":"Churipuy et al. (2024)"}],"fun_headline_variants":["Pilot data can cut rare-disease trial sample sizes 23%","Bayesian pilot priors shrink rare-disease trial size up to 23%","Rare-disease trials save 23% sample size with pilot data reuse","Pilot reuse slashes rare-disease trial enrollment by 23%","Bayesian pilot data trims rare-disease trial sample needs 23%"],"cache_read_input_tokens":12928,"weakest_assumption_plain":"The simulations generate pilot and definitive data from the same true success probabilities and assume only a fixed, low between-study heterogeneity, so the pilot prior is unbiased by construction; if real pilot-to-definitive discordance or heterogeneity is larger, the reported sample-size savings and type I error control would shrink.","fun_headline_variants_meta":{"raw":{"variants":["Pilot data can cut rare-disease trial sample sizes 23%","Bayesian pilot priors shrink rare-disease trial size up to 23%","Rare-disease trials save 23% sample size with pilot data reuse","Pilot reuse slashes rare-disease trial enrollment by 23%","Bayesian pilot data trims rare-disease trial sample needs 23%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000206,"raw_usage":{"total_tokens":1389,"prompt_tokens":932,"completion_tokens":457,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":548,"completion_tokens_details":{"reasoning_tokens":357}},"tokens_in":548,"tokens_out":457,"duration_ms":4832,"temperature":1.0,"reasoning_tokens":357,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T17:36:01.922489+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same sample-size simulation with the pilot drawn under a true risk ratio 20% lower than the definitive trial's and with the implicit between-study heterogeneity set to a moderate value rather than a low one; if the robust-MAP design then requires as many new enrollees as the no-pilot design (or more) under H1, or rejects H0 at more than the nominal rate under H0, the central efficiency-and-control claim is refuted.","supporting_citations":[{"cited_title":"C., Davis, L","cited_arxiv_id":null,"evidence_quote":"Provides the caution against pooling pilot and main-trial data directly, motivating the use of calibrated Bayesian borrowing instead of simple combination."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Original source of the recruitment prediction model, a Poisson process with Gamma-distributed rate, used to compute the probability of meeting recruitment targets."},{"cited_title":"M., Golchi, S., Hudson, M., and Hoa, S","cited_arxiv_id":null,"evidence_quote":"Recent Bayesian adaptive feasibility design whose recruitment-target probability approach the paper follows for estimating the chance of completing enrollment on time."}],"review_version":1}