REVIEW 2 major objections 5 minor 24 references
The efficiencies of pilot feasibility trials in rare diseases using Bayesian methods
T0 review · 2 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read This paper claims that reusing outcome data from a single protocol-aligned pilot feasibility study as a robust Bayesian prior can make confirmatory rare-disease trials smaller, faster, and more likely to complete recruitment without…
desk verdict A solid conditional result on sample size reduction from pilot data via robust MAP priors, but the duration-savings and type I error claims outrun the simulations. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is the robust meta-analytic-predictive (MAP) prior, approximated here as a two-component mixture for each arm's success probability: a vague component $\operatorname{Beta}(1,1)$ and an informative component $\operatorname{Beta}(a_0+y_g^{(1)}, b_0+n_g^{(1)}-y_g^{(1)})$ built from the pilot counts, with initial mixture weight $w=0.5$. After the definitive trial observes $y_g^{(2)}$ successes in $n_g^{(2)}$ patients, the posterior for $p_g$ is again a Beta mixture whose weight $\tilde{w}_g$ is updated by the marginal likelihood under each component, so the pilot's influence shrinks automatically when pilot and definitive data conflict. All updates are closed-form because Beta is conjugate to the binomial likelihood, which makes the design easy to simulate and compute sample sizes for.
What would settle it
Run the same sample-size simulation with the pilot drawn under a true risk ratio 20% lower than the definitive trial's and with the implicit between-study heterogeneity set to a moderate value rather than a low one; if the robust-MAP design then requires as many new enrollees as the no-pilot design (or more) under H1, or rejects H0 at more than the nominal rate under H0, the central efficiency-and-control claim is refuted.
Extended reading notes
Core claim
The central claim is that a single pilot study, when protocol-aligned with a confirmatory trial, can be formally incorporated into that trial's analysis through a robust MAP prior, converting early-phase data into sample-size savings and operational feasibility gains. Concretely, the simulations show that for a control event rate of $p_C=0.06$ and risk ratio $1.9$, pilot data equal to 20% of the definitive sample size reduced required new enrollees from 846 to 736 (13%), and 40% pilot data reduced it by 23%; at $p_C=0.25$ the reductions were 10% and 17%, and at $p_C=0.6$ they were 11% and 17%. Expected durations shrank by several months to over a year depending on recruitment rate, and the probability of completing recruitment within a fixed period rose by roughly 0.06 to 0.17. The paper frames this as an ethical as well as operational gain: pilot participants' data are not wasted, and the need to recruit fewer patients in a small population is itself valuable.
Load-bearing premise
The simulations generate pilot and definitive data from the same true success probabilities and assume only a fixed, low between-study heterogeneity, so the pilot prior is unbiased by construction; if real pilot-to-definitive discordance or heterogeneity is larger, the reported sample-size savings and type I error control would shrink.
Editorial extensions
If this is right
- A pilot study containing 10-20% of the definitive sample size is enough to produce meaningful reductions in required enrollment, so the added cost of a slightly larger, protocol-aligned pilot can pay for itself in a smaller confirmatory trial.
- The time savings are largest in slow-recruiting, low-event-rate settings, which are precisely the rare-disease scenarios where trial viability is most at risk.
- The recruitment-probability model, a Gamma-Poisson process with negative binomial marginal, lets planners translate sample-size reductions into concrete probabilities of finishing on time.
- Pessimistic pilot results do not erase the benefit in most settings: at $p_C=0.06$ and $0.25$, even a pilot risk ratio of $0.8$ times the true $RR$ still yields sample sizes below the no-pilot benchmark, though at $p_C=0.6$ a pessimistic pilot can push sample size slightly higher.
- The robust MAP framework extends beyond binary outcomes to any conjugate exponential-family setting and to time-to-event endpoints via MCMC, so the same pilot-reuse logic applies to other rare-disease designs.
Reading between the lines
- A testable extension would randomize the between-study heterogeneity parameter in the robust MAP prior and re-run the operating-characteristic simulations; the paper does not do this because a single pilot cannot estimate that parameter, but doing so would show how much the sample-size savings degrade as heterogeneity grows.
- The same design logic could be applied to the IVIg taper example with a two-sided or non-inferiority hypothesis, since the motivating trial actually expects a reduction in flare rates rather than the increase simulated here; the paper notes the framework is reparameterizable but does not simulate that orientation.
- A more ambitious implication, left implicit, is that funders and ethics committees could condition approval of a feasibility study on protocol alignment with the future confirmatory trial, turning a common operational step into a formal evidence-generating component.
- One could build an adaptive version where the mixture weight is re-estimated at an interim analysis and borrowing is dynamically adjusted; the paper mentions this as future work, so it is a natural next step rather than a demonstrated result.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes using robust meta-analytic-predictive (MAP) priors to incorporate data from a single protocol-aligned pilot feasibility study into the design and analysis of a definitive confirmatory trial for a binary efficacy outcome in rare diseases. The authors specify a two-component mixture prior for each arm, with an informative Beta component updated from pilot data and a vague Beta(1,1) component, and an initial weight w=0.5. Through simulations, they claim that including pilot data at 10-40% of the definitive sample size reduces the required confirmatory sample size, shortens expected trial duration, and increases the probability of meeting recruitment targets. A secondary simulation applies deterministic attenuation factors to the pilot risk ratio to assess prior-data conflict. The paper is motivated by an IVIg de-escalation feasibility trial in autoimmune inflammatory myopathies.
Significance. If the central claims are fully supported, the paper would offer practical guidance for a common and important problem: leveraging pilot data to make rare disease confirmatory trials smaller and faster without compromising validity. The methodological core is standard—the posterior mixture formulas follow Schmidli et al. (2014) and are correctly derived—and the simulation code is provided in the supplementary materials, which is a strength. However, the operational conclusions (sample size, duration, recruitment probability) rest on assumptions that are not adequately tested, particularly the fixed low between-study heterogeneity and the absence of type I error verification. The paper is therefore a useful illustration of an existing method rather than a new methodological contribution, and its practical recommendations require additional sensitivity analyses before they can be accepted.
major comments (2)
- [Section 4 (Duration and recruitment calculations)] The expected trial duration and recruitment-probability analyses use Duration = n/λ and P(N ≥ n) with n equal to the definitive trial sample size only. The enrollment time of the pilot study itself (at least n_pilot/λ months if run sequentially) is omitted from Figures 2 and 3 and from the headline time savings in Section 4.1. Since the total time from pilot start to definitive completion is the quantity of operational interest in rare disease trials, the reported duration reductions are overstated unless the pilot enrollment period is explicitly included or the analysis is clearly labeled as confirmatory-phase-only.
- [Sections 3.3, 4, and 5] The robustness property that anchors the paper's validity claims is not evaluated by the simulations. The main simulation fixes pC and pT at identical values in the pilot and definitive trials (Section 4), so there is no prior-data conflict by construction. The conflict scenarios in Figure 4 only multiply the pilot risk ratio by a deterministic factor (0.80–0.95 × RR); they do not generate pilot and definitive data from a common hierarchical model with random between-study variability, and they do not report type I error at the ϕ = 0.975 threshold. Consequently, the Discussion's claim that the robust MAP prior 'maintains type I error control even when disagreement exists' is unsupported. The authors should add a sensitivity analysis that varies the implicit low heterogeneity (e.g., a logit-normal random effect between pilot and definitive log-odds, or an explicit prior on the mixture weight w) and report both power and type I error under conflict.
minor comments (5)
- [Section 4] The pilot study proportion is defined inconsistently: at the start of Section 4 it is 'of the definitive sample size', while in Section 4.1 and Figure 1 it is 'of the total required sample size'. This ambiguity affects the interpretation of the sample-size reductions and should be harmonized.
- [Section 1] Typo: 'arising form heterogeneity' should be 'arising from heterogeneity'.
- [Section 5] The phrase 'non-informative informative component' in the Discussion is contradictory; it should read 'non-informative component' or 'vague component'.
- [Section 4] The Gamma prior λ | λ0 ~ Gamma(2λ0, 2) fixes the coefficient of variation at 1/sqrt(2λ0); the authors do not justify this choice or test sensitivity to it, although the recruitment probability results in Figure 3 depend on it.
- [Figures 2 and 3] The figures report point estimates only; adding variability across simulation replicates would help readers gauge the precision of the duration and recruitment-probability estimates.
Circularity Check
No significant circularity: the pilot-prior borrowing and sample-size reductions are simulated under openly stated exchangeability assumptions, not derived from the target result.
full rationale
The paper's central claim is that a protocol-aligned pilot study can be used, via a robust MAP prior, to reduce confirmatory sample size and improve operational feasibility. The derivation chain is not circular: the mixture prior in Eqs. (2)-(3) follows Schmidli et al. (2014) with fixed constants (w=0.5, a0=b0=1); the posterior update in Eq. (5) is a standard conjugate calculation; and the sample-size, duration, and recruitment results in Section 4 are Monte Carlo operating characteristics under stated assumptions, not fitted outputs. The no-prior-data-conflict setting is explicitly declared, conflict scenarios are simulated separately (Section 4, Figure 4), and limitations such as the unidentifiable between-study heterogeneity are openly acknowledged in the Discussion. The one self-citation (Churipuy et al. 2024, for recruitment probability methods) is not load-bearing because the paper invokes the original external method (Anisimov and Fedorov 2007) and implements the calculations itself. Any concerns about the fixed-low-heterogeneity assumption or missing type I error verification are correctness/robustness issues, not circularity.
Assumptions & free parameters
free parameters (4)
- Initial prior weight w =
0.5
- Posterior decision threshold phi =
0.975
- Vague component hyperparameters a0, b0 =
1, 1
- Recruitment Gamma prior parameters =
Gamma(2*lambda0, 2)
assumptions (4)
- domain assumption Pilot and definitive trial data are exchangeable, with an implicit fixed low between-study heterogeneity.
- ad hoc to paper Simulation generates pilot and definitive outcomes from identical event probabilities in the main analysis, so there is no prior-data conflict by construction.
- domain assumption Recruitment follows a Poisson process with a Gamma-distributed rate, Gamma(2*lambda0, 2).
- standard math Standard beta-binomial conjugacy and marginal-likelihood weight updating.
Cite this review
Pith. "Pith review of The efficiencies of pilot feasibility trials in rare diseases using Bayesian methods." pith.science (2026). https://pith.science/paper/XMSNBOUN
@misc{pith2026250710269,
author = {Pith},
title = {Pith review of: The efficiencies of pilot feasibility trials in rare diseases using Bayesian methods},
year = {2026},
howpublished = {\url{https://pith.science/paper/XMSNBOUN}},
note = {Machine review of arXiv:2507.10269}
}
read the original abstract
Pilot feasibility studies play a pivotal role in the development of clinical trials for rare diseases, where small populations and slow recruitment often threaten trial viability. While such studies are commonly used to assess operational parameters, they also offer a valuable opportunity to inform the design and analysis of subsequent definitive trials-particularly through the use of Bayesian methods. In this paper, we demonstrate how data from a single, protocol-aligned pilot study can be incorporated into a definitive trial using robust meta-analytic-predictive priors. We focus on the case of a binary efficacy outcome, motivated by a feasibility trial of intravenous immunoglobulin tapering in autoimmune inflammatory myopathies. Through simulation studies, we evaluate the operating characteristics of trials informed by pilot data, including sample size, expected trial duration, and the probability of meeting recruitment targets. Our findings highlight the operational and ethical advantages of leveraging pilot data via robust Bayesian priors, and offer practical guidance for their application in rare disease settings.
Figures
Reference graph
Works this paper leans on
-
[1]
M., Griger, Z., Moiseev, S., Oddis, C., Schiopu, E., Vencovsk \`y , J., et al
Aggarwal, R., Charles-Schoeman, C., Schessl, J., Bata-Cs \"o rg o , Z., Dimachkie, M. M., Griger, Z., Moiseev, S., Oddis, C., Schiopu, E., Vencovsk \`y , J., et al. (2022). Trial of intravenous immune globulin in dermatomyositis. New England Journal of Medicine , 387(14):1264--1278
work page 2022
-
[2]
Anisimov, V. V. and Fedorov, V. V. (2007). Modelling, prediction and adaptive adjustment of recruitment in multicentre trials. Statistics in Medicine , 26(27):4958--4975
work page 2007
-
[3]
Arain, M., Campbell, M. J., Cooper, C. L., and Lancaster, G. A. (2010). What is a pilot or feasibility study? A review of current practice and editorial policy. BMC Medical Research Methodology , 10(1):67
work page 2010
-
[4]
Balcome, S., Musgrove, D., Haddad, T., and Hickey, G. L. (2021). bayesDP: Tools for the Bayesian Discount Prior Function . R package version 1.3.4
work page 2021
-
[5]
Bi, D., Liu, M., Lin, J., and Liu, R. (2023). Beats: B ayesian hybrid design with flexible sample size adaptation for time-to-event endpoints. Statistics in Medicine , 42(30):5708--5722
work page 2023
-
[6]
Chandereng, T., Musgrove, D., Haddad, T., Hickey, G., Hanson, T., and Lystig, T. (2020). bayesCT: Simulation and Analysis of Adaptive Bayesian Clinical Trials . R package version 0.99.3
work page 2020
-
[7]
M., Golchi, S., Hudson, M., and Hoa, S
Churipuy, M. M., Golchi, S., Hudson, M., and Hoa, S. (2024). A B ayesian adaptive feasibility design for rare diseases. Contemporary Clinical Trials Communications , 42:101392
work page 2024
-
[8]
Duan, Y., Ye, K., and Smith, E. P. (2006). Evaluating water quality using power priors to incorporate historical information. Environmetrics: The Official Journal of the International Environmetrics Society , 17(1):95--106
work page 2006
Show all 24 references
-
[9]
Eggleston, B., Wilson, D., McNeil, B., Ibrahim, J., and Catellier, D. (2019). BayesCTDesign: Two Arm Bayesian Clinical Trial Design with and Without Historical Control Data . R package version 0.6.0
2019
-
[10]
Fleming, T. R. (2010). Clinical trials: discerning hype from substance. Annals of Internal Medicine , 153(6):400--406
2010
-
[11]
C., Batshaw, M., Dunkle, M., Gopal-Srivastava, R., Kaye, E., Krischer, J., Nguyen, T., Paulus, K., Merkel, P
Griggs, R. C., Batshaw, M., Dunkle, M., Gopal-Srivastava, R., Kaye, E., Krischer, J., Nguyen, T., Paulus, K., Merkel, P. A., et al. (2009). Clinical research for rare disease: opportunities, challenges, and solutions. Molecular Genetics and Metabolism , 96(1):20--26
2009
-
[12]
P., Carlin, B
Hobbs, B. P., Carlin, B. P., Mandrekar, S. J., and Sargent, D. J. (2011). Hierarchical commensurate and power prior models for adaptive incorporation of historical information in clinical trials. Biometrics , 67(3):1047--1056
2011
-
[13]
Ibrahim, J. G. and Chen, M.-H. (2000). Power prior distributions for regression models. Statistical Science , pages 46--60
2000
-
[14]
Utilisation des immunoglobulines non spécifiques intraveineuses et sous-cutanées au québec, 2020-2021
Institut National de Santé Publique du Québec (2023). Utilisation des immunoglobulines non spécifiques intraveineuses et sous-cutanées au québec, 2020-2021. Report, Institut National de Santé Publique du Québec. Available from: https://www.inspq.qc.ca/biovigilance
2023
-
[15]
C., Davis, L
Leon, A. C., Davis, L. L., and Kraemer, H. C. (2011). The role and interpretation of pilot studies in clinical research. Journal of Psychiatric Research , 45(5):626--629
2011
-
[16]
E., Saris, C
Lim, J., Eftimov, F., Verhamme, C., Brusse, E., Hoogendijk, J. E., Saris, C. G., Raaphorst, J., De Haan, R. J., van Schaik, I. N., Aronica, E., et al. (2021). Intravenous immunoglobulins as first-line treatment in idiopathic inflammatory myopathies: A pilot study. Rheumatology...
2021
-
[17]
Pan, H., Yuan, Y., and Xia, J. (2017). A calibrated power prior approach to borrow information from historical data with application to biosimilar clinical trials. Journal of the Royal Statistical Society Series C: Applied Statistics , 66(5):979--996
2017
-
[18]
Psioda, M. A. and Ibrahim, J. G. (2019). B ayesian clinical trial design using historical data that inform the treatment effect. Biostatistics , 20(3):400--415
2019
-
[19]
V., and Wason, J
Qi, Y., Hampson, L. V., and Wason, J. M. S. (2022). Sample size calculation for clinical trials using meta-analytic-predictive priors. Biostatistics , 23(1):48--63
2022
-
[20]
Schmidli, H., Gsteiger, S., Roychoudhury, S., O'Hagan, A., Spiegelhalter, D., and Neuenschwander, B. (2014). Robust meta-analytic-predictive priors in clinical trials with historical control information. Biometrics , 70(4):1023--1032
2014
-
[21]
A., and Ibrahim, J
Shen, Y., Psioda, M. A., and Ibrahim, J. G. (2021). BayesPPD : An R package for bayesian sample size determination using the power and normalized power prior for generalized linear models. arXiv preprint arXiv:2112.14616
2021 arXiv
-
[22]
P., Robson, R., Thabane, M., Giangregorio, L., and Goldsmith, C
Thabane, L., Ma, J., Chu, R., Cheng, J., Ismaila, A., Rios, L. P., Robson, R., Thabane, M., Giangregorio, L., and Goldsmith, C. H. (2010). A tutorial on pilot studies: T he what, why and how. BMC Medical Research Methodology , 10(1):1--10
2010
-
[23]
Food and Drug Administration (2010)
U.S. Food and Drug Administration (2010). Guidance for the use of B ayesian statistics in medical device clinical trials. https://www.fda.gov/regulatory-information/search-fda-guidance-documents/guidance-use-bayesian-statistics-medical-device-clinical-trials. Accessed: 2025-05-05
2010
-
[24]
Weber, S. (2021). RBesT : R Bayesian Evidence Synthesis Tools . R package version 1.6-2
2021
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.