Pith. sign in

REVIEW 2 major objections 5 minor 24 references

The efficiencies of pilot feasibility trials in rare diseases using Bayesian methods

T0 review · 2 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read This paper claims that reusing outcome data from a single protocol-aligned pilot feasibility study as a robust Bayesian prior can make confirmatory rare-disease trials smaller, faster, and more likely to complete recruitment without…

desk verdict A solid conditional result on sample size reduction from pilot data via robust MAP priors, but the duration-savings and type I error claims outrun the simulations. read the letter →

arxiv 2507.10269 v1 pith:XMSNBOUN submitted 2025-07-14 stat.AP

classification stat.AP MSC 62F1562P10
keywords Bayesianclinicaltrialspilotfeasibilitystudyrarediseasesrobustmeta-analytic-predictivepriorsamplesizedeterminationtrialdurationrecruitmentprobabilitybinaryoutcome
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Pilot feasibility studies in rare diseases are usually treated as operational checkpoints, and their outcome data are discarded at the analysis stage. This paper argues that when the pilot is deliberately designed to mirror the planned confirmatory trial, its binary outcome data can be carried forward as an informative prior using a robust meta-analytic-predictive (MAP) prior. In simulations, incorporating pilot data equal to 10-40% of the definitive sample size reduced the required number of new enrollees by roughly 10-23%, shortened expected trial duration at realistic recruitment rates, and increased the probability of meeting recruitment targets within a fixed window. The mechanism is a mixture prior that borrows from the pilot when the data agree and down-weights it when they conflict, so the efficiency gains are claimed to come without inflating the type I error rate. The practical upshot for rare-disease research is that a small, harmonized pilot can do double duty: feasibility assessment and prior evidence.

What carries the argument

The central mechanism is the robust meta-analytic-predictive (MAP) prior, approximated here as a two-component mixture for each arm's success probability: a vague component $\operatorname{Beta}(1,1)$ and an informative component $\operatorname{Beta}(a_0+y_g^{(1)}, b_0+n_g^{(1)}-y_g^{(1)})$ built from the pilot counts, with initial mixture weight $w=0.5$. After the definitive trial observes $y_g^{(2)}$ successes in $n_g^{(2)}$ patients, the posterior for $p_g$ is again a Beta mixture whose weight $\tilde{w}_g$ is updated by the marginal likelihood under each component, so the pilot's influence shrinks automatically when pilot and definitive data conflict. All updates are closed-form because Beta is conjugate to the binomial likelihood, which makes the design easy to simulate and compute sample sizes for.

What would settle it

Run the same sample-size simulation with the pilot drawn under a true risk ratio 20% lower than the definitive trial's and with the implicit between-study heterogeneity set to a moderate value rather than a low one; if the robust-MAP design then requires as many new enrollees as the no-pilot design (or more) under H1, or rejects H0 at more than the nominal rate under H0, the central efficiency-and-control claim is refuted.

Watch

Extended reading notes

Core claim

The central claim is that a single pilot study, when protocol-aligned with a confirmatory trial, can be formally incorporated into that trial's analysis through a robust MAP prior, converting early-phase data into sample-size savings and operational feasibility gains. Concretely, the simulations show that for a control event rate of $p_C=0.06$ and risk ratio $1.9$, pilot data equal to 20% of the definitive sample size reduced required new enrollees from 846 to 736 (13%), and 40% pilot data reduced it by 23%; at $p_C=0.25$ the reductions were 10% and 17%, and at $p_C=0.6$ they were 11% and 17%. Expected durations shrank by several months to over a year depending on recruitment rate, and the probability of completing recruitment within a fixed period rose by roughly 0.06 to 0.17. The paper frames this as an ethical as well as operational gain: pilot participants' data are not wasted, and the need to recruit fewer patients in a small population is itself valuable.

Load-bearing premise

The simulations generate pilot and definitive data from the same true success probabilities and assume only a fixed, low between-study heterogeneity, so the pilot prior is unbiased by construction; if real pilot-to-definitive discordance or heterogeneity is larger, the reported sample-size savings and type I error control would shrink.

Editorial extensions

If this is right

  • A pilot study containing 10-20% of the definitive sample size is enough to produce meaningful reductions in required enrollment, so the added cost of a slightly larger, protocol-aligned pilot can pay for itself in a smaller confirmatory trial.
  • The time savings are largest in slow-recruiting, low-event-rate settings, which are precisely the rare-disease scenarios where trial viability is most at risk.
  • The recruitment-probability model, a Gamma-Poisson process with negative binomial marginal, lets planners translate sample-size reductions into concrete probabilities of finishing on time.
  • Pessimistic pilot results do not erase the benefit in most settings: at $p_C=0.06$ and $0.25$, even a pilot risk ratio of $0.8$ times the true $RR$ still yields sample sizes below the no-pilot benchmark, though at $p_C=0.6$ a pessimistic pilot can push sample size slightly higher.
  • The robust MAP framework extends beyond binary outcomes to any conjugate exponential-family setting and to time-to-event endpoints via MCMC, so the same pilot-reuse logic applies to other rare-disease designs.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A testable extension would randomize the between-study heterogeneity parameter in the robust MAP prior and re-run the operating-characteristic simulations; the paper does not do this because a single pilot cannot estimate that parameter, but doing so would show how much the sample-size savings degrade as heterogeneity grows.
  • The same design logic could be applied to the IVIg taper example with a two-sided or non-inferiority hypothesis, since the motivating trial actually expects a reduction in flare rates rather than the increase simulated here; the paper notes the framework is reparameterizable but does not simulate that orientation.
  • A more ambitious implication, left implicit, is that funders and ethics committees could condition approval of a feasibility study on protocol alignment with the future confirmatory trial, turning a common operational step into a formal evidence-generating component.
  • One could build an adaptive version where the mixture weight is re-estimated at an interim analysis and borrowing is dynamically adjusted; the paper mentions this as future work, so it is a natural next step rather than a demonstrated result.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper proposes using robust meta-analytic-predictive (MAP) priors to incorporate data from a single protocol-aligned pilot feasibility study into the design and analysis of a definitive confirmatory trial for a binary efficacy outcome in rare diseases. The authors specify a two-component mixture prior for each arm, with an informative Beta component updated from pilot data and a vague Beta(1,1) component, and an initial weight w=0.5. Through simulations, they claim that including pilot data at 10-40% of the definitive sample size reduces the required confirmatory sample size, shortens expected trial duration, and increases the probability of meeting recruitment targets. A secondary simulation applies deterministic attenuation factors to the pilot risk ratio to assess prior-data conflict. The paper is motivated by an IVIg de-escalation feasibility trial in autoimmune inflammatory myopathies.

Significance. If the central claims are fully supported, the paper would offer practical guidance for a common and important problem: leveraging pilot data to make rare disease confirmatory trials smaller and faster without compromising validity. The methodological core is standard—the posterior mixture formulas follow Schmidli et al. (2014) and are correctly derived—and the simulation code is provided in the supplementary materials, which is a strength. However, the operational conclusions (sample size, duration, recruitment probability) rest on assumptions that are not adequately tested, particularly the fixed low between-study heterogeneity and the absence of type I error verification. The paper is therefore a useful illustration of an existing method rather than a new methodological contribution, and its practical recommendations require additional sensitivity analyses before they can be accepted.

major comments (2)
  1. [Section 4 (Duration and recruitment calculations)] The expected trial duration and recruitment-probability analyses use Duration = n/λ and P(N ≥ n) with n equal to the definitive trial sample size only. The enrollment time of the pilot study itself (at least n_pilot/λ months if run sequentially) is omitted from Figures 2 and 3 and from the headline time savings in Section 4.1. Since the total time from pilot start to definitive completion is the quantity of operational interest in rare disease trials, the reported duration reductions are overstated unless the pilot enrollment period is explicitly included or the analysis is clearly labeled as confirmatory-phase-only.
  2. [Sections 3.3, 4, and 5] The robustness property that anchors the paper's validity claims is not evaluated by the simulations. The main simulation fixes pC and pT at identical values in the pilot and definitive trials (Section 4), so there is no prior-data conflict by construction. The conflict scenarios in Figure 4 only multiply the pilot risk ratio by a deterministic factor (0.80–0.95 × RR); they do not generate pilot and definitive data from a common hierarchical model with random between-study variability, and they do not report type I error at the ϕ = 0.975 threshold. Consequently, the Discussion's claim that the robust MAP prior 'maintains type I error control even when disagreement exists' is unsupported. The authors should add a sensitivity analysis that varies the implicit low heterogeneity (e.g., a logit-normal random effect between pilot and definitive log-odds, or an explicit prior on the mixture weight w) and report both power and type I error under conflict.
minor comments (5)
  1. [Section 4] The pilot study proportion is defined inconsistently: at the start of Section 4 it is 'of the definitive sample size', while in Section 4.1 and Figure 1 it is 'of the total required sample size'. This ambiguity affects the interpretation of the sample-size reductions and should be harmonized.
  2. [Section 1] Typo: 'arising form heterogeneity' should be 'arising from heterogeneity'.
  3. [Section 5] The phrase 'non-informative informative component' in the Discussion is contradictory; it should read 'non-informative component' or 'vague component'.
  4. [Section 4] The Gamma prior λ | λ0 ~ Gamma(2λ0, 2) fixes the coefficient of variation at 1/sqrt(2λ0); the authors do not justify this choice or test sensitivity to it, although the recruitment probability results in Figure 3 depend on it.
  5. [Figures 2 and 3] The figures report point estimates only; adding variability across simulation replicates would help readers gauge the precision of the duration and recruitment-probability estimates.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the pilot-prior borrowing and sample-size reductions are simulated under openly stated exchangeability assumptions, not derived from the target result.

full rationale

The paper's central claim is that a protocol-aligned pilot study can be used, via a robust MAP prior, to reduce confirmatory sample size and improve operational feasibility. The derivation chain is not circular: the mixture prior in Eqs. (2)-(3) follows Schmidli et al. (2014) with fixed constants (w=0.5, a0=b0=1); the posterior update in Eq. (5) is a standard conjugate calculation; and the sample-size, duration, and recruitment results in Section 4 are Monte Carlo operating characteristics under stated assumptions, not fitted outputs. The no-prior-data-conflict setting is explicitly declared, conflict scenarios are simulated separately (Section 4, Figure 4), and limitations such as the unidentifiable between-study heterogeneity are openly acknowledged in the Discussion. The one self-citation (Churipuy et al. 2024, for recruitment probability methods) is not load-bearing because the paper invokes the original external method (Anisimov and Fedorov 2007) and implements the calculations itself. Any concerns about the fixed-low-heterogeneity assumption or missing type I error verification are correctness/robustness issues, not circularity.

Assumptions & free parameters 4 free parameters · 4 assumptions · 0 invented entities

No new scientific entities are introduced; the analysis builds on existing Bayesian machinery. The main burden is the compatibility and recruitment-modeling assumptions, which are either stated openly or hard-coded into the simulation design.

free parameters (4)
  • Initial prior weight w = 0.5
    Weight placed on the pilot-derived informative component in the robust MAP mixture; fixed by default rather than calibrated, and the magnitude of sample-size reduction depends on it.
  • Posterior decision threshold phi = 0.975
    Threshold for declaring treatment superiority; stated to be chosen to control type I error, but no type I error simulation is reported.
  • Vague component hyperparameters a0, b0 = 1, 1
    Weakly informative Beta(1,1) prior component used in the robust MAP mixture; a conventional choice rather than an empirical fit.
  • Recruitment Gamma prior parameters = Gamma(2*lambda0, 2)
    Shape and rate chosen to reflect moderate confidence around the pilot recruitment rate; variance is lambda0/2 by construction.
assumptions (4)
  • domain assumption Pilot and definitive trial data are exchangeable, with an implicit fixed low between-study heterogeneity.
    Section 3.3 and Discussion state this assumption because a single pilot cannot estimate heterogeneity; the main simulations fix pC and pT equal across phases.
  • ad hoc to paper Simulation generates pilot and definitive outcomes from identical event probabilities in the main analysis, so there is no prior-data conflict by construction.
    Section 4: 'These values were fixed across both the pilot and definitive trials, representing a setting with no prior-data conflict'.
  • domain assumption Recruitment follows a Poisson process with a Gamma-distributed rate, Gamma(2*lambda0, 2).
    Section 4, Equations (6)-(8); used to compute the probability of meeting recruitment targets, and the prior parameters are chosen by hand.
  • standard math Standard beta-binomial conjugacy and marginal-likelihood weight updating.
    Equations (2)-(5) rely on standard conjugate Bayesian updating.

how reviews work

0 comments
Cite this review

Pith. "Pith review of The efficiencies of pilot feasibility trials in rare diseases using Bayesian methods." pith.science (2026). https://pith.science/paper/XMSNBOUN

@misc{pith2026250710269,
  author       = {Pith},
  title        = {Pith review of: The efficiencies of pilot feasibility trials in rare diseases using Bayesian methods},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/XMSNBOUN}},
  note         = {Machine review of arXiv:2507.10269}
}
read the original abstract

Pilot feasibility studies play a pivotal role in the development of clinical trials for rare diseases, where small populations and slow recruitment often threaten trial viability. While such studies are commonly used to assess operational parameters, they also offer a valuable opportunity to inform the design and analysis of subsequent definitive trials-particularly through the use of Bayesian methods. In this paper, we demonstrate how data from a single, protocol-aligned pilot study can be incorporated into a definitive trial using robust meta-analytic-predictive priors. We focus on the case of a binary efficacy outcome, motivated by a feasibility trial of intravenous immunoglobulin tapering in autoimmune inflammatory myopathies. Through simulation studies, we evaluate the operating characteristics of trials informed by pilot data, including sample size, expected trial duration, and the probability of meeting recruitment targets. Our findings highlight the operational and ethical advantages of leveraging pilot data via robust Bayesian priors, and offer practical guidance for their application in rare disease settings.

Figures

Figures reproduced from arXiv: 2507.10269 by the authors.

Figure 1
Figure 1. Required total sample size to achieve 80% power across different pilot study [PITH_FULL_IMAGE:figures/full_fig_p008_1.png] view at source ↗
Figure 2
Figure 2. Expected trial duration (in months) as a function of pilot data proportion and [PITH_FULL_IMAGE:figures/full_fig_p010_2.png] view at source ↗
Figure 3
Figure 3. Probability of meeting recruitment targets within a fixed trial duration, using [PITH_FULL_IMAGE:figures/full_fig_p011_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Estimated definitive trial sample sizes required for 80% power under varying [PITH_FULL_IMAGE:figures/full_fig_p012_4.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

24 extracted references · 22 canonical work pages

  1. [1]

    M., Griger, Z., Moiseev, S., Oddis, C., Schiopu, E., Vencovsk \`y , J., et al

    Aggarwal, R., Charles-Schoeman, C., Schessl, J., Bata-Cs \"o rg o , Z., Dimachkie, M. M., Griger, Z., Moiseev, S., Oddis, C., Schiopu, E., Vencovsk \`y , J., et al. (2022). Trial of intravenous immune globulin in dermatomyositis. New England Journal of Medicine , 387(14):1264--1278

  2. [2]

    Anisimov, V. V. and Fedorov, V. V. (2007). Modelling, prediction and adaptive adjustment of recruitment in multicentre trials. Statistics in Medicine , 26(27):4958--4975

  3. [3]

    J., Cooper, C

    Arain, M., Campbell, M. J., Cooper, C. L., and Lancaster, G. A. (2010). What is a pilot or feasibility study? A review of current practice and editorial policy. BMC Medical Research Methodology , 10(1):67

  4. [4]

    Balcome, S., Musgrove, D., Haddad, T., and Hickey, G. L. (2021). bayesDP: Tools for the Bayesian Discount Prior Function . R package version 1.3.4

  5. [5]

    Bi, D., Liu, M., Lin, J., and Liu, R. (2023). Beats: B ayesian hybrid design with flexible sample size adaptation for time-to-event endpoints. Statistics in Medicine , 42(30):5708--5722

  6. [6]

    Chandereng, T., Musgrove, D., Haddad, T., Hickey, G., Hanson, T., and Lystig, T. (2020). bayesCT: Simulation and Analysis of Adaptive Bayesian Clinical Trials . R package version 0.99.3

  7. [7]

    M., Golchi, S., Hudson, M., and Hoa, S

    Churipuy, M. M., Golchi, S., Hudson, M., and Hoa, S. (2024). A B ayesian adaptive feasibility design for rare diseases. Contemporary Clinical Trials Communications , 42:101392

  8. [8]

    Duan, Y., Ye, K., and Smith, E. P. (2006). Evaluating water quality using power priors to incorporate historical information. Environmetrics: The Official Journal of the International Environmetrics Society , 17(1):95--106

Show all 24 references
  1. [9]

    Eggleston, B., Wilson, D., McNeil, B., Ibrahim, J., and Catellier, D. (2019). BayesCTDesign: Two Arm Bayesian Clinical Trial Design with and Without Historical Control Data . R package version 0.6.0

  2. [10]

    Fleming, T. R. (2010). Clinical trials: discerning hype from substance. Annals of Internal Medicine , 153(6):400--406

  3. [11]

    C., Batshaw, M., Dunkle, M., Gopal-Srivastava, R., Kaye, E., Krischer, J., Nguyen, T., Paulus, K., Merkel, P

    Griggs, R. C., Batshaw, M., Dunkle, M., Gopal-Srivastava, R., Kaye, E., Krischer, J., Nguyen, T., Paulus, K., Merkel, P. A., et al. (2009). Clinical research for rare disease: opportunities, challenges, and solutions. Molecular Genetics and Metabolism , 96(1):20--26

  4. [12]

    P., Carlin, B

    Hobbs, B. P., Carlin, B. P., Mandrekar, S. J., and Sargent, D. J. (2011). Hierarchical commensurate and power prior models for adaptive incorporation of historical information in clinical trials. Biometrics , 67(3):1047--1056

  5. [13]

    Ibrahim, J. G. and Chen, M.-H. (2000). Power prior distributions for regression models. Statistical Science , pages 46--60

  6. [14]

    Utilisation des immunoglobulines non spécifiques intraveineuses et sous-cutanées au québec, 2020-2021

    Institut National de Santé Publique du Québec (2023). Utilisation des immunoglobulines non spécifiques intraveineuses et sous-cutanées au québec, 2020-2021. Report, Institut National de Santé Publique du Québec. Available from: https://www.inspq.qc.ca/biovigilance

  7. [15]

    C., Davis, L

    Leon, A. C., Davis, L. L., and Kraemer, H. C. (2011). The role and interpretation of pilot studies in clinical research. Journal of Psychiatric Research , 45(5):626--629

  8. [16]

    E., Saris, C

    Lim, J., Eftimov, F., Verhamme, C., Brusse, E., Hoogendijk, J. E., Saris, C. G., Raaphorst, J., De Haan, R. J., van Schaik, I. N., Aronica, E., et al. (2021). Intravenous immunoglobulins as first-line treatment in idiopathic inflammatory myopathies: A pilot study. Rheumatology...

  9. [17]

    Pan, H., Yuan, Y., and Xia, J. (2017). A calibrated power prior approach to borrow information from historical data with application to biosimilar clinical trials. Journal of the Royal Statistical Society Series C: Applied Statistics , 66(5):979--996

  10. [18]

    Psioda, M. A. and Ibrahim, J. G. (2019). B ayesian clinical trial design using historical data that inform the treatment effect. Biostatistics , 20(3):400--415

  11. [19]

    V., and Wason, J

    Qi, Y., Hampson, L. V., and Wason, J. M. S. (2022). Sample size calculation for clinical trials using meta-analytic-predictive priors. Biostatistics , 23(1):48--63

  12. [20]

    Schmidli, H., Gsteiger, S., Roychoudhury, S., O'Hagan, A., Spiegelhalter, D., and Neuenschwander, B. (2014). Robust meta-analytic-predictive priors in clinical trials with historical control information. Biometrics , 70(4):1023--1032

  13. [21]

    A., and Ibrahim, J

    Shen, Y., Psioda, M. A., and Ibrahim, J. G. (2021). BayesPPD : An R package for bayesian sample size determination using the power and normalized power prior for generalized linear models. arXiv preprint arXiv:2112.14616

  14. [22]

    P., Robson, R., Thabane, M., Giangregorio, L., and Goldsmith, C

    Thabane, L., Ma, J., Chu, R., Cheng, J., Ismaila, A., Rios, L. P., Robson, R., Thabane, M., Giangregorio, L., and Goldsmith, C. H. (2010). A tutorial on pilot studies: T he what, why and how. BMC Medical Research Methodology , 10(1):1--10

  15. [23]

    Food and Drug Administration (2010)

    U.S. Food and Drug Administration (2010). Guidance for the use of B ayesian statistics in medical device clinical trials. https://www.fda.gov/regulatory-information/search-fda-guidance-documents/guidance-use-bayesian-statistics-medical-device-clinical-trials. Accessed: 2025-05-05

  16. [24]

    Weber, S. (2021). RBesT : R Bayesian Evidence Synthesis Tools . R package version 1.6-2

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.