Pith. sign in

REVIEW 4 major objections 5 minor 19 references

Fragility Index for Time-to-Event Endpoints in Single-Arm Clinical Trials

T0 review · 4 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read This paper defines a Fragility Index for survival endpoints in single-arm trials: the minimum number of censored patients that, when recoded as events, would push the posterior probability that median survival exceeds a threshold below a…

desk verdict A useful extension of the fragility index to single-arm survival trials, but the reclassification rule for censored observations is undefined and affects the index. read the letter →

arxiv 2411.16938 v1 pith:LVKD42FW submitted 2024-11-25 stat.ME stat.AP

classification stat.MEstat.AP MSC 62F1562N0162P10
keywords BayesiananalysisExponentialsurvivalmodelFragilityIndexSingle-armclinicaltrialsTime-to-eventdatacensoringposteriorprobability
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper defines a Fragility Index (FI) for time-to-event endpoints in single-arm clinical trials, a setting where there is no control arm to lean on. The FI is the smallest number of censored observations that, when reclassified as uncensored events, makes the Bayesian posterior probability that median survival exceeds a chosen threshold drop below a preset confidence level. The index is worked out under an exponential survival model with a Gamma prior, which gives a closed-form Gamma posterior and makes the reclassification count easy to compute; the authors supply an accompanying software tool. In three real single-arm datasets—advanced lung cancer, pembrolizumab in hepatocellular carcinoma, and palbociclib in breast cancer—the FI came out as 5, 6, and 6, which they interpret as moderate robustness. The point of the index is to tell clinicians how much of a survival conclusion rests on the status of a few censored patients.

What carries the argument

The central object is the Fragility Index itself, computed sequentially. The carrying identity is the exponential median formula $t_{\mathrm{med}} = \ln 2/\lambda$, which converts the clinical statement "median survival exceeds $t_0$" into the one-sided posterior tail $P(\lambda < \ln 2/t_0 \mid \mathrm{data})$. With a Gamma($\alpha, \beta$) prior and the exponential likelihood, the posterior is Gamma($\alpha + \sum \delta_i, \beta + \sum T_i$), so each reclassification of a censored observation updates only two numbers—the total event count and the total follow-up time—before the tail probability is re-evaluated. The algorithm recodes censored observations in increasing order of censoring time until the posterior probability sinks below $p_0$; the number of reclassifications needed is the FI.

What would settle it

Take a single-arm survival dataset whose Kaplan-Meier curve visibly bends (hazard changing over time), compute the FI under the paper's exponential model, and compare it with the same reclassification count evaluated under a piecewise-exponential or Weibull fit; if the index changes substantially, the exponential assumption, not the data, is driving the fragility number.

Watch

Extended reading notes

Core claim

The paper's central claim is that a single-arm trial's survival conclusion can be assigned a single fragility number. Starting from the exponential model with rate $\lambda$ and a Gamma($\alpha, \beta$) prior, the posterior is Gamma($\alpha + \sum \delta_i, \beta + \sum T_i$); since the median survival time is $t_{\mathrm{med}} = \ln 2/\lambda$, the posterior probability that median survival exceeds $t_0$ reduces to the gamma tail $P(\lambda < \ln 2/t_0 \mid \mathrm{data})$. The Fragility Index is the smallest $k$ such that recoding the $k$ censored patients with the shortest censoring times as events pushes that tail probability below the confidence level $p_0$. Under this definition, each recoding changes the posterior by incrementing the event count and the total survival time, so the index can be read off by sequential recalculation. The authors report FI values of 5 (lung cancer, threshold 7 months), 6 (pembrolizumab/HCC, threshold 3.5 months), and 6 (palbociclib/breast cancer, threshold 15 months) at $p_0 = 0.7$.

Load-bearing premise

The whole index assumes that survival times follow an exponential distribution with a constant risk of the event over time; if that risk rises or falls, the posterior probability and the reclassification count come from a misspecified model, so the FI may not reflect the true fragility of the conclusion.

Editorial extensions

If this is right

  • A low FI in a single-arm survival trial is a warning that the conclusion that median survival beats the threshold rests on the censoring status of a small number of patients.
  • The FI can be computed from summary quantities—number of events and total observed time—not just from full individual-level data, because the posterior depends only on those sums.
  • The index is meant as a complement to the posterior-probability decision rule, not a standalone significance test, since the authors note there is no universal FI threshold.
  • In the three case studies, FI values of 5 and 6 at confidence level 0.7 indicate conclusions that survive moderate reclassification but are not immune to it.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because the exponential assumption is load-bearing, an immediate testable extension is to recompute the same reclassification count under survival models where the risk changes over time; if the FI changes materially, the index is measuring model sensitivity as much as data fragility.
  • The FI as defined only moves censored observations into the event category; a symmetric version that also reclassifies events as censored would reveal whether the conclusion is fragile in both directions.
  • The paper does not quantify uncertainty in the FI itself; a bootstrap or prior-perturbation study could show how the reclassification count varies, which would help interpret single point estimates like 5 or 6.
  • The three examples all use $p_0 = 0.7$; applying the metric at conventional 0.95 or 0.975 confidence levels would likely give lower FIs, so the choice of confidence level deserves reporting alongside the index.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper introduces a Fragility Index (FI) for single-arm time-to-event trials under a Bayesian exponential survival model with a Gamma prior. The FI is defined as the smallest number of censored observations that, when reclassified as uncensored events, cause the posterior probability that the median survival time exceeds a specified threshold to fall below a confidence level. The authors derive the posterior distribution of the rate parameter, the median survival time, and the posterior probability in three theorems, describe an R package implementation, and illustrate the method on three real-world datasets. The central theoretical derivation is the Gamma posterior for the exponential rate parameter, which is correct; the main concerns are about the precise definition of the reclassification perturbation and the unexamined parametric assumption.

Significance. If the definitional ambiguity is resolved, the paper offers a simple and computationally transparent sensitivity measure that is a natural complement to existing fragility indices for binary outcomes and two-arm survival trials. The closed-form Gamma posterior makes the FI easy to compute, and the accompanying R package is a practical contribution for applied researchers. The case studies demonstrate feasibility, though not validity under model misspecification. The theoretical content is elementary, but the methodological framing is useful for single-arm oncology and chronic-disease trials, where fragility measures are largely absent. The main value of the paper lies in its applied accessibility rather than in statistical depth.

major comments (4)
  1. [Section 2.3] The definition of the FI says that censored observations are "reclassified as uncensored events" but does not specify what event time is assigned to a reclassified observation. The implementation implicitly keeps the censoring time T_i as the event time, changing only the censoring indicator. A censored observation, however, only establishes that the true event time exceeds T_i; the observational data are equally compatible with counterfactual event times smaller than T_i. Under an alternative imputation rule (for example, imputing the event at 0.5*T_i or at the smallest observed event time), the updated rate parameter beta' is smaller, and the posterior probability P(tmed > t0) decreases more rapidly, yielding a smaller FI. The FI is therefore not a uniquely defined property of the data unless the perturbation experiment is fully specified. Please state the perturbation rule explicitly and either justify keeping T_i as the event time or assess the sensitivity of the FI to reasonable alternative imputation rules.
  2. [Section 2.3] The phrase "smallest number k of censored observations with the shortest censoring times" defines the index relative to a particular ordering, but the paper does not prove that reclassifying the shortest censoring times first indeed yields the smallest possible k. Because the effect of reclassifying a censored observation on the Gamma posterior depends on its time T_i, a lemma establishing the required monotonicity (or a counterexample) is needed to justify the "smallest number" terminology. Without such a proof, the index is merely the stopping point of one deterministic reclassification sequence rather than a demonstrated minimum.
  3. [Section 2.1 and Section 3] The entire posterior probability computation and the resulting FI rest on the exponential survival model with constant hazard. The case studies apply the method without any model-checking for the three datasets, so a time-varying hazard in any of these studies would make the reported posterior probability and FI reflect a misspecified model. Since this is a load-bearing assumption for the definition, the paper should provide goodness-of-fit diagnostics (for example, a comparison with a Weibull or piecewise-exponential model, or a graphical check of exponentiality) and discuss how the FI would change under model misspecification.
  4. [Section 3.1 and Section 4] The lung cancer case study analyzes a randomly selected subset of 30 patients from the 'lung' dataset, but the selection procedure (including the random seed) is not described, which makes that particular FI value non-reproducible and dependent on an arbitrary subsample. The concluding section also identifies the prior as an influence on the FI but does not mention the exponential-model assumption or the dependence on the chosen threshold p0 and target t0. Please either analyze the full dataset or describe the subset selection and report a sensitivity analysis across subsets, thresholds, and prior parameters.
minor comments (5)
  1. [Section 3.2] The section is titled "Case Study 2: Pembrolizumab in hepatocellular carcinoma (HCC)", but the text begins "For the third case study," and the caption of Figure 2 refers to a breast cancer dataset while the text describes HCC; these inconsistencies should be corrected.
  2. [Section 2.1] The likelihood definition contains a typo: "for i ≤ i ≤ n" should read "for i = 1, . . . , n."
  3. [Section 2.3] The description of p0 = 0.7 as "a standard choice balancing statistical confidence and flexibility" is not accompanied by a citation or justification; either provide a reference or rephrase this as an arbitrary but reasonable default.
  4. [Section 3 and Appendix] For reproducibility, the R package is available only through a GitHub link without a version identifier; please provide a versioned release or archive, and include the random seed used for the lung-data subset selection.
  5. [Section 3.3] In Case Study 3, the text says "the remaining patients were censored" without giving the number; specifying that 20 of 51 patients were censored would be clearer.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: FI is a defined sensitivity metric, not a derived prediction.

full rationale

The paper's central object, the Fragility Index, is introduced by definition as the smallest number of censored observations that, when reclassified as events, drops the posterior probability below a confidence level. The posterior probability itself follows from a standard conjugate Bayesian calculation: exponential likelihood plus Gamma prior yields a Gamma posterior with updated parameters alpha' = alpha + sum(delta_i) and beta' = beta + sum(T_i). The median survival formula t_med = ln(2)/lambda is a standard property of the exponential distribution, and the probability P(t_med > t0) is just the Gamma CDF evaluated at ln(2)/t0. Reclassifying a censored observation as an event changes alpha' by 1 while leaving beta' unchanged, and the FI is computed by sequentially applying this perturbation until the threshold is crossed. This is a deterministic sensitivity analysis, not a prediction derived from fitted parameters, and no parameter is fitted to reproduce the FI values. The paper contains no load-bearing self-citation chain: its references to prior fragility-index literature provide background, not support for the derivation. The skeptical concern that reclassifying a censored time as an event at the censoring time is an arbitrary imputation convention is a validity or ambiguity caveat, not a circularity, because the FI is explicitly defined under that perturbation rule and the paper's claims are conditional on its own definition. No equation in the paper reduces by construction to its inputs, and the case-study FI values are reported as computed quantities rather than as confirmations of the method. Therefore the analysis is self-contained with respect to circularity, and the appropriate score is 0.

Assumptions & free parameters 3 free parameters · 3 assumptions · 0 invented entities

The central metric depends on three free choices: the Gamma prior parameters and the confidence threshold. The exponential model and the reclassification rule are explicit modeling assumptions. No new physical entities are introduced.

free parameters (3)
  • Gamma prior shape parameter alpha = 0.5
    Chosen by the authors as weakly informative; the FI depends on this value.
  • Gamma prior rate parameter beta = 0.5
    Chosen by the authors as weakly informative; the FI depends on this value.
  • Confidence threshold p0 = 0.7
    Chosen by the authors for the case studies and called 'standard' without a supporting reference; it directly determines the FI.
assumptions (3)
  • domain assumption Survival times follow an exponential distribution with constant hazard.
    The paper assumes this model for all case studies; if misspecified, the posterior probability and the FI are affected.
  • standard math The Gamma prior is conjugate to the exponential likelihood.
    Used to derive the closed-form posterior; this is a textbook result.
  • ad hoc to paper Reclassified censored observations are assigned event times equal to their observed censoring times.
    The FI definition requires a reclassification rule; the paper implicitly uses the censoring time as the event time when changing delta from 0 to 1.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Fragility Index for Time-to-Event Endpoints in Single-Arm Clinical Trials." pith.science (2026). https://pith.science/paper/LVKD42FW

@misc{pith2026241116938,
  author       = {Pith},
  title        = {Pith review of: Fragility Index for Time-to-Event Endpoints in Single-Arm Clinical Trials},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/LVKD42FW}},
  note         = {Machine review of arXiv:2411.16938}
}
read the original abstract

The reliability of clinical trial outcomes is crucial, especially in guiding medical decisions. In this paper, we introduce the Fragility Index (FI) for time-to-event endpoints in single-arm clinical trials - a novel metric designed to quantify the robustness of study conclusions. The FI represents the smallest number of censored observations that, when reclassified as uncensored events, causes the posterior probability of the median survival time exceeding a specified threshold to fall below a predefined confidence level. While drug effectiveness is typically assessed by determining whether the posterior probability exceeds a specified confidence level, the FI offers a complementary measure, indicating how robust these conclusions are to potential shifts in the data. Using a Bayesian approach, we develop a practical framework for computing the FI based on the exponential survival model. To facilitate the application of our method, we developed an R package fi, which provides a tool to compute the Fragility Index. Through real world case studies involving time to event data from single arms clinical trials, we demonstrate the utility of this index. Our findings highlight how the FI can be a valuable tool for assessing the robustness of survival analyses in single-arm studies, aiding researchers and clinicians in making more informed decisions.

Figures

Figures reproduced from arXiv: 2411.16938 by the authors.

Figure 1
Figure 1. Kaplan-Meier curve for the lung cancer dataset The Fragility Index was determined by sequentially reclassifying censored observations with the shortest censoring times as events and recalculating the posterior probability until it fell below the predefined confidence threshold of 0.7, a standard choice balancing statistical confidence and flexibility. For this dataset, the Fragility Index was found to be 5. This mea… view at source ↗
Figure 2
Figure 2. Kaplan-Meier curve for the breast cancer dataset treated with Pembrolizumab The Fragility Index for this dataset was determined to be 6, meaning that reclassifying six censored patients as having experienced the event (disease progression) would reduce the posterior probability of the median progression-free survival time exceeding 3.5 months to below 0.7. Given the sample size of 28, an FI of 6 indicates that the s… view at source ↗
Figure 3
Figure 3. Kaplan-Meier curve for the breast cancer dataset treated with Palbociclib The Fragility Index (FI) for this dataset was calculated to be 6, indicating that if six censored patients were reclassified as having experienced the event (disease progression), the posterior probability of the median survival time exceeding 15 months would drop below 0.7. With a sample size of 51, an FI of 6 suggests that the study’s conclu… view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

19 extracted references · 19 canonical work pages

  1. [1]

    The use and limitations of the fragility index in the interpretation of clinical trial findings

    Chittaranjan Andrade. The use and limitations of the fragility index in the interpretation of clinical trial findings. The Journal of Clinical Psychiatry, 81 0 (2): 0 21994, 2020

  2. [2]

    Fragility indices for only sufficiently likely modifications

    Benjamin R Baer, Mario Gaudino, Mary Charlson, Stephen E Fremes, and Martin T Wells. Fragility indices for only sufficiently likely modifications. Proceedings of the National Academy of Sciences, 118 0 (49): 0 e2105254118, 2021

  3. [3]

    Survival-inferred fragility index of phase 3 clinical trials evaluating immune checkpoint inhibitors

    David Bomze, Nethanel Asher, Omar Hasan Ali, Lukas Flatz, Daniel Azoulay, Gal Markel, and Tomer Meirson. Survival-inferred fragility index of phase 3 clinical trials evaluating immune checkpoint inhibitors. JAMA network open, 3 0 (10): 0 e2017675--e2017675, 2020

  4. [4]

    It's time to talk about ditching statistical significance

    Editorial. It's time to talk about ditching statistical significance. Nature, 567 0 (7748): 0 283, 2019

  5. [5]

    Truncated life tests in the exponential case

    Benjamin Epstein. Truncated life tests in the exponential case. The Annals of Mathematical Statistics, pages 555--564, 1954

  6. [6]

    Life testing

    Benjamin Epstein and Milton Sobel. Life testing. Journal of the American Statistical Association, 48 0 (263): 0 486--502, 1953

  7. [7]

    Some theorems relevant to life testing from an exponential distribution

    Benjamin Epstein and Milton Sobel. Some theorems relevant to life testing from an exponential distribution. The Annals of Mathematical Statistics, pages 373--381, 1954

  8. [8]

    Phase 2 study of pembrolizumab and circulating biomarkers to predict anticancer response in advanced, unresectable hepatocellular carcinoma

    Lynn G Feun, Ying-Ying Li, Chunjing Wu, Medhi Wangpaichitr, Patricia D Jones, Stephen P Richman, Beatrice Madrazo, Deukwoo Kwon, Monica Garcia-Buitrago, Paul Martin, et al. Phase 2 study of pembrolizumab and circulating biomarkers to predict anticancer response in advanced, unresectable hepatocellular carcinoma. Cancer, 125 0 (20): 0 3603--3614, 2019

Show all 19 references
  1. [9]

    Fragility index and fragility quotient in randomized clinical trials

    Marcos Vinicius Fernandes Garcia, Juliana Carvalho Ferreira, and Pedro Caruso. Fragility index and fragility quotient in randomized clinical trials. Jornal Brasileiro de Pneumologia, 49 0 (01): 0 e20230034, 2023

  2. [10]

    The robustness index: going beyond statistical significance by quantifying fragility

    Thomas F Heston. The robustness index: going beyond statistical significance by quantifying fragility. Cureus, 15 0 (8), 2023

  3. [11]

    A phase ii trial of an alternative schedule of palbociclib and embedded serum tk1 analysis

    Jairam Krishnamurthy, Jingqin Luo, Rama Suresh, Foluso Ademuyiwa, Caron Rigden, Timothy Rearden, Katherine Clifton, Katherine Weilbaecher, Ashley Frith, Anna Roshal, et al. A phase ii trial of an alternative schedule of palbociclib and embedded serum tk1 analysis. NPJ breast c...

  4. [12]

    Assessing and visualizing fragility of clinical results with binary outcomes in r using the fragility package

    Lifeng Lin and Haitao Chu. Assessing and visualizing fragility of clinical results with binary outcomes in r using the fragility package. PLoS One, 17 0 (6): 0 e0268754, 2022

  5. [13]

    The fragility of phase iii trials in oncology

    Y Liu, TA Lin, A Koong, C Lin, JA Jaoude, RR Patel, R Kouzy, MB El Alam, T Meirson, and EB Ludmir. The fragility of phase iii trials in oncology. International Journal of Radiation Oncology, Biology, Physics, 120 0 (2): 0 S42, 2024

  6. [14]

    Fibrinolysis for patients with intermediate-risk pulmonary embolism

    Guy Meyer, Eric Vicaut, Thierry Danays, Giancarlo Agnelli, Cecilia Becattini, Jan Beyer-Westendorf, Erich Bluhmki, Helene Bouvaist, Benjamin Brenner, Francis Couturaud, et al. Fibrinolysis for patients with intermediate-risk pulmonary embolism. New England Journal of Medicine,...

  7. [15]

    Statistical fragility of findings from randomized phase 3 trials in pediatric oncology., 2024

    Hannah Olsen, Pei-Chi Kao, Caleb Richmond, David Stephen Shulman, Wendy B London, and Steven G DuBois. Statistical fragility of findings from randomized phase 3 trials in pediatric oncology., 2024

  8. [16]

    Dismantling the fragility index: a demonstration of statistical reasoning

    Gail E Potter. Dismantling the fragility index: a demonstration of statistical reasoning. Statistics in Medicine, 39 0 (26): 0 3720--3731, 2020

  9. [17]

    The fragility index in multicenter randomized controlled critical care trials

    Elliott E Ridgeon, Paul J Young, Rinaldo Bellomo, Marta Mucchetti, Rosalba Lembo, and Giovanni Landoni. The fragility index in multicenter randomized controlled critical care trials. Critical care medicine, 44 0 (7): 0 1278--1284, 2016

  10. [18]

    The fragility index in randomized clinical trials as a means of optimizing patient care

    Christopher J Tignanelli and Lena M Napolitano. The fragility index in randomized clinical trials as a means of optimizing patient care. JAMA surgery, 154 0 (1): 0 74--79, 2019

  11. [19]

    The statistical significance of randomized controlled trial results is frequently fragile: a case for a fragility index

    Michael Walsh, Sadeesh K Srinathan, Daniel F McAuley, Marko Mrkobrada, Oren Levine, Christine Ribic, Amber O Molnar, Neil D Dattani, Andrew Burke, Gordon Guyatt, et al. The statistical significance of randomized controlled trial results is frequently fragile: a case for a frag...

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.