{"id":"d932869c-b2dc-48eb-9f50-74da4ae0b3ae","arxiv_id":"2411.16938","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A new fragility index for single-arm time-to-event trials counts how many censored observations must become events to flip a Bayesian conclusion about median survival.","lead":"The paper defines a Fragility Index for time-to-event outcomes in single-arm clinical trials: the smallest number of censored patients that, if reclassified as having the event, would push the Bayesian posterior probability of a survival threshold below a chosen confidence level. The measure is computed with a standard exponential survival model and demonstrated on three oncology datasets, with an R package provided.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"FI definition does not specify the event time assigned to reclassified censorings; using the censoring time as the event time is an arbitrary choice that can inflate the index.","rationale":"The reader's weakest assumption is the exponential distribution with constant hazard. This is a valid concern about external validity, but the FI remains a well-defined quantity under that model. A more fundamental issue is the ambiguous counterfactual used in the reclassification: the paper does not state what event time is assigned when a censored observation is changed to an event. The implementation uses the censoring time, but this is arbitrary. Because the posterior Gamma parameters depend on the total follow-up time β', shifting the assumed event time changes the posterior probability and hence the FI. The central claim—that the FI is the smallest number of censored observations causing a loss of confidence—is not well-defined unless the imputation rule is specified and justified. My concrete test would show whether this ambiguity materially changes the index. If the FI is robust to the imputation choice (e.g., changes by at most one for all three case studies), the concern is minor. If it changes substantially, the definition needs revision or explicit defense. I therefore recommend keeping a conditional verdict, with the added condition that the authors specify and justify the event-time imputation for reclassified censorings and report sensitivity of the FI to this choice. This does not require rejecting the paper; the metric may still be useful after clarification, but the current manuscript does not supply the needed specification.","tokens_in":7452,"tokens_out":6571,"duration_ms":65555,"concrete_test":"Recompute the FI for Case Study 1 (lung dataset, 30 patients, threshold 7 months, p0=0.7) under three event-time imputation rules for reclassified censorings: (a) event time = censoring time (current), (b) event time = 0.5 times censoring time, and (c) event time = the minimum observed event time in the dataset. If the FI changes by more than one event between rules, the index is not invariant to the imputation choice and the claimed definition is incomplete. Repeat for Case Study 2 and 3 to confirm.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The Fragility Index is defined as the smallest number of censored observations that, when reclassified as uncensored events, drops the posterior probability of the median survival time exceeding a threshold below a confidence level. The paper's implementation reclassifies a censored observation (T_i, δ_i=0) to (T_i, δ_i=1), effectively treating the censoring time T_i as the event time. However, for a censored observation the only information is that the event time exceeds T_i; the precise event time under the counterfactual is unobserved and could be any value greater than T_i. If the event is assumed to occur earlier than T_i (e.g., at 0.5 × T_i or at the smallest observed event time), the total follow-up time β' decreases, the posterior hazard increases, and the posterior probability P(tmed > t0) decreases more rapidly. Consequently, the FI would be smaller (more fragile) under equally plausible imputation rules. The paper does not justify the choice of T_i as the event time, nor does it assess how sensitive the FI is to this choice. This makes the FI not a uniquely defined property of the data but an artifact of an unstated imputation convention. The exponential-model assumption identified by the reader is a real limitation, but the imputation ambiguity is more load-bearing because it affects the central claim even when the exponential model is correct.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces a Fragility Index (FI) for single-arm time-to-event trials under a Bayesian exponential survival model with a Gamma prior. The FI is defined as the smallest number of censored observations that, when reclassified as uncensored events, cause the posterior probability that the median survival time exceeds a specified threshold to fall below a confidence level. The authors derive the posterior distribution of the rate parameter, the median survival time, and the posterior probability in three theorems, describe an R package implementation, and illustrate the method on three real-world datasets. The central theoretical derivation is the Gamma posterior for the exponential rate parameter, which is correct; the main concerns are about the precise definition of the reclassification perturbation and the unexamined parametric assumption.","tokens_in":7737,"tokens_out":8092,"duration_ms":81546,"significance":"If the definitional ambiguity is resolved, the paper offers a simple and computationally transparent sensitivity measure that is a natural complement to existing fragility indices for binary outcomes and two-arm survival trials. The closed-form Gamma posterior makes the FI easy to compute, and the accompanying R package is a practical contribution for applied researchers. The case studies demonstrate feasibility, though not validity under model misspecification. The theoretical content is elementary, but the methodological framing is useful for single-arm oncology and chronic-disease trials, where fragility measures are largely absent. The main value of the paper lies in its applied accessibility rather than in statistical depth.","major_comments":[{"comment":"The definition of the FI says that censored observations are \"reclassified as uncensored events\" but does not specify what event time is assigned to a reclassified observation. The implementation implicitly keeps the censoring time T_i as the event time, changing only the censoring indicator. A censored observation, however, only establishes that the true event time exceeds T_i; the observational data are equally compatible with counterfactual event times smaller than T_i. Under an alternative imputation rule (for example, imputing the event at 0.5*T_i or at the smallest observed event time), the updated rate parameter beta' is smaller, and the posterior probability P(tmed > t0) decreases more rapidly, yielding a smaller FI. The FI is therefore not a uniquely defined property of the data unless the perturbation experiment is fully specified. Please state the perturbation rule explicitly and either justify keeping T_i as the event time or assess the sensitivity of the FI to reasonable alternative imputation rules.","section":"Section 2.3"},{"comment":"The phrase \"smallest number k of censored observations with the shortest censoring times\" defines the index relative to a particular ordering, but the paper does not prove that reclassifying the shortest censoring times first indeed yields the smallest possible k. Because the effect of reclassifying a censored observation on the Gamma posterior depends on its time T_i, a lemma establishing the required monotonicity (or a counterexample) is needed to justify the \"smallest number\" terminology. Without such a proof, the index is merely the stopping point of one deterministic reclassification sequence rather than a demonstrated minimum.","section":"Section 2.3"},{"comment":"The entire posterior probability computation and the resulting FI rest on the exponential survival model with constant hazard. The case studies apply the method without any model-checking for the three datasets, so a time-varying hazard in any of these studies would make the reported posterior probability and FI reflect a misspecified model. Since this is a load-bearing assumption for the definition, the paper should provide goodness-of-fit diagnostics (for example, a comparison with a Weibull or piecewise-exponential model, or a graphical check of exponentiality) and discuss how the FI would change under model misspecification.","section":"Section 2.1 and Section 3"},{"comment":"The lung cancer case study analyzes a randomly selected subset of 30 patients from the 'lung' dataset, but the selection procedure (including the random seed) is not described, which makes that particular FI value non-reproducible and dependent on an arbitrary subsample. The concluding section also identifies the prior as an influence on the FI but does not mention the exponential-model assumption or the dependence on the chosen threshold p0 and target t0. Please either analyze the full dataset or describe the subset selection and report a sensitivity analysis across subsets, thresholds, and prior parameters.","section":"Section 3.1 and Section 4"}],"minor_comments":[{"comment":"The section is titled \"Case Study 2: Pembrolizumab in hepatocellular carcinoma (HCC)\", but the text begins \"For the third case study,\" and the caption of Figure 2 refers to a breast cancer dataset while the text describes HCC; these inconsistencies should be corrected.","section":"Section 3.2"},{"comment":"The likelihood definition contains a typo: \"for i ≤ i ≤ n\" should read \"for i = 1, . . . , n.\"","section":"Section 2.1"},{"comment":"The description of p0 = 0.7 as \"a standard choice balancing statistical confidence and flexibility\" is not accompanied by a citation or justification; either provide a reference or rephrase this as an arbitrary but reasonable default.","section":"Section 2.3"},{"comment":"For reproducibility, the R package is available only through a GitHub link without a version identifier; please provide a versioned release or archive, and include the random seed used for the lung-data subset selection.","section":"Section 3 and Appendix"},{"comment":"In Case Study 3, the text says \"the remaining patients were censored\" without giving the number; specifying that 20 of 51 patients were censored would be clearer.","section":"Section 3.3"}],"recommendation":"major_revision","confidential_remarks":"The paper's central posterior derivation is correct, and the proposed FI is potentially useful in applied settings. The main issue is the unstated imputation convention for reclassified censorings, which is fixable but must be addressed before the FI can be regarded as a well-defined index. The paper is methodologically light for a statistics journal, so the editor may also want to weigh whether the applied angle is sufficient for the journal's scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a legitimate extension of the fragility index to single-arm time-to-event trials, and the Bayesian calculations are correct. The definition, however, has an unstated imputation rule: when a censored observation is reclassified as an event, the paper keeps the censoring time as the event time. That choice is arbitrary and affects the value of the FI.\n\nWhat is new: Walsh et al. gave FI for binary outcomes, Bomze et al. gave Survival-Inferred FI for two-arm trials with log-rank. This paper does single-arm Bayesian exponential. The extension is natural but missing. The posterior of λ is textbook, and the R package is a practical addition.\n\nWhere it is soft: First, the unstated imputation rule. If you instead assign the event at, say, half the censoring time, the total follow-up time decreases, the posterior hazard increases, and the posterior probability drops faster, so the FI is smaller. The paper does not justify its choice and does not test sensitivity. Second, the exponential model is a real limitation for a robustness measure, though it is explicitly stated. Third, reproducibility: the first case study uses a random subset of lung data that is not seeded, and the other two datasets are reconstructed from published KM curves, so the exact data are not available. Fourth, there is no simulation or comparison with other measures, so calling FI=5 or 6 'moderate robustness' is an interpretation, not an established calibration.\n\nA minor internal inconsistency: the paper says it reclassifies censored observations with the shortest censoring times, but any censored observation has the same effect on the posterior because flipping δ adds 1 to α' and leaves β' unchanged. That ordering is irrelevant.\n\nWho it is for: applied statisticians and clinical trialists working with single-arm oncology studies who want a quick sensitivity check to complement Bayesian posterior probabilities. It deserves a serious referee because the concept is useful and the package is usable, but the authors should be asked to specify the imputation rule and report how the FI varies under alternative rules.","headline":"A useful extension of the fragility index to single-arm survival trials, but the reclassification rule for censored observations is undefined and affects the index.","tokens_in":610,"tokens_out":731,"would_cite":true,"duration_ms":42882,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F15","62N01","62P10"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper defines a Fragility Index for survival endpoints in single-arm trials: the minimum number of censored patients that, when recoded as events, would push the posterior probability that median survival exceeds a threshold below a…","keywords":["Bayesian analysis","Exponential survival model","Fragility Index","Single-arm clinical trials","Time-to-event data","censoring","posterior probability","survival analysis"],"falsifier":"Take a single-arm survival dataset whose Kaplan-Meier curve visibly bends (hazard changing over time), compute the FI under the paper's exponential model, and compare it with the same reclassification count evaluated under a piecewise-exponential or Weibull fit; if the index changes substantially, the exponential assumption, not the data, is driving the fragility number.","tokens_in":7253,"feed_emoji":"⏳","tokens_out":13505,"duration_ms":107936,"temperature":0.7,"pith_summary":"The paper defines a Fragility Index (FI) for time-to-event endpoints in single-arm clinical trials, a setting where there is no control arm to lean on. The FI is the smallest number of censored observations that, when reclassified as uncensored events, makes the Bayesian posterior probability that median survival exceeds a chosen threshold drop below a preset confidence level. The index is worked out under an exponential survival model with a Gamma prior, which gives a closed-form Gamma posterior and makes the reclassification count easy to compute; the authors supply an accompanying software tool. In three real single-arm datasets—advanced lung cancer, pembrolizumab in hepatocellular carcinoma, and palbociclib in breast cancer—the FI came out as 5, 6, and 6, which they interpret as moderate robustness. The point of the index is to tell clinicians how much of a survival conclusion rests on the status of a few censored patients.","feed_headline":"Fragility index counts censored patients who can flip a verdict","feed_subtitle":"In three single-arm cancer trials, recoding five or six censored patients as events flipped the verdict.","key_machinery":"The central object is the Fragility Index itself, computed sequentially. The carrying identity is the exponential median formula $t_{\\mathrm{med}} = \\ln 2/\\lambda$, which converts the clinical statement \"median survival exceeds $t_0$\" into the one-sided posterior tail $P(\\lambda < \\ln 2/t_0 \\mid \\mathrm{data})$. With a Gamma($\\alpha, \\beta$) prior and the exponential likelihood, the posterior is Gamma($\\alpha + \\sum \\delta_i, \\beta + \\sum T_i$), so each reclassification of a censored observation updates only two numbers—the total event count and the total follow-up time—before the tail probability is re-evaluated. The algorithm recodes censored observations in increasing order of censoring time until the posterior probability sinks below $p_0$; the number of reclassifications needed is the FI.","core_discovery":"The paper's central claim is that a single-arm trial's survival conclusion can be assigned a single fragility number. Starting from the exponential model with rate $\\lambda$ and a Gamma($\\alpha, \\beta$) prior, the posterior is Gamma($\\alpha + \\sum \\delta_i, \\beta + \\sum T_i$); since the median survival time is $t_{\\mathrm{med}} = \\ln 2/\\lambda$, the posterior probability that median survival exceeds $t_0$ reduces to the gamma tail $P(\\lambda < \\ln 2/t_0 \\mid \\mathrm{data})$. The Fragility Index is the smallest $k$ such that recoding the $k$ censored patients with the shortest censoring times as events pushes that tail probability below the confidence level $p_0$. Under this definition, each recoding changes the posterior by incrementing the event count and the total survival time, so the index can be read off by sequential recalculation. The authors report FI values of 5 (lung cancer, threshold 7 months), 6 (pembrolizumab/HCC, threshold 3.5 months), and 6 (palbociclib/breast cancer, threshold 15 months) at $p_0 = 0.7$.","pith_inferences":["Because the exponential assumption is load-bearing, an immediate testable extension is to recompute the same reclassification count under survival models where the risk changes over time; if the FI changes materially, the index is measuring model sensitivity as much as data fragility.","The FI as defined only moves censored observations into the event category; a symmetric version that also reclassifies events as censored would reveal whether the conclusion is fragile in both directions.","The paper does not quantify uncertainty in the FI itself; a bootstrap or prior-perturbation study could show how the reclassification count varies, which would help interpret single point estimates like 5 or 6.","The three examples all use $p_0 = 0.7$; applying the metric at conventional 0.95 or 0.975 confidence levels would likely give lower FIs, so the choice of confidence level deserves reporting alongside the index."],"forward_implications":["A low FI in a single-arm survival trial is a warning that the conclusion that median survival beats the threshold rests on the censoring status of a small number of patients.","The FI can be computed from summary quantities—number of events and total observed time—not just from full individual-level data, because the posterior depends only on those sums.","The index is meant as a complement to the posterior-probability decision rule, not a standalone significance test, since the authors note there is no universal FI threshold.","In the three case studies, FI values of 5 and 6 at confidence level 0.7 indicate conclusions that survive moderate reclassification but are not immune to it."],"supporting_citations":[{"why":"introduces the original binary-outcome Fragility Index that this paper extends to survival endpoints.","marker":"Walsh et al. (2014)"},{"why":"introduces the Survival-Inferred Fragility Index for two-arm trials, the predecessor this single-arm index is designed to complement.","marker":"Bomze et al. (2020)"},{"why":"establishes the exponential life-testing model whose likelihood and median identity the posterior calculation uses.","marker":"Epstein and Sobel, 1953, 1954; Epstein, 1954"},{"why":"supplies the pembrolizumab hepatocellular carcinoma case-study data reconstructed from its Kaplan-Meier curve.","marker":"Feun et al. (2019)"},{"why":"supplies the palbociclib breast-cancer case-study data reconstructed from its Kaplan-Meier curve.","marker":"Krishnamurthy et al. (2022)"},{"why":"provides software for binary-outcome fragility that the paper's own tool parallels.","marker":"Lin and Chu (2022)"}],"fun_headline_variants":["Fragility index: how many censored deaths can flip a trial's verdict","A single number quantifies how robust a survival conclusion is","Recoding a few censored patients as events can flip trial results","In three cancer trials, just 5 or 6 patients shift the verdict","New metric shows survival results can hinge on a handful of patients"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole index assumes that survival times follow an exponential distribution with a constant risk of the event over time; if that risk rises or falls, the posterior probability and the reclassification count come from a misspecified model, so the FI may not reflect the true fragility of the conclusion.","fun_headline_variants_meta":{"raw":{"variants":["Fragility index: how many censored deaths can flip a trial's verdict","A single number quantifies how robust a survival conclusion is","Recoding a few censored patients as events can flip trial results","In three cancer trials, just 5 or 6 patients shift the verdict","New metric shows survival results can hinge on a handful of patients"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000766,"raw_usage":{"total_tokens":3430,"prompt_tokens":1015,"completion_tokens":2415,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":631,"completion_tokens_details":{"reasoning_tokens":2322}},"tokens_in":631,"tokens_out":2415,"duration_ms":16720,"temperature":1.0,"reasoning_tokens":2322,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T12:42:50.854538+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a single-arm survival dataset whose Kaplan-Meier curve visibly bends (hazard changing over time), compute the FI under the paper's exponential model, and compare it with the same reclassification count evaluated under a piecewise-exponential or Weibull fit; if the index changes substantially, the exponential assumption, not the data, is driving the fragility number.","supporting_citations":[{"cited_title":"The statistical significance of randomized controlled trial results is frequently fragile: a case for a fragility index","cited_arxiv_id":null,"evidence_quote":"introduces the original binary-outcome Fragility Index that this paper extends to survival endpoints."},{"cited_title":"Survival-inferred fragility index of phase 3 clinical trials evaluating immune checkpoint inhibitors","cited_arxiv_id":null,"evidence_quote":"introduces the Survival-Inferred Fragility Index for two-arm trials, the predecessor this single-arm index is designed to complement."},{"cited_title":"Life testing","cited_arxiv_id":null,"evidence_quote":"establishes the exponential life-testing model whose likelihood and median identity the posterior calculation uses."},{"cited_title":"Phase 2 study of pembrolizumab and circulating biomarkers to predict anticancer response in advanced, unresectable hepatocellular carcinoma","cited_arxiv_id":null,"evidence_quote":"supplies the pembrolizumab hepatocellular carcinoma case-study data reconstructed from its Kaplan-Meier curve."},{"cited_title":"A phase ii trial of an alternative schedule of palbociclib and embedded serum tk1 analysis","cited_arxiv_id":null,"evidence_quote":"supplies the palbociclib breast-cancer case-study data reconstructed from its Kaplan-Meier curve."},{"cited_title":"Assessing and visualizing fragility of clinical results with binary outcomes in r using the fragility package","cited_arxiv_id":null,"evidence_quote":"provides software for binary-outcome fragility that the paper's own tool parallels."}],"review_version":1}