{"id":"30b3f771-1e1c-44cb-9462-1782a3adebd3","arxiv_id":"2505.16835","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"In simulations and a head-and-neck cancer case study, adding real-world survival data to short trial follow-up reduces bias in long-term survival extrapolations with the survextrap model, even when the external data are moderately biased.","lead":"A simulation study tests how well the survextrap Bayesian survival model extrapolates long-term survival when adding real-world registry and population-mortality data to short-term trial data. It finds that including even moderately biased external data improves extrapolation accuracy over trial data alone, and that treatment-effect extrapolations are reliable only when the long-term treatment effect assumption is close to the truth.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The improvement from biased external data is only demonstrated for a constant multiplicative hazard bias under the same data-generating mechanism; more realistic external-data discrepancies may invalidate the abstract's general claim.","rationale":"The central claim is that including even moderately biased external data improves long-term extrapolation. The simulation evidence for this claim is generated under a single, particularly benign form of external-data discrepancy: the external data arise from the same age distribution, disease process, and background mortality as the trial control arm, with only a constant multiplicative bias on the all-cause hazard. This is a strong assumption because the survextrap model has no mechanism to down-weight or adjust for external-data bias; by assuming the external data have the same survival as the control arm, the likelihood will pull the estimated hazard toward the external data. With a constant hazard scaling, this pull is uniform across the post-trial period and the bias in the 40-year RMST remains smaller than the bias from trial-only extrapolation. With structured discrepancies, the pull can be non-uniform and may distort the hazard trajectory in ways that are harder for the smoothness prior to correct, potentially making external data harmful. The reader's weakest_assumption identifies exactly this, and the paper's own limitations acknowledge the single-DGM design. A focused simulation with an age-shifted or time-varying-bias external cohort would directly test whether the improvement result is robust to more realistic discrepancies. If the improvement disappears, the abstract and conclusions need substantial caveats; if it persists, the constant-bias simplification is not the limiting factor. Since the reader already recommends a conditional verdict on these grounds, my assessment does not change the verdict, but it sharpens the specific check needed.","tokens_in":18781,"tokens_out":6933,"duration_ms":48805,"concrete_test":"Re-run the simulation study with the same settings (N=400, 5-year follow-up, knots at 5/10/25 years, true Gompertz background mortality) but generate the external data from a systematically different population: e.g., baseline ages drawn from N(70,9) instead of N(60,9), or with a time-varying bias v(t)=log(1.2) for years 6-15 and v(t)=log(1.0) for years 16-25, while the fitted model still assumes the external data have the same survival as the trial control arm. Compare bias and MSE of 40-year control-arm RMST for trial-only (extra knots) versus trial + external data. If the external-data model does not reduce bias or increases MSE relative to trial-only, the abstract's improvement claim is overgeneralized; if it still improves, the constant-bias simplification is not the limiting factor.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's headline claim that 'even using moderately biased external data gives improvements over using the short-term trial data alone' is supported only for external data generated from the same DGM as the trial control arm, with a constant multiplicative bias exp(v) applied to the all-cause cumulative hazard (Supplementary Methods, 'External data was generated from the same underlying data generating mechanism...'; main text, Simulation Study). The model then 'assumed the external data have the same survival as the control arm of the trial' (main text, Bayesian survival models). Consequently, the only discrepancy explored is a time-invariant scaling of the hazard. Real-world registries and EHR data differ in case mix, age distribution, secular trends, and selection over time; these induce non-proportional and non-multiplicative discrepancies (e.g., higher background mortality in an older registry population, or bias that appears only after 10 years). Under such discrepancies, the flexible M-spline can distort the hazard shape rather than shift it uniformly, potentially producing RMST bias larger than that from trial-only analysis. The limitations section (Discussion) concedes the simulation 'depends on a particular data generating mechanism,' but the abstract and conclusions state the improvement result without this caveat. This is the load-bearing assumption for the paper's central practical message.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper evaluates the extrapolation performance of the survextrap flexible Bayesian survival model when incorporating external real-world data, using a head-and-neck cancer case study and a simulation study based on a Weibull mixture disease-specific hazard plus Gompertz other-cause mortality. The simulation compares models with and without external data under varying bias levels, knot placements, and treatment-effect assumptions, with estimands being 40-year restricted mean survival time (RMST) and its treatment difference. The main claims are that including long-term external data improves control-arm RMST extrapolation even under a constant multiplicative hazard bias of up to 20%, and that treatment-effect extrapolation requires explicit assumptions about waning. The paper reports bias, mean squared error, model posterior standard deviation, and credible-interval coverage, with Monte Carlo standard errors.","tokens_in":19032,"tokens_out":8226,"duration_ms":66400,"significance":"The paper is a timely and practically relevant evaluation of a widely used tool in health technology assessment. Its strengths include a data-generating mechanism that is structurally different from the M-spline model, true estimands computed from a very large simulated cohort, Monte Carlo errors reported for all performance measures, and public availability of code and data. If the central claims hold, the paper provides useful guidance on knot placement, external-data use, and treatment-effect sensitivity analysis. However, the external-data bias mechanism is limited to a constant multiplicative hazard shift, and one simulation result contradicts the general statement that external data always improves extrapolation accuracy. These issues need to be addressed before the conclusions can be fully accepted.","major_comments":[{"comment":"The external data are generated from the same underlying data-generating mechanism as the trial control arm, with only a constant multiplicative bias exp(v) applied to the all-cause hazard. This means the only discrepancy explored is a time-invariant proportional shift; real-world registries and EHR data can differ in case mix, age distribution, secular trends, and selection, which would induce non-proportional and non-multiplicative discrepancies. The abstract and Discussion state without this caveat that 'even using moderately biased external data gives improvements' and that incorporating external data 'still improved the quality of extrapolations'. The Discussion acknowledges the DGM dependence, but the headline claims are not qualified. Please either add simulations with structurally different external-data mechanisms (e.g., different age distributions or hazard-ratio bias that changes over time) or restrict the conclusions and abstract to the constant-bias setting.","section":"Simulation Study, Data Generating Mechanism; Supplementary Methods"},{"comment":"In Scenario 2 (waning treatment effect), the PH model with unbiased external data has bias 1.10, MSE 2.37, and coverage 0.83, whereas the same model without external data has bias 0.64, MSE 1.14, and coverage 0.89. This is a direct counterexample to the Discussion statement that incorporating external data 'still improved the quality of extrapolations in comparison to relying on trial data alone' when the treatment effect is of interest. The manuscript should discuss this exception and qualify the general claim about external data improving extrapolations.","section":"Table 3, Scenario 2; Discussion"}],"minor_comments":[{"comment":"The expressions for scenarios 2 and 3 are written as 'exp[β(t)] = ...' but the right-hand sides are negative or can be negative; presumably these are meant to be β(t) (the log hazard ratio) whose exponential is the hazard ratio. Also, the displayed 'tanh' formula is actually coth(x) = (e^x+e^{-x})/(e^x-e^{-x}); the standard tanh is (e^x-e^{-x})/(e^x+e^{-x}). Please correct these formulas or clarify the notation.","section":"Supplementary Table 2"},{"comment":"The main text states that survextrap version 0.8.12 was used, while the Supplementary Methods mention a feature available from version 0.8.16; please align the version numbers.","section":"Simulation Study; Supplementary Methods"},{"comment":"The caption says 'The vertical line shows the true value', but the plots display bias; please specify whether the vertical line indicates zero bias or the true estimand value.","section":"Figure 6 caption"},{"comment":"The phrase 'In case studies and simulations' is grammatically awkward; consider 'In a case study and simulations' or 'In case-study and simulation analyses'.","section":"Abstract"},{"comment":"The sentence 'we demonstrate the robustness of survextrap when modelling external data that was imperfect' should be 'when modelling external data that are imperfect'.","section":"Discussion"}],"recommendation":"major_revision","confidential_remarks":"The authors are developers of the survextrap package and are employed by the funder, though this is disclosed. The supplementary material contains mathematical errors in the DGM definitions (the tanh formula and the sign of the hazard ratio expressions) that should be corrected before publication; these do not necessarily invalidate the results if the code is correct, but they hinder reproducibility from the written description. The more substantive concern is that the external-data bias model is a single, constant multiplicative shift, and the paper's general conclusions should be tempered or supported by additional scenarios."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a careful simulation study evaluating survextrap's long-term extrapolation with external data, and the main practical claim—that even biased external data beats trial-only extrapolation—is supported for the bias mechanisms they actually test. The paper is a good contribution but the abstract states the result more generally than the evidence supports.\n\nWhat's new: it's the first formal simulation evaluation of survextrap's long-term extrapolation with external data, covering biased external data, knot placement, treatment-effect modeling, and waning. The DGM is structurally different from the M-spline (Weibull mixture plus Gompertz), the true estimands are computed from a very large simulated cohort, and they report Monte Carlo errors. They include scenarios where the model fails—separate-arms modeling does poorly, PH overestimates under waning—which counts in their favor. Code and data are on GitHub, and the methods section is detailed enough to reproduce.\n\nSoft spots: the external-data bias is only a constant multiplicative scaling of the all-cause hazard, generated from the same DGM as the trial control arm. Real-world registries differ in case mix, age distribution, secular trends, and sometimes bias that emerges after a lag, all of which can distort a spline's shape rather than shift it uniformly. The stress-test note is right that this is the load-bearing assumption for the abstract's general claim. The paper's limitations section does concede the simulation depends on a particular DGM, but the abstract and conclusions don't carry that caveat. I'd call this a mild overclaim rather than a fatal flaw, because the qualitative direction—more long-term data, even imperfect, helps anchor extrapolation—is plausible and consistent with the case study and with other work.\n\nSecond smaller issue: the abstract's claim about quantifying structural uncertainty when no long-term data are available is approximate, and the limitations section says so. The paper should say 'approximate' in the abstract too.\n\nThe citation pattern is fine; the authors cite the prior survextrap paper, Vickers, Guyot, and the relevant methods literature. The competing interests are declared.\n\nWho this is for: HTA modellers and methodologists working on survival extrapolation with real-world data. It deserves a serious referee. I'd send it to peer review and ask for a softened abstract and a discussion of more realistic external-data discrepancy mechanisms, plus maybe a small simulation with non-proportional bias if feasible. Not a desk reject.","headline":"A careful simulation study that supports the value of even biased external data for survextrap extrapolation, but the abstract overstates the generality of the bias mechanism tested.","tokens_in":19564,"tokens_out":1822,"would_cite":true,"duration_ms":13154,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F15","62N01","62P10"],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that adding long-term real-world survival data, even when moderately biased, yields less biased 40-year restricted mean survival estimates than trial data alone, provided the long-term treatment-effect assumption is…","keywords":["survival extrapolation","health technology assessment","Bayesian evidence synthesis","real-world data","M-spline hazard","restricted mean survival time","treatment effect waning","simulation study"],"falsifier":"Simulate or reanalyse a dataset where the external registry population has a diverging disease-specific hazard over time, for example a different case mix or a time-varying rather than constant hazard ratio relative to the trial control arm, and check whether including the external data still lowers bias in 40-year restricted mean survival; if it increases bias, the paper's central claim fails in that setting.","tokens_in":18561,"feed_emoji":"📊","tokens_out":6075,"duration_ms":47783,"temperature":0.7,"pith_summary":"This paper asks whether survival extrapolations used for health technology assessment can be made more accurate by feeding longer-term real-world data into a flexible Bayesian survival model. Using the survextrap model, whose hazard is an M-spline, the authors run a case study on head-and-neck cancer trial data and a simulation study with 1,000 trial datasets. They claim that adding long-term registry and population-mortality data gives accurate 40-year restricted mean survival estimates, and that even external data whose hazard is biased by plus or minus 20 percent improves on trial-data-only extrapolations. They also claim that treatment-effect differences are estimated accurately when the long-term treatment effect assumption, such as proportional hazards or an imposed waning schedule, is reasonably correct. Health technology assessment needs defensible long-term survival estimates from trials that are too short to observe them, so this is a direct test of a widely usable tool.","feed_headline":"Even 20%-biased real-world data improve 40-year survival estimates","feed_subtitle":"A flexible Bayesian spline anchors extrapolation with registry and life-table data, cutting long-term bias.","key_machinery":"The carrying mechanism is the M-spline hazard model implemented in the survextrap package: a hazard (or excess hazard) written as a weighted sum of positive cubic basis functions with a scale parameter, a smoothness prior, and knots placed at quantiles of trial event times. Extra knots placed beyond trial follow-up let the hazard vary in the long term, so the posterior from trial data alone expresses structural uncertainty rather than forcing a constant hazard. External registry data enter as aggregate counts of survivors over annual intervals, and population mortality enters as a fixed known background hazard in an excess-hazard (relative survival) framework. Treatment effects are modelled either as proportional hazards, as flexible non-proportional hazards with a hierarchical prior on time-varying coefficients, or by fitting arms separately, with optional treatment-effect waning that linearly shrinks the log hazard ratio to zero over a specified interval.","core_discovery":"The central discovery is that incorporating long-term external data into a flexible Bayesian evidence-synthesis model removes most of the bias in extrapolated long-term survival, whereas trial data alone, even with extra spline knots, remains biased and highly uncertain. In the main simulation, unbiased external data reduce the bias in 40-year control-arm restricted mean survival from about -0.96 to -0.02 years, and even plus-or-minus 20 percent biased external data outperform models without external data. The model also recovers the 40-year restricted mean survival difference between arms under a constant treatment effect when a proportional-hazards assumption is used, while separate-arm models fail because they contain no long-term information about the active arm. When the true treatment effect wanes after trial end, estimates improve only under waning assumptions close to the truth, showing that external data on the control arm alone cannot resolve treatment-effect uncertainty.","pith_inferences":["Implicit implication: if the constant-bias simulation mechanism is representative, then health technology assessments should treat the difference between trial-only and trial-plus-external-data extrapolations as a quantitative measure of structural uncertainty, not just a sensitivity check.","Testable extension: run the same simulation design with external data generated from a shifted age distribution, a different case mix, or a time-varying hazard bias; the current design only varies a constant hazard multiplier, so it cannot tell how robust the benefit is to those discrepancies.","Possible design: the paper mentions power priors and power likelihoods as ways to downweight suspicious external data; a natural comparison is fixed-weight inclusion versus power-likelihood weighting under the plus-or-minus 20 percent bias scenarios, with coverage of 40-year restricted mean survival as the outcome."],"forward_implications":["Health technology assessments can anchor extrapolations to 40-year restricted mean survival using registry and life-table data rather than relying on short-term trial data alone.","Moderately biased external data, with hazard rates up to 20 percent higher or lower than the truth, still reduce bias in long-term control-arm survival compared with trial-only models.","Without long-term data, extra knots beyond trial follow-up honestly reflect structural uncertainty instead of giving falsely narrow constant-hazard extrapolations.","Treatment-effect differences are trustworthy only under a correct long-term assumption, such as proportional hazards or an appropriate waning schedule; separate-arm modelling without active-arm long-term data is unreliable.","Including unbiased external data can compensate for shorter trial follow-up: three-year trial data plus external data gave unbiased 40-year restricted mean survival estimates, while eight-year follow-up was needed without it."],"supporting_citations":[{"why":"Defines the survextrap model and the flexible M-spline Bayesian evidence-synthesis approach that the paper evaluates.","marker":"[6]"},{"why":"Supplies the registry and life-table external data used in the case study.","marker":"[9]"},{"why":"Provides the head-and-neck cancer trial data that anchor both the case study and the simulation control-arm mechanism.","marker":"[16]"},{"why":"Describes the method used to reconstruct individual patient data from published Kaplan-Meier curves for the case study.","marker":"[17]"},{"why":"Previous simulation study by the same group establishing short-term fit of survextrap models, which this paper extends to long-term extrapolation.","marker":"[13]"},{"why":"Earlier simulation of flexible spline-based models with external data that this study builds on and improves in terms of flexibility and transparency.","marker":"[8]"},{"why":"Provides the mortality parameters used to generate other-cause survival in the simulation and population rates in the case study.","marker":"[18]"},{"why":"Supplies the cumulative hazard inversion method used to simulate survival times in the simulation study.","marker":"[21]"}],"fun_headline_variants":["Bayesian model with real-world data sharpens 40-year survival forecasts","Even biased external data beat trial-only survival extrapolation","Flexible Bayesian spline cuts long-term survival extrapolation bias","Real-world data, even 20% off, improve survival extrapolation"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The main load-bearing premise is that the external real-world data are generated by the same all-cause hazard process as the trial control arm, differing only by a constant multiplicative bias factor; if real-world discrepancies are more complex, the improvement from including external data may not carry over.","fun_headline_variants_meta":{"raw":{"variants":["Bayesian model with real-world data sharpens 40-year survival forecasts","Even biased external data beat trial-only survival extrapolation","Flexible Bayesian spline cuts long-term survival extrapolation bias","Real-world data, even 20% off, improve survival extrapolation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000513,"raw_usage":{"total_tokens":2514,"prompt_tokens":990,"completion_tokens":1524,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":606,"completion_tokens_details":{"reasoning_tokens":1451}},"tokens_in":606,"tokens_out":1524,"duration_ms":8653,"temperature":1.0,"reasoning_tokens":1451,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T14:54:01.684818+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate or reanalyse a dataset where the external registry population has a diverging disease-specific hazard over time, for example a different case mix or a time-varying rather than constant hazard ratio relative to the trial control arm, and check whether including the external data still lowers bias in 40-year restricted mean survival; if it increases bias, the paper's central claim fails in that setting.","supporting_citations":[{"cited_title":"survextrap: a package for flexible and transparent survival extrapolation","cited_arxiv_id":null,"evidence_quote":"Defines the survextrap model and the flexible M-spline Bayesian evidence-synthesis approach that the paper evaluates."},{"cited_title":"Extrapolation of Survival Curves from Cancer Trials Using External Information","cited_arxiv_id":null,"evidence_quote":"Supplies the registry and life-table external data used in the case study."},{"cited_title":"An Evaluation of Survival Curve Extrapolation Techniques Using Long - Term Observational Cancer Data","cited_arxiv_id":null,"evidence_quote":"Earlier simulation of flexible spline-based models with external data that this study builds on and improves in terms of flexibility and transparency."}],"review_version":1}