{"id":"91041973-a605-44b5-b366-0a71b3f56c56","arxiv_id":"2509.10383","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"A new M-spline network meta-analysis model with a weighted random walk shrinkage prior handles flexible, non-proportional survival hazards without needing model selection.","lead":"This paper proposes a flexible Bayesian network meta-analysis model for survival data, using M-splines for the baseline hazard and a weighted random walk prior to prevent overfitting. It allows non-proportional hazards between treatments and is implemented in the R package multinma, demonstrated on progression-free survival data from non-small cell lung cancer trials.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The practical claim of knot/timescale invariance rests on a prior-predictive figure, not on posterior evidence; even if the prior were invariant, posterior estimates can still depend on the chosen M-spline space, and the non-proportionality random walk is never checked.","rationale":"The reader's weakest assumption is the invariance of the weighted random walk prior (Eqs. 5–9). I agree that this is the load-bearing part of the central claim, but I would sharpen the concern in two ways. First, the paper's own evidence for invariance is prior predictive only; the abstract's workflow claim is about posterior estimates, and prior invariance is not sufficient for posterior knot-insensitivity because the M-spline approximation space changes with the knots. Second, the non-proportionality effects γ_k are the main new modelling feature, yet no prior or posterior check of their invariance is presented; Figure 4 only covers the baseline hazard prior. The case study's 8-vs-11-knot comparison is a single dataset and cannot establish the general claim, especially since the authors had to manually adjust the default knots for sampling reasons. These concerns do not invalidate the method, but they make the central marketing of the paper—'invariant to knots and timescale'—overstated relative to the evidence. The appropriate verdict remains conditional: the method is promising and reproducible, but the central invariance/practical-insensitivity claim needs either a formal argument or a systematic posterior check across knot configurations. This does not move the reader's verdict, so I mark it unchanged.","tokens_in":17691,"tokens_out":21524,"duration_ms":264281,"concrete_test":"Using the provided code/data, rerun model (11) on the NSCLC data with L = 6, 8, 11, and 15 internal knots under at least two defensible placement rules (quantiles of pooled event times; per-study quantiles combined; the paper's manually adjusted vector). Compare posterior medians and 95% intervals for landmark survival (12, 24, 52 weeks) and marginal log hazard ratios over time. Pre-specify a tolerance, e.g. 0.05 absolute survival or 0.1 log HR. If estimates differ beyond this, the knot-invariance claim is not supported in a realistic setting. Also run a simulation with a known non-PH hazard (e.g., delayed effect) over the same knot choices; this isolates prior vs likelihood contributions to any observed sensitivity.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that the weighted random walk prior makes the practitioner's choice of knots/timescale irrelevant ('simply needs to specify a sufficiently-large number of knots'). The main support is Figure 4, which shows prior predictive distributions of the baseline hazard, not posterior estimates. This is a logical gap: even a perfectly invariant prior does not imply posterior insensitivity, because changing L or knot locations changes the M-spline function space. The likelihood can only represent structure that the local basis can express; shrinkage cannot create flexibility where the basis has no support. Hence posterior survival curves, time-varying hazard ratios, and LOOIC can still depend on knots when the true hazard has features not captured by a particular knot sequence or when follow-up is sparse in some intervals. The only posterior evidence is one NSCLC dataset comparing 8 vs 11 knots (Table A.1, Fig. A.7), and in that same case study the default knot vector had to be manually modified to avoid sampling problems (Fig. A.2 vs A.3), undercutting the 'no tuning' message. Moreover, Figure 4 does not examine the multivariate random walk on the non-proportionality effects γ_k (Eq. 12), so the primary new capability has no invariance check at all. The prior invariance itself is also asserted rather than derived, and Appendix A.2 concedes it is only approximate for piecewise exponential hazards; applying Li & Cao weights through the inverse-softmax transform is a heuristic. If any of these sensitivities materialize, the headline benefit (no model selection) fails.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops a Bayesian network meta-analysis (NMA) model for survival outcomes in which the baseline hazard is modelled with M-splines and a weighted random walk prior is placed on the inverse-softmax transformed spline coefficients. The prior is designed to shrink towards a constant baseline hazard and to be invariant to the number, location, and timescale of knots, so that the analyst need only choose a sufficiently large number of knots. Non-proportional hazards are handled either by stratifying the baseline hazard by treatment arm or by adding treatment-specific coefficients to the spline coefficients under a symmetric multivariate random walk prior. The method is implemented in the multinma R package and applied to a four-study network of progression-free survival in non-small cell lung cancer, with model comparison by LOOIC and an assessment of sensitivity to the number of knots.","tokens_in":18102,"tokens_out":3339,"duration_ms":39217,"significance":"If the claimed invariance holds, this is a practically important contribution: it would remove model selection from flexible survival NMA and make non-proportional-hazards models routinely usable in Bayesian decision-making contexts. The manuscript is strengthened by a fully implemented R package, reproducible code and data, a realistic case study, and careful LOOIC-based model comparisons. The prior-predictive simulations in Figure 4 are informative and clearly demonstrate a deficiency in previously proposed priors. However, the central invariance claim is supported mainly by prior simulation rather than by analytic derivation or posterior sensitivity analysis, and the non-proportionality part of the model is not checked for invariance. These gaps are load-bearing for the paper's main practical message and need to be addressed before the claims can be accepted as stated.","major_comments":[{"comment":"The invariance of the prior to knot number, location, and timescale is demonstrated only by prior-predictive simulation, not derived analytically. More importantly, a prior-invariant property does not imply posterior insensitivity: changing the knot vector changes the M-spline function space, and the likelihood can only represent structure expressible by that basis. The single posterior comparison in the case study (8 vs 11 knots, Table A.1 and Fig. A.7) is reassuring but limited to one dataset and does not vary timescale or knot locations. I recommend adding a posterior sensitivity study, either simulated or on the case study, varying L, knot placement rules, and time units, and reporting differences in survival curves, time-varying hazard ratios, and LOOIC. Without such evidence, the abstract's unqualified claim of invariance is overstated.","section":"Random walk shrinkage prior, Eqs. (5)-(9), Fig. 4"},{"comment":"The multivariate weighted random walk prior on the non-proportionality effects gamma_k is the main new capability of the paper, yet no invariance check, either prior or posterior, is provided for these parameters. The weights and normalisation in Eq. (9) were designed for the baseline hazard coefficients, and their behaviour on the inverse-softmax treatment effect vectors is not self-evident. I ask for at least a prior-predictive figure analogous to Fig. 4 for the implied time-varying hazard ratios under varying knot configurations and timescales, and ideally posterior sensitivity results for the gamma_k estimates in the case study.","section":"Treatment effects on spline coefficients, Eq. (12)"},{"comment":"The paper concedes that for degree-zero M-splines (piecewise exponential hazards) the invariance is only approximate, but it does not quantify how close the approximation is. Since piecewise exponential hazards are a common practical choice, the abstract's blanket statement that the weighted random walk prior 'is invariant to the choice of knots and timescale' is too strong. Please either prove the invariance for kappa=1 under the modified weights (Eq. A.4) or state the qualified claim in the abstract and assess the approximation error across uneven knot placements and short intervals.","section":"Appendix A.2"},{"comment":"The default common knot vector produced by the proposed quantile-based rule led to sampling problems, and the authors had to manually add a knot at the end of follow-up in the ERACLE study. This manual adjustment undercuts the practical message that the analyst simply needs to specify a sufficiently large number of knots with no tuning. The discussion acknowledges this limitation, but it deserves to be raised as a major point because it directly affects the central usability claim. A robust automatic knot-placement algorithm, or at least a clear warning in the abstract and methods about when manual adjustment is needed, would resolve this.","section":"Case study, Fig. A.2 vs A.3"}],"minor_comments":[{"comment":"Typo: 'symetrically' should be 'symmetrically'.","section":"Treatment effects on spline coefficients"},{"comment":"In the text comparing ERACLE with the other study, 'PROSPECT' appears to be a typo for 'PRONOUNCE'.","section":"Case study"},{"comment":"The notation \\eta_{jk}(\\mathbf{x}_{ijk}) uses subscript i, but the linear predictor was defined without an individual subscript. Please clarify how individual-level covariates enter the log hazard rate, especially when only aggregate data are available.","section":"Including covariates, Eq. (13)"},{"comment":"The claim that this is 'the first time that M-splines have been incorporated within the NMA framework' may be stronger than necessary and could be softened, given prior work on M-splines in related multilevel settings (e.g., Phillippo et al.) and the possibility of unpublished or parallel work.","section":"Discussion"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is within the journal's scope and the applied machinery is sound. The main issue is a mismatch between the strength of the invariance claim and the evidence presented: the proof is lacking, the posterior sensitivity evidence is thin, and the non-proportionality feature is not checked. The requested additions are feasible within the paper's framework and would substantially raise confidence. I do not see grounds for rejection, provided the authors are willing to either prove the invariance in an appropriate setting or qualify the claims and add the missing sensitivity analyses."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dan — here's my take. The method is a real advance: first NMA built on M-splines, with a normalized weighted random walk prior on the inverse-softmax coefficients, and it ships in multinma with code and data. That combination is what makes the paper worth reading, not the theory alone. The case study is well done: three models compared with LOOIC within studies, survival curves and hazard ratios shown, and the non-proportional hazards model with treatment effects gives sensible results. The authors are honest about limitations: no loops so no inconsistency check, only fixed effects, reconstructed IPD used without accounting for reconstruction uncertainty. That last one is a real limitation, not fatal—most applied NMA of reconstructed data does the same—but it's worth flagging.\n\nWhere the paper oversells: the abstract and discussion claim the prior is 'invariant to the choice of knots and timescale.' What's actually shown is prior predictive simulation for a few knot configurations (Figure 4), with only approximate invariance for degree-zero splines. The posterior could still depend on the M-spline space; more knots means a richer basis, and shrinkage can't create flexibility the basis doesn't have. The case study shows only 8 vs 11 knots, and the common knot vector had to be manually adjusted (Figure A.2 vs A.3) to avoid sampling problems. That undercuts the 'simply specify a large enough number of knots' message. Also, the multivariate random walk on non-proportionality effects is never checked for knot/timescale sensitivity at all. None of this kills the method—it's still practical and likely works well with reasonable knot choices—but the invariance claim should be softened to 'approximately invariant in the prior, robust in practice' or backed with posterior evidence across more knot placements and a derivation for the normalized weights.\n\nThe stress-test note about the logical gap between prior invariance and posterior invariance is fair. I would ask the authors to prove or substantially qualify the invariance claim, and to add a small simulation study varying knots and timescale for the non-proportional effects.\n\nWho gets value: applied statisticians in HTA doing survival NMA, and methodologists working on P-spline priors. Worth a serious referee; conditional acceptance is the right call. I'd cite it and bring it to our reading group.","headline":"A genuinely useful M-spline NMA, well implemented and tested, but the headline claim of knot/timescale invariance is stronger than the evidence supports.","tokens_in":18587,"tokens_out":1623,"would_cite":true,"duration_ms":17959,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F15","62N01","62P10"],"pacs":[],"model":"deepseek-v4-flash","headline":"A weighted random-walk prior makes flexible M-spline network meta-analysis insensitive to knot choice and timescale, so analysts can fit one large-knot model and let shrinkage do the work.","keywords":["network meta-analysis","survival analysis","non-proportional hazards","M-splines","random walk prior","shrinkage","Bayesian inference","multinma"],"falsifier":"Run the M-spline NMA on the same dataset with two very different knot placements (e.g., quantile-based vs. evenly spaced) and with 5 vs. 20 internal knots, then compare posterior estimates of the baseline hazard and treatment effects; material differences would indicate the invariance claim fails. Alternatively, compute the prior covariance function of the log hazard under the weighted random-walk prior for two different knot vectors with the same number of knots; if the covariances differ, the prior is not strictly invariant.","tokens_in":17633,"feed_emoji":"📊","tokens_out":3465,"duration_ms":37995,"temperature":0.7,"pith_summary":"The paper proposes a Bayesian network meta-analysis model for survival outcomes that uses M-splines to flexibly model the baseline hazard, and introduces a novel weighted random-walk prior on the spline coefficients. The central claim is that this prior makes the model invariant to the number and placement of knots and to the timescale, so the analyst only needs to specify a sufficiently large number of knots and the prior shrinks away unnecessary complexity. The paper also shows how to relax proportional hazards by adding treatment effects on the spline coefficients, and demonstrates the method on progression-free survival in non-small cell lung cancer. If correct, this removes the model-selection burden of fractional polynomials and makes flexible non-proportional hazards NMA practical for routine Bayesian analysis.","feed_headline":"One Bayesian prior removes knot tuning from survival meta-analysis","feed_subtitle":"M-spline NMA with a weighted random-walk prior handles non-proportional hazards and needs only a large-enough knot count.","key_machinery":"The central mechanism is the M-spline basis for the baseline hazard (integrated to I-splines for survival), with coefficients mapped through a softmax transform to the unit simplex. A weighted random-walk prior is placed on the inverse-softmax coefficients, with step variances weighted by normalized inter-knot distances (Equation 9). These weights make the prior on the baseline hazard invariant to knot spacing, knot count, and timescale, while centering the prior on a constant hazard provides shrinkage. For non-proportional hazards, a multivariate weighted random-walk prior on treatment-specific spline-coefficient effects plays the same role, symmetrically across treatments.","core_discovery":"On the paper's own terms, the discovery is that a weighted random-walk prior on inverse-softmax transformed M-spline coefficients induces a prior on the baseline hazard that is invariant to knot locations, number of knots, and timescale. The prior is centered on a constant hazard and provides shrinkage, preventing overfitting even when many knots are used. This invariance is achieved by normalizing the random-walk step variances by inter-knot distances and total follow-up time. The paper extends the same prior to treatment effects on spline coefficients, yielding a symmetric non-proportional hazards model that shrinks toward proportional hazards when data are sparse, and it demonstrates that","pith_inferences":["If the invariance property holds generally, it is a portable tool: the same weighted random-walk prior could be applied to other Bayesian spline-based survival models outside network meta-analysis, such as single-arm extrapolation models or joint modeling.","The symmetry of the non-proportionality effects suggests a way to define consistency equations that do not depend on the choice of reference treatment; this could inspire new diagnostic tools for networks with multi-arm trials.","A formal analytic derivation of the prior's invariance—for instance, computing the induced prior covariance of the log hazard between two time points under different knot vectors—would turn the simulation-based evidence into a theorem and could provide guidance on how many knots are 'large enough'.","The method's local fit property (a spline fit between knots is informed only by data up to a few knots later) could be exploited to combine RCT evidence with external long-term data smoothly, potentially improving extrapolation in health technology assessment."],"forward_implications":["Analysts no longer need to run model selection over fractional polynomial powers; a single M-spline model with a sufficiently large number of knots suffices.","Bayesian fitting is tractable because the M-spline formulation ensures the cumulative hazard is monotonically increasing by construction, avoiding the sampling difficulties of Royston-Parmar models.","Non-proportional hazards can be modeled with treatment effects on spline coefficients, enabling predictions for all treatments in any target population, with automatic shrinkage toward proportional hazards where data are sparse.","The method is implemented in the multinma R package, which supports aggregate data, individual participant data, or mixtures of both, making the approach readily available.","Extrapolation beyond the observed follow-up shrinks toward a constant proportional hazard, avoiding the unrealistic polynomial tails of fractional polynomial models."],"fun_headline_variants":["One prior, any knot count: survival NMA simplified","Knot-free flexibility: new prior for survival networks","Survival NMA with a prior that ignores knots","Survival NMA: invariant prior ends knot selection","Adaptive Bayesian prior for non-proportional hazards"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The invariance of the weighted random-walk prior to knot placement and timescale is the load-bearing premise; it is demonstrated only by prior predictive simulation for selected configurations, not derived analytically, and is acknowledged to be only approximate for degree-zero M-splines.","fun_headline_variants_meta":{"raw":{"variants":["One prior, any knot count: survival NMA simplified","Knot-free flexibility: new prior for survival networks","Survival NMA with a prior that ignores knots","Survival NMA: invariant prior ends knot selection","Adaptive Bayesian prior for non-proportional hazards"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001173,"raw_usage":{"total_tokens":4716,"prompt_tokens":800,"completion_tokens":3916,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":544,"completion_tokens_details":{"reasoning_tokens":3840}},"tokens_in":544,"tokens_out":3916,"duration_ms":31784,"temperature":1.0,"reasoning_tokens":3840,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T17:51:45.816264+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the M-spline NMA on the same dataset with two very different knot placements (e.g., quantile-based vs. evenly spaced) and with 5 vs. 20 internal knots, then compare posterior estimates of the baseline hazard and treatment effects; material differences would indicate the invariance claim fails. Alternatively, compute the prior covariance function of the log hazard under the weighted random-walk prior for two different knot vectors with the same number of knots; if the covariances differ, the prior is not strictly invariant.","supporting_citations":[],"review_version":1}