{"id":"58b3c161-f127-4159-9afe-9b9ad100fa95","arxiv_id":"2501.13525","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"low","formal_verification":"none","parameter_count":1,"one_line_summary":"Each subgroup gets its own smooth time-varying effect of a covariate on survival, fitted with penalized splines inside piecewise exponential additive mixed models.","lead":"This paper extends survival models so a drug's or biomarker's effect on risk can change over time in different ways for different patient groups, instead of assuming one fixed effect or one shared time trend. The authors show this improves fit on simulations and on brain tumor survival data, revealing effects that ordinary averages miss.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Exchangeability of the group random effects, not the spline construction, is the load-bearing premise; the simulations only exercise the model's own assumption.","rationale":"The reader's weakest assumption and my main concern coincide: the exchangeable random-effects structure in Eq. (4) is load-bearing for the paper's claim. The method's ability to capture heterogeneous time-variation depends on the grouping levels being exchangeable draws from a common distribution; the paper's own framing, however, explicitly targets fixed subpopulations such as gender or subdiagnoses, and the case study uses five diagnosis groups. With only five groups, the variance component is imprecisely estimated and shrinkage can distort the very heterogeneity the method aims to expose. The simulation study, while internally coherent, generates scenario (I) from the same functional random coefficient model, so it validates estimation under the assumed mechanism but not under a misspecified grouping structure. A fixed-effects competitor is absent from both the simulation and the case study, leaving a genuine gap in the evidence. I do not regard this as an internal inconsistency; the construction is mathematically sound and the Poisson-likelihood inference is standard. Rather, it is an external-validity limitation that a sensitivity analysis could settle. The AIC labeling and the first-to-propose novelty claim are legitimate concerns, but they do not threaten the central mechanism as directly as the exchangeability assumption. The reader's conditional verdict already captures this uncertainty, so I recommend no change: the paper should be accepted only if the authors add the sensitivity analysis or clearly restrict the claim to exchangeable grouping structures.","tokens_in":11590,"tokens_out":3857,"duration_ms":41847,"concrete_test":"Refit the brain tumor model replacing s(diagnosis, t, by=FGA, bs=\"fs\") with a fixed group-specific smooth interaction, e.g., s(t, by=diagnosis, bs=\"ps\") with per-smooth penalties or te(t, diagnosis, bs=c(\"ps\",\"re\"), by=FGA), and compare the estimated FGA curves and predictive performance using a holdout log-likelihood or 5-fold cross-validated Brier score. In addition, simulate scenario (I) with non-exchangeable fixed group curves (monotone increase, U-shape, flat, decrease) and estimate bias/coverage of f_g(t). If the functional random coefficient no longer outperforms fixed-smoothing competitors or shows substantial bias, the central claim must be qualified to exchangeable group structures.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that the functional random coefficient f_g(t) captures subgroup-specific time-varying covariate effects, yet the construction in Eq. (4) treats the grouping index g as an i.i.d. random intercept with a common variance and a shared smoothness penalty, implemented in mgcv via bs=\"fs\". This exchangeability assumption is natural for patient-level frailty or for groups sampled from a super-population, but the paper's advertised targets include fixed subpopulations, such as gender or subdiagnoses, and the case study uses five fixed diagnosis groups from TCGA. With G=5, the single variance component is weakly identified, and all five curves are shrunk toward a common zero curve. If the true diagnosis curves are systematically different, as the case study claims (FGA increasing for four diagnoses and decreasing for glioblastoma), the shrinkage is misspecified and the estimated curves are biased toward each other. The simulation evidence does not address this: scenario (I) generates the group curves from the same functional random coefficient family, so it only demonstrates recovery when exchangeability holds. No fixed-effects competitor, such as s(t, by=diagnosis, bs=\"ps\") or a tensor-product interaction with separate penalties, is fitted to the case study. Thus the claim that the method reliably reveals heterogeneous time-variation for fixed subgroups, rather than merely for exchangeable random groups, is not yet supported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes modeling subgroup-specific time-varying covariate effects in survival analysis via functional random coefficients of the form f_g(t) * x_ik, embedded in piecewise exponential additive mixed models (PAMMs). The functional random effect is implemented as a factor-smooth tensor product interaction using the mgcv specification s(g, t, by = x, bs = \"fs\"), with penalized P-splines providing regularization. The authors argue that this is the first unified framework for heterogeneously time-varying effects in hazard regression. They support the proposal with a simulation study comparing four nested models under four data-generating scenarios, and with a brain tumor case study examining the diagnosis-specific time-varying effect of fraction genome altered (FGA) on survival.","tokens_in":11739,"tokens_out":3702,"duration_ms":35933,"significance":"The methodological core is sound and practically valuable: the paper connects a recently proposed functional random coefficient construction to the well-established PAMM/Poisson-likelihood framework, and the implementation relies on existing, widely used software (mgcv, pammtools), which lowers the barrier for adoption. The simulation study is extensive (1000 repetitions, three sample sizes, four scenarios) and the authors appropriately acknowledge the one case of slight overfitting. The case study is clinically relevant and produces interpretable diagnosis-specific effect curves. If the exchangeability assumption of the group random effects is accepted, the contribution is a useful extension of PAMMs. However, the principal evidence for the central claim is a simulation whose data-generating process is the proposed model itself, and the case study uses five fixed diagnosis groups, so the paper does not yet demonstrate that the method is reliable when the exchangeability assumption is violated.","major_comments":[{"comment":"The simulation DGP in scenario (I) generates the group-specific time-varying effect f_g(t) from the same functional random coefficient family that the proposed model assumes. Therefore, scenario (I) only demonstrates that the estimator can recover curves generated under its own exchangeability assumption; it does not test the method's ability to capture heterogeneous time-variation for fixed subgroups, which are an explicit target in the Introduction (Section 1, final paragraph) and in the case study. I recommend adding a fixed-effects competitor such as s(t, by = diagnosis, bs = \"ps\") or a tensor-product smooth with diagnosis as a fixed factor, and adding a simulation scenario in which the true group curves are deterministic and systematically different across groups to assess shrinkage bias.","section":"Section 4.2, Eq. (5)"},{"comment":"With only G = 5 diagnosis groups, the single common variance component and shared smoothness penalty in Eq. (4) imply potentially strong shrinkage of all five curves toward a common mean, and the paper does not report the estimated variance component, the effective degrees of freedom, or pointwise confidence intervals for the five curves. Without a comparison to a fixed-effects specification, the conclusion that the FGA effect 'strongly varies' between diagnoses and that this variation is not a shrinkage artifact is not yet supported. Please report the variance estimate and the effective degrees of freedom, add pointwise confidence intervals to Figure 3, and fit a fixed-effects interaction model to the case study as a sensitivity check.","section":"Section 5, Figure 3 and Table 3"},{"comment":"The abstract states that 'using a penalized basis prevents overfitting in case of absence of such effects,' but the simulation results in scenario (III) show that the proposed model fits slightly better than the true nested model, which the authors acknowledge as slight overfitting. The suggested remedy of visual inspection is subjective and not quantified. This weakens the overfitting-prevention claim as stated. Please provide a quantitative measure, for example the proportion of simulation runs in which the functional random coefficient is estimated to have non-negligible time-variation under a null DGP, or an explicit discussion of how the abstract should be read in light of the scenario (III) finding.","section":"Section 4.3 and Abstract"}],"minor_comments":[{"comment":"Typo: 'the reader is refereed to' should be 'the reader is referred to.'","section":"Section 3"},{"comment":"Typo: 'implemented in in the R package mgcv' contains a duplicated 'in.'","section":"Section 4.1"},{"comment":"The conclusion contains 'the simulation study outlines the superior fit of your approach'; 'your' should be 'our.'","section":"Section 6"},{"comment":"The model labeled 'Heterogeneity and time-variation' is not fully specified in Section 5; please clarify its predictor structure, particularly whether it is the same as model (ii) in Section 4.2.","section":"Section 5, Table 3"},{"comment":"Adding pointwise confidence intervals to the estimated FGA curves would help the reader judge the statistical evidence for the between-diagnosis differences, which is currently only visually implied.","section":"Section 5, Figure 3"},{"comment":"The claim that penalization 'mostly solves' the choice of the number of intervals is not accompanied by a sensitivity analysis with respect to the number of intervals or the number of inner knots; a brief simulation or a discussion would make this statement more precise.","section":"Section 2.2 and 4.3"},{"comment":"The statement that this is the first proposal of heterogeneously time-varying covariate effects in hazard regression models should be tempered, since Hagemann et al. (2024) already propose the same effect type in conditional logit models; the novelty is the application to hazard regression, not the effect type itself.","section":"Section 1 and 3"}],"recommendation":"major_revision","confidential_remarks":"The paper is within the scope of the journal and the technical machinery is standard and correctly presented. The main weakness is that the simulation and case study do not challenge the exchangeability assumption that underpins the functional random coefficient; adding fixed-effects comparisons and a misspecified simulation scenario should determine whether a major revision is sufficient or whether the claim needs to be restricted to exchangeable groups. I see no reason for rejection if the authors can address the empirical support for the central claim."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear X,\n\nQuick take: this is a competent extension of Hagemann et al.'s functional random coefficients to PAMMs. The model term s(g, t, by=x, bs=\"fs\") is clearly specified, the PAMM/Poisson machinery is standard, and the simulation study shows what the authors claim: better fit when heterogeneous time-variation is present, and shrinkage toward the simpler model when it is not. Credit where due: the paper is honest about the slight overfitting in scenario (III), ships R code and data, and the brain tumor case study is a nice practical illustration of effects canceling across subgroups.\n\nThe main soft spot is the exchangeability assumption in Eq. (4). Treating the group index as an i.i.d. random intercept with a common variance and shared smoothness is natural for many random groupings (patients, centers), but the advertised targets include fixed subpopulations like gender or diagnoses, and the case study uses five fixed diagnosis groups. With G=5, the single variance component is weakly identified and all curves shrink toward a common zero curve. If the true curves are systematically different, the shrinkage can bias the very heterogeneity the method is meant to expose. The simulation does not address this: scenario (I) generates curves from the same functional random coefficient family, so it only demonstrates recovery when exchangeability holds. A fixed-effects competitor, such as s(t, by=diagnosis, bs=\"ps\") with separate penalties, is never fitted to the case study. That is the load-bearing gap.\n\nOther concerns are minor. Calling AIC an 'out-of-sample' approximation is sloppy; it is an in-sample criterion with a complexity penalty. The 'first to propose' claim is plausible within the cited literature but could be softened. The simulation being its own DGP is standard practice for methods papers, not a flaw by itself.\n\nNet: the paper is a genuine, if incremental, contribution. The central modeling idea holds up under the machinery it uses, but the authors should be pushed to show what happens when the groups are fixed and the exchangeability assumption is violated. A serious referee should ask for that sensitivity analysis.\n\nI'd send it to review. It is not a breakthrough, but it is a clean, citable method for applied survival analysts.\n\nBest.","headline":"A solid, incremental transfer of functional random coefficients to survival analysis; the machinery is standard, the simulation supports the main claims, but the exchangeability assumption for fixed subgroups is the real soft spot and is not stress-tested.","tokens_in":12346,"tokens_out":1920,"would_cite":false,"duration_ms":17482,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62N01","62N02","62G08"],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper introduces a hazard-regression term $f_g(t)\\cdot x$ that gives each subgroup its own non-linear, time-varying covariate effect, and shows it fits better than nested alternatives while shrinking when the effect is absent.","keywords":["survival analysis","hazard regression","piecewise exponential additive mixed models","time-varying covariate effects","functional random effects","factor smooths","non-proportional hazards","penalized splines"],"falsifier":"Simulate data in which the four group effects are fixed, widely separated curves with no shared variance structure, fit both the functional random coefficient and a fixed group-by-time interaction, and check whether the random-effect version visibly shrinks the outer curves toward a common mean; if it does, the exchangeability assumption fails in exactly the cases the method is meant to expose.","tokens_in":43,"feed_emoji":"⏳","tokens_out":5952,"duration_ms":113222,"temperature":0.7,"pith_summary":"Survival analysts often need covariate effects that change over time and differ across subgroups, sometimes with the time pattern itself differing across subgroups. The paper proposes to capture that third layer—heterogeneously time-varying effects—by adding a functional random coefficient to a piecewise exponential additive mixed model. The term $f_g(t)\\cdot x_{ik}$ lets each group have its own non-linear time curve for a covariate's effect, built as a tensor-product interaction of a random group effect and a penalized spline in time. Simulations show this term fits better when such effects exist and is penalized toward the simpler nested model when they do not; the brain-tumor example finds that fraction genome altered affects survival with direction and timing that differ by diagnosis.","feed_headline":"Hazard model lets covariate effects differ by group—and change over time","feed_subtitle":"A functional random coefficient learns each group's time curve and shrinks to simpler models when no such curve exists.","key_machinery":"The central object is the functional random coefficient, implemented as $s(g, t, \\text{by}=x, \\text{bs}=\\text{fs})$: a functional random effect (factor smooth) used as a varying coefficient. It is an anisotropic tensor-product interaction of a group-index random effect (indicator basis with identity penalty) and a penalized B-spline in time, built with the numerically stable reparameterization of tensor-product smooths. This object carries the argument because it is the only term in the predictor that lets a covariate's effect have both a group-specific level and a group-specific non-linear time curve, while its penalty structure automatically shrinks the term toward simpler nested alternatives.","core_discovery":"The central claim is that a piecewise exponential additive mixed model (PAMM) can include covariate effects of the form $f_g(t)\\,x_{ik}$, where $g$ indexes a subgroup and $t$ is time, using a functional random effect—a tensor product of an i.i.d. random intercept in $g$ and a P-spline in $t$. Because the effect is a separate curve per group, it can capture not only subgroup-specific time variation but subgroup-specific changes in that time variation, something no single existing survival framework did. The paper argues that the penalized spline construction keeps the term from overfitting when the true effect is simpler, shrinking toward the main-effect-plus-time model, and demonstrates the point with four simulation scenarios and a diagnosis-specific analysis of fraction genome altered in gliomas. On the case study, modeling FGA with this term improves fit over nested alternatives and reveals effects that would cancel out if the time variation were pooled across diagnoses.","pith_inferences":["Inference: the random-effects assumption behind the group term is the part most in need of stress-testing; fixed subgroups with systematically different curves may be over-shrunk, so a fixed-effects version or a prior that allows curve similarity is a natural comparison.","Inference: the construction could transfer to settings where the grouping variable is not categorical—for example a continuous moderator—by replacing the random intercept in $g$ with a smooth function of that moderator.","Inference: one testable extension is to use the fitted group curves to build a clustering or classification rule, since the method yields one full hazard-effect trajectory per subgroup that could feed downstream analyses."],"forward_implications":["If correct, researchers can fit non-proportional hazard models where a covariate's effect curve is allowed to differ per subgroup without specifying which subgroup has which curve; the data choose.","The penalization means the same software can serve as a model-selection device: when heterogeneous time-variation is absent, the term shrinks toward the simpler model, so choosing interval cut-points or testing for such effects becomes less delicate.","The brain-tumor analysis suggests that pooling diagnosis-specific time-varying effects can hide real signals—FGA's effect declines for glioblastoma but rises for other diagnoses—so a single average time-varying coefficient can be misleading.","Because the term lives inside the Poisson-likelihood reparameterization of PAMMs, it inherits existing GAM software and can be applied to large survival datasets."],"supporting_citations":[{"why":"Establishes piecewise exponential additive mixed models and the Poisson-likelihood equivalence that the proposed term is added to.","marker":"Bender et al. (2018)"},{"why":"Supplies the pammtools software for the piecewise exponential data transformation used in the implementation.","marker":"Bender and Scheipl (2018)"},{"why":"Defines functional random effects as tensor-product interactions of smooth and random effects, the basis for the proposed coefficient.","marker":"Kneib et al. (2019)"},{"why":"Introduces functional random coefficients in conditional logit models, which this paper generalizes to hazard regression.","marker":"Hagemann et al. (2024)"},{"why":"Gives the numerically stable tensor-product reparameterization underlying the t2/fs basis used for estimation.","marker":"Wood et al. (2013)"},{"why":"Provides the REML-based smoothing parameter estimation used to fit the model without a full mixed-model formulation.","marker":"Wood (2011)"},{"why":"Describes the algorithm used to simulate survival times in the simulation study.","marker":"Bender et al. (2005)"},{"why":"Provides the brain tumor dataset used in the case study.","marker":"Ceccarelli et al. (2016)"}],"fun_headline_variants":["Survival model lets covariate effects vary by group and time","New hazard model learns group-specific time-varying effects","PAMM captures heterogeneous time variation in covariate effects","Flexible spline model handles subgroup-specific time-varying effects","Survival analysis with group-specific time curves for covariates"],"cache_read_input_tokens":14464,"weakest_assumption_plain":"The method treats the subgroup labels as if they are random draws from one common distribution with a shared spread and shared smoothness, so each subgroup's curve is pulled toward the average rather than estimated as its own fixed curve.","fun_headline_variants_meta":{"raw":{"variants":["Survival model lets covariate effects vary by group and time","New hazard model learns group-specific time-varying effects","PAMM captures heterogeneous time variation in covariate effects","Flexible spline model handles subgroup-specific time-varying effects","Survival analysis with group-specific time curves for covariates"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000217,"raw_usage":{"total_tokens":1469,"prompt_tokens":1010,"completion_tokens":459,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":626,"completion_tokens_details":{"reasoning_tokens":380}},"tokens_in":626,"tokens_out":459,"duration_ms":4474,"temperature":1.0,"reasoning_tokens":380,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T15:52:00.649758+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate data in which the four group effects are fixed, widely separated curves with no shared variance structure, fit both the functional random coefficient and a fixed group-by-time interaction, and check whether the random-effect version visibly shrinks the outer curves toward a common mean; if it does, the exchangeability assumption fails in exactly the cases the method is meant to expose.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines functional random effects as tensor-product interactions of smooth and random effects, the basis for the proposed coefficient."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces functional random coefficients in conditional logit models, which this paper generalizes to hazard regression."},{"cited_title":"N., Scheipl, F., and Faraway, J","cited_arxiv_id":null,"evidence_quote":"Gives the numerically stable tensor-product reparameterization underlying the t2/fs basis used for estimation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the REML-based smoothing parameter estimation used to fit the model without a full mixed-model formulation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Describes the algorithm used to simulate survival times in the simulation study."},{"cited_title":"P., Malta, T","cited_arxiv_id":null,"evidence_quote":"Provides the brain tumor dataset used in the case study."}],"review_version":1}