{"id":"0a49d443-1817-4359-be20-84295a00644a","arxiv_id":"2607.09577","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.5,"correctness_risk":"low","formal_verification":"none","parameter_count":4,"one_line_summary":"Jackknife pseudo-RMST targets let multiple tabular foundation-model backbones do horizon-specific censored survival regression without survival-specific fine-tuning, with competitive recovery and real-data performance.","lead":"SurvFM-RMST turns right-censored survival outcomes into jackknife pseudo-observation labels for restricted mean survival time so ordinary tabular foundation models can predict event-free time without survival-specific training. It matters because many clinical datasets are modest, heterogeneous tables with censoring, and reusable tabular models currently assume fully observed labels.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified beyond the reader's already-flagged finite-sample pseudo-observation premise.","rationale":"The reader's strongest claim matches the paper's actual contribution: a target-side interface, not a new survival foundation model. The same-backbone ablation plus known conditional RMST is the right primary evidence; public-data results correctly support competitiveness within close bands rather than dominance. The weakest assumption the reader names (Methods §4.3 / Eq. 7) is indeed the softest load-bearing premise, and the manuscript already bounds its scope (independent/non-informative censoring, finite-sample regularity, horizon-specific static RMST). I did not find a more central flaw—e.g., leakage, mis-specified estimand, or circular ranking—that would force a harsher verdict. CONDITIONAL remains appropriate pending archival code/data and broader external validation; no adjustment is warranted from this stress pass.","tokens_in":22284,"tokens_out":604,"duration_ms":5698,"concrete_test":"Re-run the SurvFM-TabPFN v3 target ablation (Fig. 2d protocol) under informative censoring (C depends on a covariate that also enters the event hazard) and under n=50 high-censoring blocks; if oracle conditional-RMST RMSE for pseudo-RMST no longer improves on min(Y,τ) by the reported margins (~0.04 at tau80/tau90), the finite-sample interface claim weakens in the regimes the Discussion already flags.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that jackknife pseudo-RMST targets form a portable interface letting frozen tabular foundation-model backbones do horizon-specific RMST regression without survival-specific fine-tuning. That claim is internally supported by (i) known-truth oracle conditional-RMST recovery, (ii) a same-backbone target ablation showing pseudo-RMST beats min(Y,τ) and event-only labels (Results 2.2, Fig. 2d; Methods 4.3, Eq. 7), and (iii) competitive SurvSet discrimination/RMST-error bands. The classical approximation E{Zi(τ;D)|Xi=x}=μ\tau(x)+op(1) under independent/non-informative censoring is the softest premise, but the authors already state it, clip targets to [0,τ], and show the ablation gains. No stronger internal inconsistency, circular evaluation, or hidden assumption that would overturn the portability claim was found. Remaining limits (static scope, eligible SurvSet subset, fixed configs, incomplete public artifacts) are scope and reproducibility caveats already reflected in a CONDITIONAL verdict, not load-bearing failures of the interface argument.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The manuscript introduces SurvFM-RMST, a target-interface framework that converts right-censored survival outcomes into jackknife pseudo-observation labels for restricted mean survival time (RMST) at a fixed horizon τ, then feeds those labels to interchangeable tabular foundation-model backbones (TabPFN-v3, TabICL, TabDPT, TabH2O, MITRA) for ordinary horizon-specific regression without survival-specific fine-tuning or new output heads. Controlled simulations with known conditional RMST (24 settings, five replicates) show accurate oracle recovery and a same-backbone ablation in which pseudo-RMST targets reduce oracle RMSE and improve Uno’s C-index relative to min(Y,τ) and event-only labels. On 36 eligible static SurvSet datasets the SurvFM backbones are competitive with RMST-regression and classical survival comparators on discrimination and tau-normalized IPCW RMST error, with performance varying by horizon, endpoint and practical constraints; predicted-RMST tertiles also separate held-out Kaplan–Meier trajectories and enrich events. The authors position the contribution as a portable censoring-aware interface rather than a universally superior survival model.","tokens_in":22625,"tokens_out":1031,"duration_ms":11848,"significance":"If the results hold, the paper supplies a practical, theory-grounded route for applying existing tabular foundation models to right-censored biomedical data without redesigning their architectures or losses. The combination of known-truth oracle recovery, a clean same-backbone target ablation, multi-backbone public benchmarks with explicit denominators, and transparent practical profiles (column limits, runtimes, reduced-training ranks) is stronger evidence than is typical for early foundation-model survival adaptations. The work is complementary to survival-native and representation-side approaches and is likely to be useful to methodologists and applied groups that already rely on TabPFN-style tools. Scope limitations (static covariates, fixed horizons, eligible SurvSet subset, fixed hyperparameters) are stated clearly and do not erase the portability claim.","major_comments":[{"comment":"Methods §4.3, Eq. (7) and the target-ablation design (Results 2.2, Fig. 2d): the central portability claim rests on the classical approximation E{Zi(τ;D)|Xi=x}=μτ(x)+op(1) under independent/non-informative censoring. The same-backbone ablation already shows clear gains over naive targets, and the authors clip labels to [0,τ], but the manuscript still lacks a finite-sample diagnostic (e.g., bias of leave-one-out RMST as a function of risk-set size or censoring rate) that would let readers judge when the approximation is adequate. A short simulation or real-data sensitivity panel quantifying this would strengthen the load-bearing premise without changing the framework.","section":null},{"comment":"Results 2.4–2.5 and Methods §4.11–4.12: public-data conclusions are drawn from an eligible static SurvSet subset (36 of 77 screened) plus a limited tau80 scalability sensitivity on 11 scale-excluded datasets. While denominators and non-imputation are reported carefully, the main text still risks over-generalization (“competitive \\ldots across 36 eligible static SurvSet datasets”). Explicitly framing the benchmark as a static, size- and dimensionality-filtered subset in the abstract and Results opening, and stating that dynamic/longitudinal/competing-risk settings remain out of scope, would keep the claim proportionate to the evidence.","section":null}],"minor_comments":[{"comment":"Figure 2a caption notes a zoomed error axis; the main text should also flag that full-range tails appear only in Supplementary Fig. 1a so readers do not misread dispersion.","section":null},{"comment":"Methods §4.5–4.6 list package versions and fixed hyperparameters; a single compact table of all backbone and comparator configurations would improve reproducibility scanning.","section":null},{"comment":"The distinction between oracle conditional-RMST RMSE and realized restricted-time error (Eqs. 11–12) is clear in Methods but could be restated briefly when first used in Results 2.2.","section":null},{"comment":"Supplementary Table 6 design-choice audits are appropriately scoped as non-main, but a one-sentence pointer in the Discussion would help readers locate them.","section":null},{"comment":"Minor notation consistency: both Zi(τ;D) and zi(τ) appear; standardizing on one form would reduce friction.","section":null}],"recommendation":"minor_revision","confidential_remarks":"The paper is a solid methods contribution for a statistics/ML journal that values practical interfaces and careful benchmarking. The concurrent TabSA and SurvivalPFN literature is cited and positioned fairly; I see no novelty or citation-pattern concerns that would affect the editorial decision. Fit with a methods-oriented venue is good; a pure clinical-prognosis journal would require stronger individual-calibration and prospective claims that the authors correctly avoid."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The one thing worth knowing is that this is a clean target-side fix, not a new survival architecture. They convert right-censored outcomes into jackknife pseudo-RMST labels so ordinary tabular foundation-model regressors (TabPFN-v3, TabICL, TabDPT, TabH2O, MITRA) can do horizon-specific restricted event-free-time prediction without survival losses or fine-tuning.\n\nWhat is actually new is the packaging and the evidence, not the ingredients. Pseudo-observation RMST regression is classical (Andersen et al.); tabular foundation models are recent. The contribution is showing that the same pseudo-RMST interface is portable across several frozen backbones, with known conditional-RMST recovery in 24 simulation settings and a same-backbone ablation where pseudo-RMST beats min(Y,τ) and event-only labels on both oracle RMSE and Uno C. Public results on 36 eligible static SurvSet datasets put several SurvFM backbones in close leading bands with RSF, AFT, DeepSurv and RMST-regression baselines, with explicit denominators for TabH2O column limits and non-evaluable IPCW blocks. They also report reduced-training ranks, practical runtime/coverage profiles, and held-out tertile stratification with ordered KM gaps. Citation pattern is appropriate; concurrent survival-native and representation-side work is positioned as complementary rather than ignored.\n\nSoft spots are real but proportionate. The load-bearing premise remains the classical finite-sample pseudo-observation approximation under independent censoring (their Eq. 7); they state it, clip targets to [0,τ], and the ablation helps, but informative censoring or tiny risk sets would still break the labels. Scope is static baseline covariates and fixed horizons; the SurvSet subset is eligibility-filtered, not the full repository; configs are fixed rather than exhaustively tuned; public code/data are promised rather than fully released. None of that overturns the portability claim.\n\nThis is for people who already want tabular foundation models on modest censored cohorts and need a defensible label construction. It deserves a serious referee. I would engage with it and expect it to clear peer review with ordinary methods revisions.","headline":"Solid methods bridge: jackknife pseudo-RMST as a portable target interface for frozen tabular foundation models, with known-truth recovery and honest multi-backbone benchmarks.","tokens_in":23231,"tokens_out":542,"would_cite":true,"duration_ms":6336,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Jackknife pseudo-RMST targets convert right-censored survival outcomes into ordinary regression labels so tabular foundation models can predict restricted event-free time without survival-specific training.","keywords":["restricted mean survival time","pseudo-observations","tabular foundation models","right censoring","survival prediction","jackknife","RMST regression","censoring-aware targets"],"falsifier":"Build a controlled simulation with known conditional RMST but informative censoring (censoring depends on unobserved factors tied to the event). If pseudo-RMST targets then fail to cut oracle RMST error relative to naive min(Y,τ) labels and discrimination collapses, the claim that this target interface is a general portable bridge fails in that regime.","tokens_in":23147,"feed_emoji":"⏳","tokens_out":1104,"duration_ms":28076,"temperature":0.7,"pith_summary":"Right-censored follow-up blocks ordinary regression because observed time is not a complete event-time label. This paper claims that the survival-specific work can live in the target rather than the model: construct jackknife pseudo-observation labels for restricted mean survival time (RMST) at a chosen horizon, then hand those labels to ordinary tabular foundation-model backbones. In simulations with known conditional RMST, those pseudo-RMST targets recovered restricted event-free time more accurately than naive restricted observed-time or event-only labels, and multiple SurvFM backbones stayed competitive with classical survival and RMST-regression methods on an eligible static SurvSet subset. Predicted RMST also sorted held-out patients into groups with ordered observed event-free time and event enrichment. A sympathetic reader would care because many biomedical tables are modest, heterogeneous, and censored, and a portable target interface reuses reusable tabular predictors instead of rebuilding a survival-specific model for every cohort.","feed_headline":"Pseudo-RMST labels let tabular models handle censored survival","feed_subtitle":"Jackknife restricted-time targets turn right-censored outcomes into ordinary regression without survival-specific training.","key_machinery":"Jackknife pseudo-RMST targets: patient-level labels Zi(τ) = n θ̂τ(D) − (n−1) θ̂τ(D−i) formed from the full-sample and leave-one-out Kaplan–Meier RMST estimators. They convert a fold-level censored-data functional into regression responses whose conditional expectation approximates conditional RMST, so standard tabular regression becomes a meaningful horizon-specific survival task.","core_discovery":"SurvFM-RMST establishes that censoring can be handled through a target interface: jackknife pseudo-observations for RMST at a prespecified horizon turn right-censored follow-up into patient-level regression responses, so interchangeable tabular foundation-model backbones can perform horizon-specific restricted event-free-time prediction without survival-specific fine-tuning, risk-set losses, or new survival-output heads. Controlled simulations with known conditional RMST and a target-definition ablation support that the pseudo-RMST construction is substantive rather than cosmetic, and public-data benchmarks show competitive discrimination and RMST-scale error across several backbones, with r","pith_inferences":["The same target-side pattern could be applied to competing risks by building pseudo-observations for cause-specific restricted mean time or cumulative incidence, extending the interface without redesigning foundation backbones.","If the interface is truly portable, it should also drop into classical tabular regressors and AutoML stacks, not only foundation models—a low-cost check of how much of the gain is the target versus the backbone.","Refreshing pseudo-RMST at successive landmarks as new covariates arrive would test whether the static interface can become a lightweight route to time-updated prognosis.","A diagnostic that flags thin risk sets or near-informative censoring before target construction would be a practical guardrail for the load-bearing assumption."],"forward_implications":["Censored survival tables can be fed to existing tabular foundation-model regressors by changing only the outcome construction, not the backbone architecture.","Horizon-specific RMST predictions give restricted event-free-time outputs on the original time scale that can stratify held-out patients into ordered risk groups with event enrichment.","The survival-specific component becomes a reusable target recipe, so multiple tabular backbones can be swapped under the same censoring-aware interface.","Pseudo-RMST labels outperform naive restricted observed-time and event-only labels for recovering known conditional RMST.","Relative model ranking remains endpoint-, horizon-, and constraint-dependent, so validation inside the intended setting still decides backbone choice."],"fun_headline_variants":["Jackknife RMST targets turn censoring into plain regression for tabular FMs","Pseudo-RMST interface lets tabular backbones predict censored event-free time","Horizon-specific RMST pseudo-labels free tabular models from survival fine-tuning","Censoring-aware jackknife targets unlock RMST regression on tabular foundation models","SurvFM-RMST: portable pseudo-RMST labels for censored survival without new heads"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The method assumes that leave-one-out RMST pseudo-labels act as usable noisy patient-level targets for true restricted mean survival time when censoring is independent of the event process and finite-sample risk sets are not too thin.","fun_headline_variants_meta":{"raw":{"variants":["Jackknife RMST targets turn censoring into plain regression for tabular FMs","Pseudo-RMST interface lets tabular backbones predict censored event-free time","Horizon-specific RMST pseudo-labels free tabular models from survival fine-tuning","Censoring-aware jackknife targets unlock RMST regression on tabular foundation models","SurvFM-RMST: portable pseudo-RMST labels for censored survival without new heads"]},"model":"grok-4.5","effort":"low","cost_usd":0.00193,"raw_usage":{"total_tokens":935,"prompt_tokens":824,"num_sources_used":0,"completion_tokens":111,"cost_in_usd_ticks":19300000,"prompt_tokens_details":{"text_tokens":824,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":0,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":824,"tokens_out":111,"duration_ms":4588,"temperature":1.0,"reasoning_tokens":0,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-13T02:00:36.598802+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Build a controlled simulation with known conditional RMST but informative censoring (censoring depends on unobserved factors tied to the event). If pseudo-RMST targets then fail to cut oracle RMST error relative to naive min(Y,τ) labels and discrimination collapses, the claim that this target interface is a general portable bridge fails in that regime.","supporting_citations":[],"review_version":1}