{"id":"5cbc5ff7-097e-4fe3-809b-5b1dbe0155ec","arxiv_id":"2411.18253","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A transformer-based temporal attention network that integrates longitudinal blood, imaging, and medication data predicts mortality in immunotherapy patients with AUCs around 0.81 to 0.84, slightly beating blood-only models.","lead":"This paper combines blood tests, CT-measured organ volumes, and medication records over time to predict survival in 694 immunotherapy patients. The best model reaches AUCs of 0.81 to 0.84 for 3, 6, 9, and 12 month mortality, with modest gains over using blood data alone.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Best-variant selection on the test set may inflate the reported MMTSimTA improvement over MMSimTA.","rationale":"The reader's weakest assumption (exclusion of 280 patients, selection bias) is acknowledged in the paper and primarily affects generalizability; it does not directly threaten the internal validity of the model-to-model comparison within the selected cohort. The most load-bearing concern for the central claim is the unacknowledged selection of the best-performing variant on the same test data used for significance testing. This is a concrete, correctable methodological issue: the DeLong tests and AUCs in Table 3 are reported for the best variant, but the choice of 'best' was made from the test-set results, producing optimistic estimates. It is also distinct from the reader's stated weakest assumption, so the agreement is partial. A re-analysis with pre-specified variants or nested selection would settle whether the improvement is real or an artifact. The verdict remains CONDITIONAL, matching the reader, because external validation and pre-specification are needed before the claim can be accepted as robust.","tokens_in":14828,"tokens_out":11959,"duration_ms":107132,"concrete_test":"Pre-specify MMTSimTA-concat+SA as the experimental model and MMSimTA-concat+SA as the baseline, and within each cross-validation fold select the variant (Concat vs Concat+SA) using only training-fold validation AUCs before computing held-out test AUCs and DeLong P-values. If the pooled test AUC difference shrinks or loses significance at 6/9/12 months, the reported improvement is partly a test-set selection artifact.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central comparison in Table 3 and Figure 2 labels MMTSimTA-concat+SA as the best variant of the extended model and MMSimTA-concat+SA as the best baseline variant, but the variant labels were evidently chosen after inspecting test-fold AUCs. No pre-registration, nested variant selection, or multiple-comparison correction is described. Because the same test sets are reused to pick the best of Concat, Concat+SA, and late-fusion variants, the reported AUCs and DeLong P-values are optimistically biased. The headline differences are small: at 3 months 0.84 vs 0.83 (P>.05), at 6 months 0.83 vs 0.80 (P=.01), at 9 months 0.82 vs 0.78 (P=.001), and at 12 months 0.81 vs 0.78 (P=.01). A selection effect of only 0.01-0.02 in the chosen variant could change which endpoints remain significant. The claim that the extended architecture improves multimodal prognosis therefore rests on an inference procedure that does not account for the model-selection step.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript proposes MMTSimTA, a multimodal architecture that embeds SimTA temporal-attention blocks inside transformer encoder blocks, and evaluates it on a pan-cancer cohort of 694 immunotherapy-treated patients using longitudinal blood tests, CT-derived organ volumes, and medication records. The models predict 3-, 6-, 9-, and 12-month mortality as a multitask binary classification problem. Using 3-fold cross-validation, the authors compare unimodal SimTA/TSimTA, multimodal MMSimTA/MMTSimTA under intermediate fusion (Concat and Concat+SA variants), and a model-agnostic late-fusion strategy (MA-MMSimTA and MA-MMTSimTA). The headline result is that the best MMTSimTA variant (Concat+SA) attains mean AUCs of 0.84, 0.83, 0.82, and 0.81, with reported DeLong improvements over the best MMSimTA variant at 6-, 9-, and 12-month endpoints, and over late-fusion MMTSimTA at early endpoints. The authors acknowledge the exclusion of 280 patients and the lack of external validation as limitations.","tokens_in":14968,"tokens_out":8820,"duration_ms":79039,"significance":"If the reported improvements survive a properly controlled model-selection protocol, the paper would be a useful contribution to multimodal longitudinal prognosis: it extends SimTA with a transformer-style block, integrates three noninvasive modalities in a large real-world cohort, and compares intermediate versus late fusion in a systematic way. Strong points include the cohort size and modality coverage (11,249 blood tests, 14,849 medication records, 1,337 CT scans), the honest discussion of limitations, and the stated intention to release code. The significance is currently limited by evaluation design: the best variant appears to be selected on the same test folds used to report AUCs and p-values, and at least one table entry appears to contradict the abstract's 'strongest prognostic performance' claim. These issues need to be resolved before the conclusions can be taken at face value.","major_comments":[{"comment":"The abstract's claim that 'the strongest prognostic performance was demonstrated using a variant of the MMTSimTA model' is not supported by the results as reported. Table 3 lists MA-MMSimTA (late-fusion baseline) with mean AUCs of 0.85, 0.81, 0.81, and 0.81, which is higher than the MMTSimTA-Concat+SA values of 0.84, 0.83, 0.82, and 0.81 at the 3-month endpoint and equal at 12 months. The results text compares MMTSimTA with MMSimTA and with MA-MMTSimTA but does not report a pairwise test of MMTSimTA-Concat+SA versus MA-MMSimTA at 3 months; if that test is reported in Table S8, it should be moved to the main text, and the abstract should be qualified accordingly.","section":"Abstract / Results, Table 3"},{"comment":"The selection of the 'best' variant appears to be made from test-fold AUCs. The same three test folds are used to choose among Concat, Concat+SA, and late-fusion configurations and then to compute the reported AUCs and DeLong p-values. No nested cross-validation, separate validation set, or multiple-comparison correction is described. Because the improvement of MMTSimTA-Concat+SA over MMSimTA-Concat+SA is small (0.01 at 3 months, 0.03 at 6 and 9 months, 0.03 at 12 months), part or all of the reported advantage could be due to selection. Please provide a nested variant-selection procedure or corrected p-values covering all variants and endpoints, and report the results for all variants rather than only the best.","section":"Model training and validation strategies / Results, Table 3"},{"comment":"The exclusion of 280 patients with insufficient follow-up is acknowledged as a possible source of selection bias, and Table S2 shows significant differences between included and excluded patients on several variables. Because the task requires a 12-month follow-up window, the excluded patients are likely to have shorter survival and may be non-ignorably missing. The manuscript should include a sensitivity analysis (for example, inverse-probability weighting or a survival-analysis formulation that can censor the excluded patients) or, at minimum, a quantitative bound on how much the model ranking could shift under plausible outcome assumptions for the excluded group.","section":"Study cohort and ethical approval / Results, Patient characteristics"},{"comment":"The statistical analysis section pools three per-fold DeLong p-values using Fisher's method but does not report effect sizes or confidence intervals for the AUC differences, and the 3-fold design has limited power. At 3 months the headline difference is not significant (P>.05), while at 6, 9, and 12 months the p-values (0.01, 0.001, 0.01) are obtained after the same test data have been used for variant selection. Reporting per-fold AUCs with confidence intervals and adjusting for the four endpoints (and the number of variants) would make the strength of the evidence easier to assess.","section":"Statistical analysis"}],"minor_comments":[{"comment":"The phrase 'with area under the curves (AUCs)' should be 'with areas under the receiver operating characteristic curves'; the plural form should be used consistently.","section":"Abstract"},{"comment":"The relationship between the Mann-Whitney U test and the DeLong test is unclear; please state explicitly which test is used for which comparison and why both are needed.","section":"Statistical analysis"},{"comment":"The 3-month row is labeled 'No death' and reports only three AUC values; the absence of a three-month value and the meaning of 'No death' should be explained in the table or text.","section":"Table 4"},{"comment":"The abbreviations 'T' and 'N' in the figure are not defined in the figure legend; the legend should define these symbols.","section":"Figure 1"},{"comment":"The manuscript states the code 'will be publicly accessible,' but the provided GitHub link is not resolved in the current version; please ensure the availability statement is accurate at the time of publication.","section":"Data availability"},{"comment":"The notation 'Concat+SA' and 'Concat þ SA' is used interchangeably; please standardize the notation across the text, tables, and figures.","section":"Tables and text"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is honest about its limitations and the code-sharing intent is good, but the headline claim needs to be reconciled with Table 3 and the variant-selection issue addressed before I can support acceptance. The self-citations are appropriate as direct methodological predecessors, but the phrase 'novel architecture' should be calibrated because the main contribution is an incremental extension of SimTA."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe paper is a reasonable, incremental application of an existing attention module (SimTA) inside a transformer encoder, applied to a pan-cancer immunotherapy cohort with longitudinal blood, organ-volume, and medication data. That combination is new, and the clinical question is worth asking. The authors also run a sensible set of baselines, including a model-agnostic late fusion, and they are unusually honest in the Discussion about the selection bias from excluding 280 patients and the lack of external validation. Credit where due: the data curation is substantial (694 patients, 11,000+ blood tests, 1,300+ CTs) and the paper clearly states its limitations.\n\nThe soft spots are real, and the stress-test note lands. The headline claim—MMTSimTA beats MMSimTA—is based on picking the best variant (Concat+SA) after looking at the test-set AUCs, with no nested selection or multiple-comparison correction. The differences are small: at 3 months 0.84 vs 0.83 (not significant), and the model-agnostic late fusion baseline actually gets 0.85 at 3 months. Only the 6- and 12-month comparisons survive with P<.05. A selection effect of 0.01–0.02 could easily change which endpoints stay significant. On top of that, there is no external validation, only 3 folds of cross-validation, and the code isn't released yet (despite the data availability statement saying it will be). These are not fatal—the paper is honest about most of them—but they mean the central claim should be read as 'promising' rather than established.\n\nThe paper is not a paradigm shift. For someone working on multimodal longitudinal prognosis in oncology, it's a useful data point and a decent benchmark reference. For the general reader, the incremental gain over the blood-only model is modest, so I wouldn't push it. It deserves a serious referee because the question is important and the methods are competently executed; I'd recommend engaging with it critically, mainly to push for external validation and a more disciplined model-selection procedure.","headline":"Solid incremental application paper whose headline improvement rests on test-set variant selection and a single-center cohort; worth a careful read, not a practice-changer.","tokens_in":15564,"tokens_out":1983,"would_cite":false,"duration_ms":22848,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A multimodal transformer that fuses longitudinal blood tests, medication records, and CT-derived organ volumes predicts 3-, 6-, 9-, and 12-month mortality in immunotherapy patients, with the best variant reaching an AUC of 0.84 at 3 months.","keywords":["artificial intelligence","deep learning","immunotherapy","longitudinal study","multimodal data integration","survival prediction","temporal attention","transformer"],"falsifier":"Run the same MMTSimTA concat+SA and blood-only TSimTA models on an external cohort that includes patients with shorter follow-up, and compare their 3-month mortality AUCs; if the multimodal model does not beat the blood-only model, or its AUC drops substantially below 0.84, the claimed advantage of multimodal longitudinal integration fails.","tokens_in":14607,"feed_emoji":"🩺","tokens_out":8293,"duration_ms":68172,"temperature":0.7,"pith_summary":"The paper sets out to show that a deep learning model trained on routinely collected, noninvasive longitudinal data—blood tests, prescribed medications, and CT-derived organ volumes—can predict short-term mortality in cancer patients treated with immunotherapy. The proposed architecture, called MMTSimTA, embeds a simple temporal attention module inside a transformer encoder so that each modality is encoded separately and then fused at the feature level. In a pan-cancer cohort of 694 patients, the best variant reaches area under the receiver operating characteristic curve (AUC) values of 0.84, 0.83, 0.82, and 0.81 for 3-, 6-, 9-, and 12-month survival prediction, outperforming the baseline multimodal model at most endpoints and the best unimodal blood-based model at the earlier endpoints. If this holds, it would mean that data already collected in routine care could support personalized prognosis for immunotherapy patients without additional invasive procedures.","feed_headline":"Fusing blood, scans, and drug data predicts immunotherapy survival","feed_subtitle":"A temporal-attention transformer beats single-modality blood models for short-term mortality in 694 patients.","key_machinery":"The load-bearing object is the MMTSimTA network, which replaces the self-attention layer of a transformer encoder with the simple temporal attention (SimTA) module—an attention mechanism that linearly encodes the time intervals between examinations and assumes the most recent examination is the most informative. Each modality is processed by its own TSimTA block, the resulting representations are concatenated, optionally passed through a multi-head self-attention block, and then fed to a multilayer perceptron that jointly predicts all four survival endpoints; multimodal dropout lets the network train on patients with missing modalities.","core_discovery":"On the paper's own terms, the central discovery is that fusing longitudinal noninvasive modalities with a transformer-based temporal attention network yields better short-term mortality prediction than the baseline multimodal architecture or the best single modality. Concretely, the best variant—MMTSimTA with concatenated modality representations followed by a self-attention block—reaches AUCs of $0.84 \\pm 0.04$, $0.83 \\pm 0.02$, $0.82 \\pm 0.02$, and $0.81 \\pm 0.03$ for 3-, 6-, 9-, and 12-month survival prediction in a cohort of 694 immunotherapy-treated cancer patients. The paper interprets this as evidence that the added nonlinearity and skip connections of the transformer extension let the model exploit feature-level interactions between blood, medication, and imaging data, particularly for near-term endpoints.","pith_inferences":["If these results replicate in an external cohort, the architecture could support early mortality-risk stratification in routine oncology practice using data already collected for every patient, with no added invasive tests.","Because the SimTA module is hard-wired to weight the most recent examinations most heavily, a learnable temporal-attention variant might find informative patterns in earlier time points; testing that variant on the same cohort would clarify how much the recency prior matters.","The large gap between blood markers and organ-volume features suggests that replacing crude volume summaries with learned imaging features could change the size of the multimodal gain, a question the paper leaves open.","A time-to-event formulation that uses all available follow-up, rather than excluding patients with short follow-up, could directly test whether the acknowledged selection bias affects the reported model ranking."],"forward_implications":["The best MMTSimTA variant significantly outperforms the baseline MMSimTA at the 6-, 9-, and 12-month endpoints, so the transformer extension itself adds prognostic value beyond the original temporal attention module.","The multimodal model beats the best unimodal blood model at the 3- and 6-month endpoints, indicating that the added imaging and medication modalities contribute early prognostic signal rather than only noise.","Intermediate fusion outperforms late fusion for the extended transformer models, while the simpler baseline models do better with late fusion, suggesting that the right fusion strategy depends on the capacity of the unimodal encoders.","Feeding the model 6 months of on-treatment data instead of 3 months improves AUCs on the same patients, so longer observation windows strengthen mortality prediction."],"supporting_citations":[{"why":"Introduces the simple temporal attention module that the paper replaces self-attention with.","marker":"[31]"},{"why":"Supplies the transformer encoder architecture being extended.","marker":"[43]"},{"why":"Provides the automatic segmentation method used to derive organ volumes from CT scans.","marker":"[47]"},{"why":"Demonstrates an earlier multimodal longitudinal application of SimTA for immunotherapy response prediction.","marker":"[33]"},{"why":"Supplies the model-agnostic late fusion strategy and the earlier finding that blood markers dominate.","marker":"[40]"},{"why":"Motivates the intermediate fusion approach used in the multimodal network.","marker":"[50]"},{"why":"Introduces multimodal dropout, which lets the network train with incomplete modalities.","marker":"[51]"},{"why":"Provides the embedding layers used to encode medication codes.","marker":"[49]"}],"fun_headline_variants":["AI fuses routine data to forecast immunotherapy survival","Multimodal AI predicts survival from blood, scans, meds","Transformer improves short-term survival prediction in immunotherapy","Temporal-attention network beats blood-only models for survival","Deep learning fuses blood, CT, meds for immunotherapy survival"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The model comparison assumes that the 280 patients excluded for having less than three months of follow-up—who differed significantly from the included 694 on several variables—would not overturn the observed ranking of models if they were kept in the study.","fun_headline_variants_meta":{"raw":{"variants":["AI fuses routine data to forecast immunotherapy survival","Multimodal AI predicts survival from blood, scans, meds","Transformer improves short-term survival prediction in immunotherapy","Temporal-attention network beats blood-only models for survival","Deep learning fuses blood, CT, meds for immunotherapy survival"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.002259,"raw_usage":{"total_tokens":8766,"prompt_tokens":1022,"completion_tokens":7744,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":638,"completion_tokens_details":{"reasoning_tokens":7664}},"tokens_in":638,"tokens_out":7744,"duration_ms":45220,"temperature":1.0,"reasoning_tokens":7664,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T11:22:02.367099+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same MMTSimTA concat+SA and blood-only TSimTA models on an external cohort that includes patients with shorter follow-up, and compare their 3-month mortality AUCs; if the multimodal model does not beat the blood-only model, or its AUC drops substantially below 0.84, the claimed advantage of multimodal longitudinal integration fails.","supporting_citations":[{"cited_title":"MIA-prognosis: a deep learning framework to predict therapy response","cited_arxiv_id":null,"evidence_quote":"Introduces the simple temporal attention module that the paper replaces self-attention with."},{"cited_title":"Attention is all you need","cited_arxiv_id":null,"evidence_quote":"Supplies the transformer encoder architecture being extended."},{"cited_title":"TotalSegmentator: robust segmentation of 104 anatomic structures in CT images","cited_arxiv_id":null,"evidence_quote":"Provides the automatic segmentation method used to derive organ volumes from CT scans."},{"cited_title":"A multi-omics-based serial deep learning approach to predict clinical outcomes of single-agent anti- PD-1/PD-L1 immunotherapy in advanced stage non-small-cell lung cancer","cited_arxiv_id":null,"evidence_quote":"Demonstrates an earlier multimodal longitudinal application of SimTA for immunotherapy response prediction."},{"cited_title":"Integrated non - invasive diagnostics for prediction of survival in immunotherapy","cited_arxiv_id":null,"evidence_quote":"Supplies the model-agnostic late fusion strategy and the earlier finding that blood markers dominate."},{"cited_title":"Multimodal machine learn - ing: a survey and taxonomy","cited_arxiv_id":null,"evidence_quote":"Motivates the intermediate fusion approach used in the multimodal network."},{"cited_title":"ModDrop: adaptive multi-modal gesture recognition","cited_arxiv_id":null,"evidence_quote":"Introduces multimodal dropout, which lets the network train with incomplete modalities."}],"review_version":1}