{"id":"7c4516c3-e73e-44cd-bae5-eb4ec0802b51","arxiv_id":"2508.21581","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Multimodal fusion of CT and pathology images improves recurrence risk prediction in kidney cancer, with the best model approaching the clinical Leibovich score.","lead":"This paper combines CT scans and digitized tumor tissue slides to predict whether kidney cancer will return after surgery. It finds that tissue-based models beat CT alone, and merging both improves predictions, though a clinical score still remains competitive.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No significance testing means the central fusion-advantage claim is not established; a permutation test is needed.","rationale":"The reader identified the filtered-cohort representativeness as the weakest assumption. That is a legitimate external-validity concern, but the more load-bearing issue is internal: the primary quantitative claim—that intermediate fusion improves over unimodal WSI—is supported only by point estimates without any significance test. With 40 events and fold-level standard deviations of ~0.04–0.05, a 0.03 difference in C-index is well within noise. The paper acknowledges variance in Section 4 ('these gains should be interpreted with caution due to their reported standard deviations') but the abstract and conclusion nonetheless assert the improvement. The concern is not that the authors are dishonest; they are careful and hedge in places. The issue is that the inferential step from point estimates to a claim is not secured. However, the paper is explicitly a feasibility study with modest claims; the conditional verdict already reflects the need for external validation and significance testing. Adding a specific significance test would strengthen the paper without necessarily overturning its conclusions. I agree partially with the reader: the cohort-filtering concern is real but secondary; the main load-bearing weakness is the lack of statistical testing on the primary comparison. My recommended verdict remains CONDITIONAL (UNCHANGED relative to the reader), because the reader already conditioned acceptance on addressing these limitations, and the missing significance test falls squarely under that condition.","tokens_in":8467,"tokens_out":1529,"duration_ms":15227,"concrete_test":"Run a paired permutation test: for the best intermediate fusion model (TITAN-CONCH + ResNet-18) vs unimodal TITAN-CONCH, permute the model-assignment labels within each outer test fold (or use bootstrap resampling of patients) and recompute the C-index difference under the null; report the two-sided p-value. If p > 0.05 (likely given n=156, 40 events, and reported fold-wise stds), the paper should downgrade the fusion-advantage claim. As a secondary check, re-evaluate the Leibovich comparison using the same tie-handling for all models.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that intermediate fusion improves over unimodal WSI models, with the best intermediate model (TITAN-CONCH + ResNet-18) at C-index 0.775 vs TITAN-CONCH unimodal at 0.745 (Table 1). The paper reports mean C-index over five outer folds with standard deviations, but no statistical test comparing models is provided. Differences of ~0.03–0.07 in C-index are well within plausible cross-validation noise for n=156 with 40 events, especially given the reported fold-standard deviations (e.g., ±0.044 for the best model, ±0.046 for unimodal TITAN-CONCH). Without paired significance testing (e.g., permutation test across folds or bootstrap over patients), the claim that fusion adds value is not established. The authors' own caution about standard deviations (Section 4) is acknowledged, but the abstract/conclusion still assert that intermediate fusion improved performance. This is a correctness-risk issue: the conclusion rests on a difference that may be sampling noise. Additionally, the random tie-breaking analysis (Leibovich RT C-index 0.749) is used to claim the Leibovich score's discretization overstates performance, but the same tie-breaking should have been applied to the learned models' continuous predictions if the claim is about fair comparison.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a modular deep learning framework that integrates preoperative CT and postoperative whole-slide images (WSIs) for recurrence risk prediction in clear cell renal cell carcinoma (ccRCC). Using a TCGA-KIRC subset of 156 patients, it compares unimodal WSI and CT models, late fusion, and intermediate fusion, all based on frozen foundation-model encoders plus a Cox-based MLP prediction head. The adjusted Leibovich score is used as a clinical baseline. The paper reports that WSI-based models outperform CT-only models, that intermediate fusion gives the best learned model (TITAN-CONCH + ResNet-18, C-index 0.775 ± 0.044), approaching the adjusted Leibovich score (0.805 ± 0.035), and that random tie-breaking lowers the baseline to 0.749 ± 0.043.","tokens_in":8775,"tokens_out":4909,"duration_ms":58439,"significance":"If the claims hold, the study is a useful feasibility demonstration of multimodal imaging integration for a clinical task where structured scores remain the standard. Strengths include the use of a public dataset, nested cross-validation with fixed outer folds, and honest reporting of fold-level standard deviations. The main limitations are the small cohort (156 patients, 40 events), manual cohort filtering, and the absence of statistical testing for model comparisons, which makes the central fusion-advantage claim suggestive rather than established. The paper's methodological framework is sound in structure but needs additional statistical rigor before the conclusions can be accepted at face value.","major_comments":[{"comment":"The central claim that intermediate fusion improves over unimodal WSI models is not statistically supported. The best intermediate model (TITAN-CONCH + ResNet-18) achieves C-index 0.775 ± 0.044 versus 0.745 ± 0.046 for TITAN-CONCH unimodal. With n=156 and 40 events, a 0.030 difference is well within the reported fold-level standard deviations. No paired significance test (e.g., permutation over patients or outer folds, bootstrap confidence intervals) is provided. Please add such tests for all pairwise comparisons of interest, report p-values/confidence intervals, and adjust the abstract/conclusion wording accordingly if the difference is not significant.","section":"Section 4, Table 1"},{"comment":"The 'best model' is identified from 18 learned configurations (2 WSI encoders × 3 CT encoders × 3 fusion strategies) using the same outer folds. No multiple-comparison correction or independent model-selection procedure is described. The reported best result may partly reflect selection bias. Please report results for all configurations (already in Table 1) and additionally provide a model-selection rule, e.g., choosing the fusion strategy on inner-fold performance only, or correct for the number of comparisons when making inferential claims.","section":"Section 3, Experimental Protocol"},{"comment":"The manual exclusion of 31 out of 187 patients based on CT quality, contrast phase, and kidney visibility is a potential source of selection bias. No comparison of excluded versus included patients' clinical characteristics or outcomes is provided. If the exclusions correlate with outcome, the relative model performance may not generalize. Please report baseline demographics, stage, grade, event rates, and follow-up for the excluded patients, and discuss or quantitatively assess the impact of this filtering.","section":"Section 3, Datasets"},{"comment":"There is an internal inconsistency regarding necrosis. Section 3 states that necrosis was omitted from the adjusted Leibovich score 'due to its absence in the dataset,' but Section 4 Case A states that 'the pathology report noted necrosis and high Fuhrman grade.' Clarify whether structured necrosis data were unavailable despite pathology reports containing the information, and discuss how this affects the validity of the adjusted baseline. Additionally, random tie-breaking is applied only to the Leibovich baseline; if the claim is that discretization overstates baseline performance, provide the same analysis for any tied risk scores produced by the learned models (or justify their absence).","section":"Sections 3 and 4 (Adjusted Leibovich score; Case A)"}],"minor_comments":[{"comment":"The text states that 'all WSI-CT combinations matching or exceeding their WSI-only baselines' in intermediate fusion, but Table 1 shows TITAN-CONCH + ResNet-10 intermediate fusion at 0.742 versus 0.745 unimodal TITAN-CONCH. Please correct this overstatement.","section":"Section 4, 'Performance Analysis of Unimodal and Multimodal Strategies'"},{"comment":"The naming of CT encoders is unclear: Table 1 lists 'ResNet-10' and 'ResNet-18,' but Section 2 mentions MedicalNet and SwinUNETR. Please clarify which architecture corresponds to each named model and whether ResNet-10/18 are MedicalNet variants.","section":"Section 2/Table 1"},{"comment":"The 'random tie-breaking' result is reported as a single mean C-index. Please specify the number of random repetitions and the seed or variance across repetitions to make the result reproducible.","section":"Section 4, Leibovich (RT)"},{"comment":"The paper does not mention code or feature-extraction pipeline availability. Given the public dataset, providing code would strengthen reproducibility, even if only for preprocessing and evaluation.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is honest and well-structured, but the abstract's 'intermediate fusion further improved performance' is not yet supported by statistical testing. I would recommend requiring a permutation/bootstrap analysis and a discussion of model-selection multiplicity before acceptance. The paper is within scope and the empirical framework is usable, but the central claim needs either stronger evidence or more cautious wording."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague, quick take on arXiv:2508.21581. The genuinely new thing here is the systematic comparison: first to combine CT and WSIs for ccRCC recurrence prediction in TCGA-KIRC, testing unimodal, late, and intermediate fusion with foundation-model encoders. The design is honest—nested CV, fixed folds, hyperparameter tuning in the inner loop, and the authors report standard deviations and explicitly caution about them. The WSI-over-CT result is robust in direction, and the finding that weak CT encoders can still contribute through fusion is a useful empirical data point. The adjusted Leibovich baseline is computed independently from clinical variables, so there is no circularity.\n\nThe soft spots are real but proportionate. The cohort is small (n=156, 40 events) after manual CT quality filtering, and no significance testing is done. The best intermediate fusion (TITAN–CONCH + ResNet-18, C-index 0.775±0.044) vs unimodal TITAN–CONCH (0.745±0.046) is well within fold noise; you cannot conclude fusion improves over WSI alone from these numbers. The authors are more careful in Section 4 than in the abstract/conclusion, which overstate the fusion benefit. This is load-bearing for the paper's main claim. Also, the adjusted Leibovich without necrosis is a weakened baseline, acknowledged but still a limitation. No code or processed data are released, which hurts reproducibility despite using a public dataset.\n\nOne point in the stress-test I would push back on: the tie-breaking critique. Learned models produce continuous risk scores, so there are no ties to break; random tie-breaking for the discrete Leibovich score is a sensible sensitivity analysis, not an unfair comparison. The absence of significance testing is the real issue.\n\nWho is this for? Researchers working on multimodal survival prediction with foundation models; it is a solid feasibility reference. It deserves a serious referee—conditional accept with requested significance testing or external validation, and ideally code release.","headline":"Careful feasibility benchmark for CT+WSI fusion in ccRCC recurrence, but the fusion-advantage claim rests on differences within one standard deviation without significance testing.","tokens_in":9253,"tokens_out":1878,"would_cite":true,"duration_ms":20061,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper shows that combining a patient's tumor pathology slide with their preoperative CT scan predicts kidney cancer recurrence better than either image type alone, and nearly as well as the standard clinical risk score.","keywords":["clear cell renal cell carcinoma","recurrence risk prediction","multimodal imaging fusion","whole-slide images","CT imaging","Cox survival model","foundation models","Leibovich score"],"falsifier":"Take the same intermediate-fusion pipeline to an external ccRCC cohort with more recurrence events and compare, under identical nested cross-validation, fused WSI+CT against WSI-only: if the fused model does not exceed the WSI-only C-index, the paper's central claim that CT adds complementary value through fusion is refuted. Separately, if the adjusted Leibovich score keeps a large lead over learned models after randomized tie-breaking in a larger sample, the claim that discretization overstates its individualized performance would be weakened.","tokens_in":8418,"feed_emoji":"🔬","tokens_out":7093,"duration_ms":72314,"temperature":0.7,"pith_summary":"The paper tries to establish that recurrence risk in clear cell renal cell carcinoma can be predicted from routine imaging alone, and that combining the two image types—the postoperative tumor slide and the preoperative CT—gives more accurate risk rankings than either modality on its own. If true, this matters because current practice relies on a coarse clinical score that assigns many patients the same risk and ignores imaging altogether. The authors report that pathology-based models consistently beat CT-only models, that joining the two feature vectors before prediction (intermediate fusion) beat both late fusion and unimodal models, and that the best fused model's concordance index (0.775) approaches the adjusted Leibovich score (0.805). They also find that the clinical score's apparent advantage shrinks when its tied scores are broken randomly, suggesting discretization inflates its apparent individual-level accuracy.","feed_headline":"Fused pathology and CT scan nears clinical kidney-cancer benchmark","feed_subtitle":"Combining whole-slide tumor images with CT beats either alone and rivals the clinical risk score.","key_machinery":"The pipeline's engine is intermediate fusion by embedding concatenation inside a Cox-based survival model. Whole-slide images are compressed by frozen pathology foundation models (TITAN–CONCH or CHIEF–CTransPath) into patient-level vectors; CT volumes are cropped to the kidneys, downsampled, and encoded by fine-tuned 3D encoders (MedicalNet ResNet variants or SwinUNETR) into vectors of the same dimension. Concatenating the two vectors and feeding them through a small multilayer perceptron trained with the Cox partial-likelihood loss lets the model learn cross-modality feature interactions before risk scoring. The same setup is also run unimodally and with late fusion (a weighted average of s","core_discovery":"The central claim is that whole-slide histopathology and preoperative CT carry complementary prognostic information for clear cell renal cell carcinoma recurrence, and that a simple fusion of their learned representations extracts it. In the authors' experiments, frozen pathology foundation-model embeddings consistently produced higher C-index values than fine-tuned CT encoders when used alone, so pathology dominates the signal; but every WSI–CT combination that concatenated the two feature vectors before the survival head matched or exceeded its WSI-only baseline, with the best result from TITAN–CONCH plus ResNet-18 (C-index 0.775 ± 0.044). The adjusted Leibovich score remained the highest","pith_inferences":["If the fusion gain replicates in a larger cohort, it would imply that current CT encoders, pretrained mainly for segmentation, are leaving prognostic information on the table; a CT foundation model trained for survival tasks might close the gap with pathology.","The paper's case analysis hints that slide sampling can miss aggressive regions; a natural extension is to test multiple slides per tumor or attention-based slide selection, which the authors did not do.","A practical translation would be a continuous imaging-based risk score that could be thresholded flexibly for surveillance intensity, rather than locked into three clinical risk groups; this follows from the paper's tie-breaking argument but is not tested here."],"forward_implications":["If the reported ordering holds, future recurrence-risk tools could be built from routine imaging alone, with no extra tests beyond slides and CT scans already acquired in standard care.","CT's prognostic value in this setting is conditional on fusion: the same CT encoders that scored near chance alone contributed to the best fused model, so radiology should be evaluated in combination, not in isolation.","The adjusted Leibovich score's lead narrows substantially under random tie-breaking, implying that discrete clinical scores may overstate their ability to rank individual patients; continuous learned risk scores may be fairer comparators.","Because simple concatenation already improved every WSI–CT combination, more expressive fusion (cross-attention or co-learning) is a plausible next step the authors explicitly leave open."],"supporting_citations":[{"why":"Supplies the public cohort with CT scans, whole-slide images, and recurrence outcomes used in all experiments.","marker":"[21, 2]"},{"why":"Defines the Leibovich score, the clinical baseline against which all learned models are benchmarked.","marker":"[17, 18]"},{"why":"Prior multimodal recurrence scoring work showing WSI-derived features contribute strongly, motivating the multimodal approach.","marker":"[10]"},{"why":"Provides the Cox proportional hazards model that underlies the survival loss used by all prediction heads.","marker":"[7]"},{"why":"DeepSurv, the deep-learning Cox partial-likelihood framework the prediction architecture adapts.","marker":"[16]"},{"why":"CONCH pathology foundation model, the patch encoder in the best-performing WSI pipeline.","marker":"[19]"},{"why":"TITAN slide-level encoder that aggregates CONCH patch embeddings into patient-level vectors in the best model.","marker":"[9]"},{"why":"MedicalNet 3D ResNet pretrained on CT/MRI, the CT encoder family (ResNet-10 and ResNet-18) used in unimodal and fusion models.","marker":"[6]"}],"fun_headline_variants":["Fusing CT and tumor slides improves kidney cancer risk prediction","Pathology trumps CT, but fusion boosts kidney cancer prediction","Simple fusion of CT and slides approaches clinical kidney risk score","Multimodal model nearly matches clinical kidney cancer recurrence score"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The 156-patient cohort used for all comparisons is the subset of patients whose CT scans passed a manual quality review; if the excluded 31 patients differ systematically in recurrence risk, the relative performance of the models may not generalize.","fun_headline_variants_meta":{"raw":{"variants":["Fusing CT and tumor slides improves kidney cancer risk prediction","Pathology trumps CT, but fusion boosts kidney cancer prediction","Simple fusion of CT and slides approaches clinical kidney risk score","Multimodal model nearly matches clinical kidney cancer recurrence score"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000765,"raw_usage":{"total_tokens":3225,"prompt_tokens":737,"completion_tokens":2488,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":481,"completion_tokens_details":{"reasoning_tokens":2420}},"tokens_in":481,"tokens_out":2488,"duration_ms":18060,"temperature":1.0,"reasoning_tokens":2420,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T14:08:53.318143+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the same intermediate-fusion pipeline to an external ccRCC cohort with more recurrence events and compare, under identical nested cross-validation, fused WSI+CT against WSI-only: if the fused model does not exceed the WSI-only C-index, the paper's central claim that CT adds complementary value through fusion is refuted. Separately, if the adjusted Leibovich score keeps a large lead over learned models after randomized tie-breaking in a larger sample, the claim that discretization overstates its individualized performance would be weakened.","supporting_citations":[],"review_version":1}