{"id":"68cddd7e-8544-4b7f-b899-3607d43c79ce","arxiv_id":"2607.08503","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Frozen CT-CLIP representations plus a lightweight DeepSurv head outperform CoxPH and match or beat other multimodal baselines for lung-cancer survival on a 242-patient real-world cohort.","lead":"A frozen CT-CLIP foundation model plus a small survival head predicts lung-cancer death risk from CT scans and clinical notes better than a standard clinical model on 242 real patients. This shows pretrained medical image-text features can help when labeled outcome data are scarce.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.5","headline":"Small held-out test set (N=50) and single-center private data make the reported superiority of frozen CT-CLIP over CoxPH and multimodal baselines statistically fragile.","rationale":"The reader’s weakest-assumption (templated clinical notes vs real radiology reports) is real and correctly identified in §2.3, yet it is not the most load-bearing threat to the strongest claim. The claim is an empirical performance statement on N=50; that statement is under-powered and single-center. The frozen-encoder result is interesting and the ablations (Table 2) are clean, but without external validation or a larger test set the numbers cannot be taken as reliable evidence that “frozen CT-CLIP … outperforms the clinical baseline.” Keeping the CONDITIONAL verdict is therefore appropriate; the condition should be strengthened to require either multi-center confirmation or a properly powered resampling analysis of the discrimination gap. I partially agree with the reader because the note-distribution issue remains a secondary concern that should still be checked, but the primary soft spot is statistical fragility of the headline metric.","tokens_in":8577,"tokens_out":556,"duration_ms":7002,"concrete_test":"Bootstrap the held-out test set 1 000 times (or perform 5\times repeated stratified 80/20 splits of the full 242-patient cohort) and recompute the paired difference in Harrell’s C between frozen CT-CLIP and CoxPH; if the 95 % CI of the difference includes zero or the mean difference falls below 0.02, the superiority claim is not robust.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim rests on Harrell’s C = 0.755 ± 0.006 (frozen CT-CLIP) vs 0.721 ± 0.007 (CoxPH) and vs other multimodal models on a held-out test set of only 50 patients (Table 3, §3.3). With 189 events overall and ~39 events expected in the test split, the effective sample for ranking is tiny; the reported standard deviations come from three CV folds of the training set rather than from bootstrap or repeated test-set resampling. Consequently the Bonferroni-corrected t-test significance markers (*) and the claim of “comparable or improved performance” are highly sensitive to the particular 50-patient draw. The reader correctly flags the templated-note distribution shift, but that is secondary: even if the notes were perfect, the discrimination numbers themselves lack the precision needed to support a general claim about frozen foundation-model representations in data-scarce survival settings.","agreement_with_reader":"partial"},"referee_report":{"model":"grok-4.5","summary":"The paper evaluates CT-CLIP (a vision-language foundation model pretrained on chest CT volumes and radiology reports) as a feature extractor for multimodal overall-survival prediction on a private single-center cohort of 242 lung-cancer patients (pretreatment CT + clinical variables). Clinical variables are converted into template-generated notes so that the frozen CXR-BERT text encoder can be used; image and text embeddings are concatenated and fed to a lightweight DeepSurv head. Adaptation strategies (frozen, full fine-tuning, LoRA) and modality ablations are compared; the best configuration is a fully frozen CT-CLIP backbone plus trainable survival head. On a held-out test set of 50 patients this model reports Harrell’s C = 0.755 ± 0.006, Uno’s C = 0.744 ± 0.006 and 5-year AUC = 0.853 ± 0.016, outperforming CoxPH (0.721 ± 0.007) and matching or exceeding several published multimodal baselines, while also producing statistically significant high/low-risk Kaplan–Meier separation.","tokens_in":8851,"tokens_out":1060,"duration_ms":9786,"significance":"If the result holds under external validation, the work supplies a practical, low-parameter recipe for leveraging domain-specific foundation-model representations in data-scarce clinical survival settings: freeze the pretrained encoders, train only a small Cox head. This is directly relevant to real-world lung-cancer cohorts that rarely exceed a few hundred cases with reliable outcome labels. Strengths include a clear experimental design (adaptation ablations, modality ablations, multiple published multimodal baselines, Bonferroni-corrected tests) and an explicit demonstration that freezing outperforms full fine-tuning on this small cohort. The contribution is therefore of genuine applied interest even if the absolute performance numbers remain provisional.","major_comments":[{"comment":"Table 3 and §3.3: the central claim that frozen CT-CLIP “outperforms the clinical baseline and achieves comparable or improved performance relative to other multimodal approaches” rests on a single held-out test set of N = 50 (≈39 events). The reported standard deviations are obtained from three training folds rather than from bootstrap or repeated test-set resampling; consequently the Bonferroni-corrected significance markers and the ranking versus ResNet+Tabular (0.757) and Interactive-Model (0.741) are highly sensitive to the particular 50-patient draw. At minimum the authors should supply bootstrap confidence intervals on the test set and, ideally, an external multi-center cohort before the superiority claim can be regarded as robust.","section":null},{"comment":"§2.3 and the modality-ablation results (Table 2): the text branch relies on the untested assumption that template-generated clinical notes (“Male lung cancer patient. The patient is diagnosed with adenocarcinoma, overall stage IIA …”) lie sufficiently close in distribution to the real radiology reports used to pre-train CXR-BERT. No quantitative check of embedding-space shift or ablation that replaces the templated notes with a conventional tabular encoder is provided. Because the multimodal gain over the image-only branch is modest (0.755 vs 0.741), this distributional assumption is load-bearing for the claim that the frozen text encoder contributes survival-informative features.","section":null}],"minor_comments":[{"comment":"Figure 1 caption and §2.3: the exact template used to generate clinical notes is only partially illustrated; a complete example (including handling of missing variables) should be supplied in the appendix for reproducibility.","section":null},{"comment":"Table 1: the validation C-indices for frozen versus LoRA configurations are numerically close and overlapping; a short discussion of why full fine-tuning systematically under-performs would strengthen the interpretation of the frozen-backbone result.","section":null},{"comment":"§3.1: the precise definition of overall survival (time origin, censoring rules) and the distribution of follow-up times should be stated more explicitly, as they affect interpretation of the 5-year AUC.","section":null},{"comment":"References: the preprint arXiv:2403.17834 (CT-CLIP) and the PEFT library citation are appropriate, but a brief note on the exact checkpoint (CT-CLIP_v2.pt) and any preprocessing differences from the original CT-RATE pipeline would aid replication.","section":null}],"recommendation":"major_revision","confidential_remarks":"The single-center private cohort and N=50 test set are the main reasons I cannot recommend acceptance in the present form; the experimental design itself is careful and the frozen-backbone finding is interesting. If the authors can add bootstrap CIs and at least one external validation set (or a clear multi-center plan), the paper would become a solid contribution for a methods-oriented clinical-AI venue. Novelty relative to other frozen-encoder survival papers is incremental but the concrete CT-CLIP adaptation is useful."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The one thing worth knowing is that a frozen CT-CLIP (image + templated clinical notes) plus a tiny DeepSurv head beats CoxPH and matches or edges the usual multimodal CNN baselines on a real 242-patient single-center cohort. That is new empirical evidence, not a new method.\n\nWhat they do well: they actually run the adaptation ablations (frozen / full FT / LoRA), modality ablations, and several published multimodal baselines with the same protocol. The frozen configuration wins cleanly on their internal numbers (Harrell C 0.755 vs 0.721 for CoxPH), the Kaplan–Meier separation is clean, and they are honest that full fine-tuning overfits. The math is standard Cox partial likelihood; no circularity. Citations are appropriate and they correctly note that Xing et al. cannot be re-run.\n\nSoft spots, in proportion. The held-out test set is only 50 patients (~39 events). The reported stds come from the three training folds, not from resampling the test set, so the Bonferroni stars and the “outperforms / comparable” language are sensitive to that particular split. That is a real limitation for any general claim about foundation-model representations, but it is the same limitation every scarce-data clinical paper faces; they do not hide it. The template-note assumption (that CXR-BERT still works on “Male lung cancer patient, stage IIA…”) is untested and secondary. No code, no external validation.\n\nThis is useful for groups sitting on 100–300-patient CT + tabular cohorts who want a drop-in feature extractor rather than training a 3-D ViT from scratch. It is not a methods advance and will not change how anyone designs foundation models. I would still send it to peer review: the experiment is clean enough and the practical question is real enough that referees should see it. Minor revision for clearer uncertainty quantification and a stronger caveat on the test-set size would be enough. Worth a look if you work in this niche; otherwise skim the tables and move on.","headline":"Solid, careful application paper showing frozen CT-CLIP embeddings work for scarce-data lung-cancer survival; the N=50 test set makes the superiority claim fragile but does not erase the practical value.","tokens_in":9480,"tokens_out":519,"would_cite":false,"duration_ms":6572,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"A frozen CT-CLIP backbone plus a small survival head beats clinical baselines for lung-cancer prognosis on a 242-patient cohort.","keywords":["multimodal survival analysis","lung cancer","CT-CLIP","foundation models","frozen encoders","DeepSurv","prognosis prediction"],"falsifier":"On a larger independent lung-cancer cohort, replace the template notes with actual free-text clinical reports (or ablate the text branch entirely) and check whether the frozen multimodal C-index still significantly exceeds the pure clinical baseline.","tokens_in":9477,"feed_emoji":"🫁","tokens_out":900,"duration_ms":8719,"temperature":0.7,"pith_summary":"Lung-cancer survival models are hard to train because large, well-curated imaging cohorts with reliable outcome labels are scarce. This paper shows that a domain-specific vision-language foundation model, CT-CLIP, already encodes useful prognostic information in its frozen image and text representations. By turning clinical variables into short template notes and feeding both the pretreatment CT volume and the note through the frozen CT-CLIP encoders, then training only a lightweight DeepSurv head on the concatenated embeddings, the authors obtain better discrimination than a pure clinical Cox model and match or exceed several task-specific multimodal baselines. Freezing the backbone also prevents the overfitting that appears when the same model is fully fine-tuned or adapted with LoRA on only 192 training cases. The resulting risk scores cleanly separate patients into high- and low-risk groups on Kaplan–Meier curves. The practical message is that pretrained medical foundation-model features can be used off-the-shelf for survival modelling when labelled data are limited.","feed_headline":"Frozen CT-CLIP beats clinical baselines for lung-cancer survival","feed_subtitle":"On 242 patients, a tiny trainable head on frozen medical embeddings separates high- and low-risk groups","key_machinery":"Frozen CT-CLIP dual encoders (CT-ViT image branch + CXR-BERT text branch) that map a preprocessed CT volume and a template-generated clinical note into a shared 1024-dimensional space; only the subsequent lightweight DeepSurv head is trained with the Cox partial log-likelihood.","core_discovery":"On a real-world cohort of 242 lung-cancer patients, a completely frozen CT-CLIP model whose 512-dimensional image and text embeddings are simply concatenated and passed through a two-layer DeepSurv head yields a held-out Harrell’s C-index of 0.755, outperforming the clinical CoxPH baseline (0.721) and matching or exceeding several trained multimodal CNN and fusion models, while also producing statistically significant high- versus low-risk separation.","pith_inferences":["The same frozen-feature recipe may transfer to other scarce-label oncology tasks (response prediction, recurrence) that already possess CT-CLIP-compatible inputs.","If template notes already work, modest domain-adaptive pre-training of the text encoder on clinical notes rather than pure radiology reports could further close the remaining gap to larger multimodal models.","The observed superiority of frozen over fine-tuned configurations suggests that many medical foundation models are currently under-utilised as pure feature extractors in low-resource settings."],"forward_implications":["Hospitals with only a few hundred labelled CT cases can obtain competitive multimodal survival models by freezing a medical foundation model and training a small head.","Full fine-tuning or LoRA of large CT encoders is often unnecessary and can degrade performance under severe data scarcity.","Template-based conversion of tabular variables into natural-language notes is a viable way to reuse radiology-pretrained text encoders for survival tasks.","Risk scores from the frozen model can be used directly for clinically meaningful high/low-risk stratification without additional calibration steps."],"fun_headline_variants":["Frozen CT-CLIP tops clinical baseline for lung-cancer survival","Frozen CT-CLIP with light head beats CoxPH on 242 patients","CT-CLIP frozen embeddings lift C-index past clinical Cox model","Tiny head on frozen CT-CLIP separates lung-cancer risk groups","Frozen CT-CLIP multimodal features outperform clinical survival baseline"],"cache_read_input_tokens":128,"weakest_assumption_plain":"The short template sentences built from tabular clinical variables are close enough in style and content to the real radiology reports used to pre-train the text encoder that the frozen embeddings remain informative for survival.","fun_headline_variants_meta":{"raw":{"variants":["Frozen CT-CLIP tops clinical baseline for lung-cancer survival","Frozen CT-CLIP with light head beats CoxPH on 242 patients","CT-CLIP frozen embeddings lift C-index past clinical Cox model","Tiny head on frozen CT-CLIP separates lung-cancer risk groups","Frozen CT-CLIP multimodal features outperform clinical survival baseline"]},"model":"grok-4.5","effort":"low","cost_usd":0.003902,"raw_usage":{"total_tokens":1191,"prompt_tokens":715,"num_sources_used":0,"completion_tokens":71,"cost_in_usd_ticks":39020000,"prompt_tokens_details":{"text_tokens":715,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":405,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":715,"tokens_out":71,"duration_ms":54155,"temperature":1.0,"reasoning_tokens":405,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-10T06:22:17.105570+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"On a larger independent lung-cancer cohort, replace the template notes with actual free-text clinical reports (or ablate the text branch entirely) and check whether the frozen multimodal C-index still significantly exceeds the pure clinical baseline.","supporting_citations":[],"review_version":1}