{"id":"eed56f20-58f9-4f9c-9c7e-589c83087616","arxiv_id":"2505.02529","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":9,"one_line_summary":"RobSurv uses vector quantization to learn noise-resistant discrete features alongside continuous features, improving CT/PET survival prediction and noise robustness on three cancer datasets.","lead":"This paper introduces RobSurv, a deep learning model that uses vector quantization to combine CT and PET images for cancer survival prediction. The authors report higher accuracy and stronger resistance to synthetic image noise than existing methods on three public datasets.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The robustness claim lacks statistical support: the 3.8–4.5% degradation is computed from subsets with overlapping error bars and a noise level selected after seeing RobSurv's own degradation.","rationale":"The paper's novelty and clinical value rest on the noise-robustness claim, not on the modest clean-data C-index gains. That claim depends on a single table whose degradation percentages are internally inconsistent and whose uncertainty is large enough to erase the effect. The supplementary's explicit statement that the noise level was chosen after observing RobSurv's degradation makes the baseline comparison vulnerable to selection bias, and the absence of any baseline data across the noise grid prevents checking whether the 8–12% figure is representative. The reader identified the synthetic-noise/post-hoc-selection issue; I add that the reported numbers do not even establish the effect statistically within their own experimental protocol. This is addressable with a systematic grid evaluation and proper uncertainty quantification, so the appropriate verdict remains conditional.","tokens_in":14739,"tokens_out":7152,"duration_ms":88545,"concrete_test":"Re-run all methods on each dataset at every noise level in the supplementary grid (CT σ = 0.01/0.05/0.1, PET low/medium/high) using one fixed test set with 50% corrupted samples, and report per-fold paired clean-vs-noisy Ctd-index differences with bootstrap confidence intervals for RobSurv and each baseline. If the 3.8–4.5% vs 8–12% margin appears at every level with non-overlapping intervals and the H&N1 drop is ≤4.5%, the robustness claim is supported; otherwise the margin is not established.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim—that RobSurv degrades only 3.8–4.5% under severe noise versus 8–12% for baselines—is supported only by Table 1, which compares Ctd-index on 'clean samples' with Ctd-index on 'noisy samples' in a test set where 50% of samples are corrupted. That comparison is not statistically established: RobSurv's HECKTOR numbers are 0.771 ± 0.13 (clean) and 0.742 ± 0.14 (noisy); the 0.029 gap is far smaller than the reported fold-to-fold spread, and no paired test or confidence interval is reported for any degradation. The H&N1 degradation is 0.040/0.742 = 5.4%, not within the claimed 3.8–4.5% range, and the text's 4.3% does not match either absolute or relative arithmetic. Moreover, the supplementary ('Dataset for Robust Model') states the noise level was selected after observing RobSurv's degradation across a grid, and no baseline degradation curves across that grid are shown, so the 8–12% baseline figure is a single-point comparison at a level chosen after seeing the proposed model's behavior. Because Ctd-index is a rank statistic, scores computed on different patient subsets (clean vs noisy) are not directly comparable without controlling for the valid-pair distribution. Together these issues mean the headline robustness margin could be a subset-composition artifact or noise-level selection effect; the paper does not currently rule that out.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces RobSurv, a multi-modal survival prediction framework that processes CT and PET images through two parallel pathways: a continuous feature stream and a vector-quantized discrete-token stream, with a patch-wise cross-modal fusion module (DualPatchFuse) and a competing-risks survival head. The authors evaluate on HECKTOR, H&N1, and NSCLC Radiogenomics, reporting state-of-the-art concordance indices on clean data and claiming that the model degrades by only 3.8–4.5% under synthetic noise, versus 8–12% for baselines. The central claims are robustness to noise and architectural novelty in applying vector quantization to multi-modal survival prediction.","tokens_in":15110,"tokens_out":5376,"duration_ms":62847,"significance":"If the reported robustness margins were statistically supported, the paper would be a useful contribution to multi-modal medical image analysis: the dual continuous/discrete representation is a sensible architectural idea, and the evaluation spans three public datasets with competing-risks survival modeling. The paper also draws attention to an important practical problem, namely performance degradation of imaging-based survival models under realistic acquisition noise. However, as presented, the quantitative claims are not internally consistent, and the robustness comparison lacks statistical support. The architectural contribution is plausible and worth further development, but the current evidence is not strong enough to support the headline '3.8–4.5% versus 8–12%' claim.","major_comments":[{"comment":"The headline degradation range is internally inconsistent. From Table 1, H&N1 drops from 0.742 to 0.702, which is (0.742−0.702)/0.742 = 5.4%, not the reported 4.3%. The values for NSCLC (0.734→0.701, 4.5%) and HECKTOR (0.771→0.742, 3.8%) happen to fall inside the Abstract's stated 3.8–4.5% range, but H&N1 does not. The authors should recompute all degradations on a single stated basis (absolute or relative) and correct the abstract, text, and range accordingly.","section":"Abstract; Table 1"},{"comment":"The robustness comparison is not statistically established. For HECKTOR, the clean score is 0.771 ± 0.13 and the noisy score is 0.742 ± 0.14; the 0.029 difference is far smaller than the fold-to-fold standard deviation, and no paired test, bootstrap confidence interval, or within-fold comparison is reported for any degradation. Moreover, the Ctd-index is a rank-based statistic computed on different patient subsets (clean samples versus noisy samples), so the observed gaps may reflect differences in the valid-pair distribution rather than model robustness. At minimum, the authors should report paired per-fold differences with confidence intervals and verify that the result is not a subset-composition artifact.","section":"§5, Table 1"},{"comment":"The noise-level selection procedure raises a risk of selection bias. The supplementary states that the highest noise configuration (CT σ = 0.1, PET high Poisson noise) was selected after observing RobSurv's degradation across a grid of noise levels, and only RobSurv's degradation curve is shown; no baseline model is evaluated across the same grid. Consequently, the claimed 8–12% baseline degradation comes from a single noise level chosen post hoc relative to the proposed model's performance. The authors should either pre-specify the noise levels or report all competing methods over the full noise grid, including baseline degradation curves.","section":"Supplementary, 'Dataset for Robust Model'; §4.1"},{"comment":"The claimed margins over baselines are not reproduced by Table 1. The text reports clean-data margins over TMSS of 5.2%, 4.1%, and 3.9%, and the introduction claims a 4–8% improvement; from Table 1, the clean margins over the best baseline (TMSS) are 0.734−0.708 = 0.026 (3.7%) on NSCLC, 0.771−0.751 = 0.020 (2.7%) on HECKTOR, and 0.742−0.725 = 0.017 (2.3%) on H&N1. Under noise, the margins over TMSS are 0.701−0.661 = 0.040 (6.1%), 0.742−0.712 = 0.030 (4.2%), and 0.702−0.689 = 0.013 (1.9%). The authors should clarify whether margins are absolute or relative and make every reported value consistent with Table 1.","section":"§5, Table 1; Introduction"},{"comment":"It is unclear whether synthetic noise is added only to the test sets or also to the training data. The text says the dataset is 'augmented' with noise, while Table 1 evaluates 'clean samples' and 'noisy samples' within the same test split. If noisy samples were used in training, the comparison against baselines trained on clean data is not a clean test-time robustness comparison; if they were not, the protocol should be stated explicitly. This distinction affects the interpretation of the entire robustness claim.","section":"§4.1, Table 1"}],"minor_comments":[{"comment":"The fusion module is called 'DualPathFuse' in the introduction and 'DualPatchFuse' elsewhere; use a single name consistently throughout.","section":"§1 and §3.2"},{"comment":"Around Eq. (4), 'referes' is a typo; in Eq. (7), the quantity A is called 'attention weights' but is defined as a value-weighted sum, so the distinction between attention weights and attention output should be clarified.","section":"Eq. (4) and Eq. (7)"},{"comment":"The reference [Farooq et al., 2024] cites a placeholder arXiv identifier (arXiv:2401.00000), and [Ma et al., 2024] appears to duplicate [De Biase et al., 2024] with the same title; the authors should supply correct and distinct citations.","section":"References"},{"comment":"The main text says HECKTOR data were collected from seven centers, while the supplementary says six; this discrepancy should be reconciled.","section":"§4.1 and Supplementary, 'Dataset Specifications'"},{"comment":"The entry for Multimodal Dropout on HECKTOR clean samples reports a standard deviation of 0.66, which is implausibly large relative to all other entries; please check and correct this value.","section":"Table 1"},{"comment":"The phrase 'resampling CT and PET scand' contains a typo ('scand' should be 'scans').","section":"Supplementary, 'Dataset Specifications'"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses a worthwhile problem and the proposed architecture is reasonable, but the central quantitative claims currently fail internal consistency checks and lack statistical support. The noise-level selection and the clean/noisy subset comparison need to be reworked. There are also citation issues (placeholder arXiv ID and an apparent duplicate reference) that should be corrected. I do not see grounds for rejection on scope, but the authors must substantively revise the evidence before the robustness claim can be accepted."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nQuick take: this is a new application of vector quantization to CT/PET survival prediction, with a sensible dual discrete-continuous design and ablations suggesting the quantization path matters. Worth a serious look, but the central robustness claim—3.8–4.5% degradation vs 8–12% for baselines—is not supported by the numbers as reported.\n\nWhat's actually new: applying VQ-GAN style codebooks to survival prediction is new in this literature. The dual-path architecture (discrete tokens plus continuous features, patch-wise bidirectional cross-attention, CBAM-style continuous fusion) is a reasonable engineering contribution. Ablations show removing DualVQ drops HECKTOR Ctd-index by roughly 13% clean and 18.7% noisy; that's a meaningful signal. The paper evaluates on three public datasets and compares against a fair set of baselines.\n\nSoft spots. First, the arithmetic: H&N1 degradation is 0.040/0.742 = 5.4%, not 4.3% as stated, and outside the abstract's 3.8-4.5% range. The claimed 4-8% improvement margins don't match Table 1 (margins are about 2.3-6.0%). Second, no significance testing anywhere; error bars overlap substantially (e.g., HECKTOR clean 0.771±0.13 vs noisy 0.742±0.14). A paired test or confidence interval on degradation is needed before claiming robustness. Third, the noise level for the main experiments was selected after seeing RobSurv's degradation (supplementary), and baseline degradation curves across the noise grid aren't shown. That's a selection effect; the headline margin could be an artifact of where they drew the line. Fourth, no code released, which matters for a reproducibility-heavy claim. Fifth, the Ctd-index comparisons between clean and noisy subsets may not be apples-to-apples if valid-pair distributions differ; the paper doesn't address this.\n\nI want to be fair: none of these are lethal to the underlying idea. The architecture is plausible, and the ablation evidence is genuine. But the paper as written oversells the robustness advantage. The fix is straightforward: report paired statistics, show full noise-response curves for all methods, correct the arithmetic, and release code.\n\nVerdict: this deserves peer review, but with heavy revision expected. A serious referee could make this a solid paper.","headline":"RobSurv is a plausible new VQ-based multi-modal survival architecture whose headline robustness claim doesn't yet survive contact with its own table.","tokens_in":15642,"tokens_out":1767,"would_cite":false,"duration_ms":20452,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"RobSurv claims that quantizing CT-PET image features into discrete codebooks preserves survival-prediction accuracy under noise, losing only 3.8–4.5% compared with 8–12% for baselines.","keywords":["cancer survival prediction","vector quantization","multi-modal learning","CT-PET imaging","noise robustness","competing risks","transformer fusion","concordance index"],"falsifier":"A reader could take any of the three datasets, add Gaussian noise to CT (standard deviation 0.1 relative to the intensity range) and high Poisson noise to PET as described in the paper's supplementary, and check whether RobSurv's concordance index drops only 3.8–4.5%; a larger drop, or a drop comparable to baseline methods, would show the robustness claim does not reproduce. A stronger test would use real low-count PET reconstructions and multi-center CT scans instead of synthetic noise.","tokens_in":14531,"feed_emoji":"🩻","tokens_out":10674,"duration_ms":112939,"temperature":0.7,"pith_summary":"The paper sets out to show that cancer survival prediction from paired CT and PET scans can be made resilient to image noise by quantizing learned features into discrete codebooks while keeping a parallel continuous path for fine detail. If correct, survival models could work in real clinical settings where scan quality varies across centers, without collapsing when scans are noisy. On three datasets (HECKTOR, H&N1, and NSCLC Radiogenomics), the proposed RobSurv model reports concordance indices (a standard ranking-accuracy measure for survival models) of 0.771, 0.742, and 0.734 on clean data, and under the paper's synthetic noise protocol it degrades only 3.8–4.5%, compared with 8–12% degradation for baseline methods. The paper also claims the model keeps statistically significant separation between high- and low-risk patient groups under noise, which is what would matter for treatment planning.","feed_headline":"Discrete codebooks keep survival prediction accurate amid noisy scans","feed_subtitle":"RobSurv's quantized CT-PET features lose only 3.8–4.5 percent under noise; baselines lose 8–12 percent.","key_machinery":"The load-bearing object is DualVQ, a pair of modality-specific vector-quantization modules (one for CT, one for PET), each with a learned codebook of 1,024 vectors. Each module maps continuous latent features to their nearest codebook entry, turning an 8 by 8 by 8 latent volume into a set of discrete tokens; the quantization is trained with codebook, commitment, and reconstruction losses and outputs both discrete tokens and continuous features. A second module, DualPatchFuse, splits the discrete tokens into 2 by 2 by 2 patches and applies bidirectional CT-to-PET and PET-to-CT cross-attention, while the continuous path uses channel-spatial attention; alignment and preservation losses tie the two streams together. The fused representation feeds a DeepHit-style network that estimates hazards for multiple competing risks over time intervals. The mechanism's work is to make the learned representation insensitive to high-frequency noise through tokenization while keeping detail through the continuous path.","core_discovery":"The paper's central claim is that discretizing CT and PET image features through learned vector-quantization codebooks gives survival prediction a noise-resistant representation, and that fusing this discrete representation with a continuous one preserves the fine-grained intensity information needed for prognosis. The architecture, RobSurv, processes each modality through separate encoders and codebooks, producing discrete tokens that capture stable anatomical and metabolic patterns, while a continuous branch retains detail. A patch-wise bidirectional cross-attention mechanism fuses CT and PET discrete tokens, and channel-spatial attention fuses the continuous streams; the combined representation feeds a discrete-time competing-risk survival network. On clean data the model reports concordance indices of 0.734 (NSCLC), 0.771 (HECKTOR), and 0.742 (H&N1), and under 50% noisy samples it reports drops of 4.5%, 3.8%, and 4.3%, respectively, while baseline models drop more steeply, with some exceeding 10 percentage points. The ablation study attributes most of the robustness to the vector-quantization module: removing it drops clean performance by about 13% and noisy performance by about 18.7%.","pith_inferences":["Editorial inference: if the quantized codebook is the main source of noise resistance, similar discrete-bottleneck designs could stabilize other medical imaging prediction tasks (e.g., segmentation, staging, treatment-response) under acquisition noise, provided a continuous path remains for detail.","Editorial inference: the paper's synthetic noise protocol (zero-mean Gaussian on CT, Poisson on PET) is unlikely to cover structured artifacts such as motion, metal streaks, or reconstruction differences; testing on real multi-center noisy scans would reveal whether the 3.8–4.5% degradation bound holds outside that protocol.","Editorial inference: because the continuous branch contributes little on clean data but more under noise, a dynamic weighting that shifts toward discrete tokens as noise increases could reduce the computational overhead that the paper's limitations section acknowledges.","Editorial inference: per-modality codebooks may offer an interpretability surface: examining which codebook entries are activated for high- versus low-risk patients could suggest imaging biomarkers for future study."],"forward_implications":["If RobSurv's robustness holds, survival models can be deployed in settings where CT and PET quality varies across scanners and centers, with smaller performance penalties than current methods.","Risk stratification remains statistically significant under noise: the paper reports log-rank p <= 0.05 separating high- and low-risk groups even with 50% noisy samples (and, per the supplementary, up to 90% noise).","The discrete token path alone can preserve most prognostic signal: ablating the continuous branch costs only about 0.8% clean-data performance but about 2.7–2.8% under noise, showing the discrete path carries the core survival information.","Vector quantization is the main stabilizer: removing DualVQ degrades clean performance by about 13% and noisy performance by about 18.7%, which suggests quantization, not attention alone, supplies the noise resistance.","The method generalizes across three cancer types (head-and-neck and lung) and different imaging protocols, supporting its use across disease sites."],"supporting_citations":[{"why":"Provides the HECKTOR PET/CT dataset with recurrence-free survival outcomes used to evaluate clean and noisy performance.","marker":"[Oreiller et al., 2022]"},{"why":"Provides the H&N1 head-and-neck dataset used as a second evaluation cohort.","marker":"[Wee and Dekker, 2019]"},{"why":"Provides the NSCLC Radiogenomics dataset used as the third evaluation cohort.","marker":"[Bakr et al., 2017; Bakr et al., 2018]"},{"why":"Supplies the DeepHit discrete-time competing-risk survival network that RobSurv's prediction head extends with likelihood and ranking losses.","marker":"[Lee et al., 2018]"},{"why":"Supplies the 3D VQ-GAN architecture that DualVQ adapts into parallel modality-specific encoders, codebooks, and decoders.","marker":"[Zhou and Khalvati, 2024]"},{"why":"Supplies the channel-spatial attention (CBAM) used in the continuous fusion path.","marker":"[Woo et al., 2018]"},{"why":"Defines the time-dependent concordance index used to measure all reported clean and noisy performance.","marker":"[Antolini et al., 2005]"},{"why":"TMSS baseline, the closest comparator on clean data and a reference point for noise degradation.","marker":"[Saeed et al., 2022]"},{"why":"SurvRNC baseline whose risk-stratification power the paper shows is lost under noise.","marker":"[Saeed et al., 2024]"},{"why":"XSurv baseline, a transformer-based CT-PET survival model the paper reports degrades sharply under noise.","marker":"[Meng et al., 2023]"}],"fun_headline_variants":["Quantized CT-PET features cut noise loss to under 5%","Vector quantization steadies cancer survival prediction amid noise","RobSurv: discrete codes halve noise-induced prediction errors","Discrete imaging codes keep survival prognosis stable under scan noise","Noise-robust survival prediction via quantized multimodal features"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole robustness result depends on the assumption that the synthetic noise protocol defined in the supplementary material — zero-mean Gaussian noise for CT and Poisson noise for PET, with the highest tested levels used in the main experiments — faithfully represents the noise encountered in real clinical scans.","fun_headline_variants_meta":{"raw":{"variants":["Quantized CT-PET features cut noise loss to under 5%","Vector quantization steadies cancer survival prediction amid noise","RobSurv: discrete codes halve noise-induced prediction errors","Discrete imaging codes keep survival prognosis stable under scan noise","Noise-robust survival prediction via quantized multimodal features"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000266,"raw_usage":{"total_tokens":1664,"prompt_tokens":1051,"completion_tokens":613,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":667,"completion_tokens_details":{"reasoning_tokens":529}},"tokens_in":667,"tokens_out":613,"duration_ms":7441,"temperature":1.0,"reasoning_tokens":529,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T00:48:45.605872+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A reader could take any of the three datasets, add Gaussian noise to CT (standard deviation 0.1 relative to the intensity range) and high Poisson noise to PET as described in the paper's supplementary, and check whether RobSurv's concordance index drops only 3.8–4.5%; a larger drop, or a drop comparable to baseline methods, would show the robustness claim does not reproduce. A stronger test would use real low-count PET reconstructions and multi-center CT scans instead of synthetic noise.","supporting_citations":[{"cited_title":"Merging-diverging hybrid transformer networks for survival prediction in head and neck cancer","cited_arxiv_id":null,"evidence_quote":"XSurv baseline, a transformer-based CT-PET survival model the paper reports degrades sharply under noise."},{"cited_title":"Head and neck tumor segmenta- tion in pet/ct: the hecktor challenge","cited_arxiv_id":null,"evidence_quote":"Provides the HECKTOR PET/CT dataset with recurrence-free survival outcomes used to evaluate clean and noisy performance."},{"cited_title":"Wee and A","cited_arxiv_id":null,"evidence_quote":"Provides the H&N1 head-and-neck dataset used as a second evaluation cohort."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the NSCLC Radiogenomics dataset used as the third evaluation cohort."},{"cited_title":"Deephit: A deep learning approach to survival analysis with competing risks","cited_arxiv_id":null,"evidence_quote":"Supplies the DeepHit discrete-time competing-risk survival network that RobSurv's prediction head extends with likelihood and ranking losses."},{"cited_title":"Conditional generation of 3d brain tumor regions via vq- gan and temporal-agnostic masked transformer","cited_arxiv_id":null,"evidence_quote":"Supplies the 3D VQ-GAN architecture that DualVQ adapts into parallel modality-specific encoders, codebooks, and decoders."},{"cited_title":"A time-dependent discrimination index for survival data","cited_arxiv_id":null,"evidence_quote":"Defines the time-dependent concordance index used to measure all reported clean and noisy performance."},{"cited_title":"Tmss: an end-to- end transformer-based multimodal network for segmenta- tion and survival prediction","cited_arxiv_id":null,"evidence_quote":"TMSS baseline, the closest comparator on clean data and a reference point for noise degradation."},{"cited_title":"Survrnc: Learn- ing ordered representations for survival prediction using rank-n-contrast","cited_arxiv_id":null,"evidence_quote":"SurvRNC baseline whose risk-stratification power the paper shows is lost under noise."}],"review_version":1}