{"id":"3295a972-04e8-4e7e-80b4-12a2b05de019","arxiv_id":"2509.01161","paper_version":1,"verdict":"REJECT","confidence":"HIGH","novelty_score":3.0,"correctness_risk":"high","formal_verification":"none","parameter_count":4,"one_line_summary":"XGBoost combining MRI radiomics and clinical biomarkers reportedly reaches C-index 0.782 for early brain tumor recurrence, but the paper's methods describe a liver-cancer cohort and no evaluation of its claimed temporal module appears.","lead":"The paper claims a multi-modal machine learning framework that combines MRI radiomics with clinical biomarkers to predict early brain tumor recurrence, reporting an XGBoost C-index of 0.782. A critical reader should know that the Methods section describes a liver cancer cohort while the Results report brain tumor patients, so the study as written does not support its headline claim.","discovery_kind":"unclear","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Methods describe an HCC cohort (Section 3.1) but Results report brain tumor patients (Section 5.1); the reported metrics are unsupported by the described data.","rationale":"The reader's strongest claim—that the framework predicts early recurrence in brain tumor patients—depends on the analyzed cohort actually being the brain tumor cohort described in Section 5.1. The Methods, however, describe a completely different population: Section 3.1 states inclusion criteria for patients undergoing curative-intent hepatic resection for suspected HCC; Section 3.2 specifies liver MRI protocols; Section 3.3 uses alpha-fetoprotein for surveillance and cites HCC references [104,106,107]; and the collected variables (Edmondson grade, resection type, AFP-based surveillance) are HCC-specific. The Results then report glioblastoma and anaplastic astrocytoma patients with MGMT methylation, IDH mutation, and Ki-67—markers that are not part of routine HCC workup. No passage reconciles this discrepancy. As written, the described methods cannot generate the reported results. This is an internal inconsistency, not a disagreement with external consensus. The reader's identification of this as the weakest assumption is accurate. Other issues (temporal encoder defined but never evaluated; duplicated Results sections with inconsistent model counts; Section 7 clustering on 'simulated' immune scores) are secondary to this cohort mismatch. The paper's candid limitations in Section 8 (retrospective, single-center, no external validation) are worth noting but do not address the mismatch. Therefore the central claim remains unsupported as written, and the reader's REJECT verdict stands.","tokens_in":14866,"tokens_out":3280,"duration_ms":35222,"concrete_test":"Request the de-identified dataset or the IRB-approved protocol and verify the cohort: list the 186 patients' primary tumor diagnoses, imaging acquisition protocols (brain vs. liver), and the exact variables collected. Specifically, confirm whether MGMT methylation, IDH mutation, and Ki-67 were measured in these patients and whether the MRI sequences are T1-Gd/ADC brain protocols or liver sequences. If the dataset contains HCC patients, the reported brain-tumor metrics cannot be reproduced. Alternatively, attempt to reproduce Table 2 by re-deriving univariate Cox hazards from the described HCC variables (tumor size, Edmondson grade, resection type, AFP); the variables MGMT methylation and GLCM entropy would not exist.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3.1 defines the study population as patients who underwent curative-intent hepatic resection for suspected HCC, with inclusion criteria for HCC and exclusion criteria about extrahepatic metastasis; Section 3.2 describes liver MRI protocols; Section 3.3 uses alpha-fetoprotein surveillance and references [104,106,107] on HCC. Yet Section 5.1 reports 186 patients with glioblastoma (65.6%) and anaplastic astrocytoma (34.4%), and the molecular biomarkers (MGMT, IDH1/2, Ki-67) are brain-tumor markers not part of HCC standard care. No passage reconciles this discrepancy. If the Methods do not describe the cohort actually analyzed, then the reported C-index, AUC, calibration, and KM split in Tables 1-5 and Figures 2-3 are not results for the described study, and the central claim is unsupported. This is an internal inconsistency, not a disagreement with external consensus.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a multi-modal machine learning framework that integrates MRI radiomic features with clinical and molecular biomarkers to predict early recurrence in high-grade brain tumors. It reports XGBoost as the best model with C-index 0.782, 1-year AUC 0.804, and 2-year AUC 0.767 (Table 1), plus Kaplan-Meier risk stratification with median RFS 9.6 vs. 21.2 months (log-rank p < 0.001), SHAP feature importance, calibration, and decision curve analysis. The central claim is that this framework is a usable risk-stratification tool. However, the Methods describe a hepatocellular carcinoma (HCC) cohort with liver MRI protocols and alpha-fetoprotein surveillance, while the Results report glioblastoma and anaplastic astrocytoma patients with MGMT/IDH/Ki-67 markers. Sections 5 and 6 duplicate results and disagree on the number of evaluated models. The risk-stratification analysis uses the median in-sample predicted score from the same cohort used for feature selection and training, making the survival separation largely a restatement of model fit. These issues leave the central claim unsupported.","tokens_in":15102,"tokens_out":2641,"duration_ms":32543,"significance":"If the reported results were valid and properly evaluated, the framework could be a practically useful tool for postoperative brain tumor risk stratification, because it combines easily available MRI and clinical markers, uses standard survival metrics, and provides interpretability via SHAP. The paper also includes algorithmic pseudocode and a clear experimental setup in principle. However, the significance cannot be assessed from the manuscript as written: the cohort mismatch and the duplication/inconsistency between the two Results sections mean that the reported performance numbers cannot be attributed to a well-defined study population or evaluation protocol. The in-sample survival stratification further undermines the predictive claim. The work therefore does not currently make a sound contribution to the literature.","major_comments":[{"comment":"The study population is described as patients who underwent curative-intent hepatic resection for suspected hepatocellular carcinoma, with liver MRI protocols in Section 3.2 and alpha-fetoprotein surveillance in Section 3.3. Yet Section 5.1 reports 186 patients with glioblastoma (65.6%) and anaplastic astrocytoma (34.4%), with molecular biomarkers MGMT, IDH1/2, and Ki-67 that are not part of HCC standard care. No passage reconciles this discrepancy. If the Methods do not describe the cohort actually analyzed, then all reported metrics in Tables 1-5 and Figures 2-3 are not attributable to the described study, and the central claim is unsupported. This is an internal inconsistency, not a minor presentation issue.","section":"Section 3.1 vs. Section 5.1"},{"comment":"The two Results sections duplicate each other but do not agree on the evaluated models. Section 5.3, Table 1 reports four models (XGBoost, CoxBoost, RSF, GBM), while Section 6.1 states that six models were compared, including CoxPH and CNN-based unimodal baselines; Table 3 also includes CoxPH. The text in Section 5.4 also refers to calibration for 'XGBoost and RSF' without mentioning the other models. No details are given for the CNN baseline, the training/validation splits, or the internal cross-validation procedure mentioned in Algorithm 1. This inconsistency makes it impossible to know which model set generated the reported numbers and prevents any reproducibility assessment.","section":"Sections 5 and 6"},{"comment":"Patients are stratified into high- and low-risk groups based on the median XGBoost predicted recurrence score, and Kaplan-Meier analysis is then performed on the same cohort that was used for univariate Cox feature selection (step 2 of Algorithm 1) and model training. The reported log-rank p < 0.001 and median RFS difference of 9.6 vs. 21.2 months therefore reflect in-sample discrimination rather than an out-of-sample validation of the risk-stratification tool. A proper evaluation would require a held-out test set or nested cross-validation, and ideally an independent cohort, before such survival separation can be claimed as predictive evidence.","section":"Sections 5.5 and 6.3, Algorithm 1 steps 16-17"},{"comment":"The temporal self-attention framework described in Section 3.4 (z(t) = SelfAttn(x(t) + PE(t))) is not used in Algorithm 1, in the model training description, or in any reported result. The Discussion in Section 8 even admits the 'time-series representation was relatively shallow,' contradicting the claimed temporal modeling component. Similarly, Section 7 introduces six immunological clusters based on 'simulated immune cell enrichment scores' and radiomic intensity distributions, but these analyses are not connected to the recurrence prediction cohort, are not mentioned in the Abstract or Introduction, and appear to be exploratory simulations rather than results from the study data. These disconnected components should either be integrated or removed.","section":"Section 3.4, Section 7"}],"minor_comments":[{"comment":"The manuscript contains placeholders such as '[Institution Name]' in Section 3.1 and '[software name]' in Section 3.2. These must be filled before any submission.","section":"Throughout"},{"comment":"The caption reads 'Kapian-Meier' instead of 'Kaplan-Meier'; please correct the typo.","section":"Figure 3 caption"},{"comment":"References [1]-[68] are almost entirely unrelated to brain tumors, HCC, or imaging; they appear to be a large block of self-citations or topic-diverse citations. This is inappropriate and should be replaced with relevant literature.","section":"References"},{"comment":"Having two 'Results' sections with overlapping content is confusing. They should be merged into one coherent Results section, with a single set of tables and figures.","section":"Section 5 vs. Section 6"},{"comment":"The feature selection result mentions 'GLSZM zone variance' but the methods in Section 3.2 only list GLCM and GLRLM texture features; please clarify whether GLSZM was included in the 107 features.","section":"Section 5.2"},{"comment":"The text refers to 'three identified immunological clusters' while Section 7 describes six clusters. This inconsistency needs to be resolved.","section":"Section 7.1"}],"recommendation":"reject","confidential_remarks":"The manuscript has a fundamental cohort mismatch: the Methods describe an HCC cohort while the Results report brain tumor patients. This is not a presentational issue; it invalidates the attribution of every reported metric. The duplicated Results sections with different model counts further indicate that the paper is not internally consistent. The risk-stratification claim is circular because it uses the training cohort's median score to produce the survival split. The reference list contains a large block of unrelated citations, and the temporal and immunological components appear disconnected from the main analysis. I do not see a feasible revision within the scope of this manuscript that would resolve these issues; a rewrite with a correctly described cohort, a single consistent evaluation, and proper out-of-sample validation would be needed. I therefore recommend rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The one thing you need to know: the Methods describe a hepatocellular carcinoma cohort and the Results report glioblastoma/anaplastic astrocytoma patients. Section 3.1 explicitly says patients underwent curative-intent hepatic resection for suspected HCC, Section 3.2 describes liver MRI protocols, Section 3.3 uses alpha-fetoprotein surveillance. Section 5.1 then reports 186 patients with glioblastoma and anaplastic astrocytoma, with MGMT/IDH/Ki-67. No passage reconciles this. Every reported C-index and AUC is uninterpretable because we don't know what data was actually analyzed.\n\nWhat does the paper do well? The Discussion lists honest limitations (retrospective, single-center, moderate sample size, no external validation). The general pipeline of univariate Cox selection plus standard survival learners is clearly described. But nothing here is new: the cited literature (Lao, Ren, Gomaa) already covers radiomics plus clinical features for glioma survival. The only apparently new component, the temporal self-attention encoder in Section 3.4, never appears in the evaluated models. Section 7 clusters patients on \"simulated immune cell enrichment scores\" and then draws radiomic associations with those clusters; that's circular and not a real analysis.\n\nThe other soft spots: there are two Results sections (5 and 6) that duplicate each other and disagree on model counts (four vs six). The KM stratification uses the median XGBoost risk score from the same cohort used for feature selection and training, so the log-rank p<0.001 largely restates the in-sample fit. No code, no data, no external validation.\n\nThis does not hold up as a scientific claim. A serious editor should desk reject it, not because the topic is uninteresting but because the manuscript's internal inconsistency makes it unfixable without redoing the study and rewriting the Methods.\n\nFor you, it's not worth reading group time. Might be worth showing to students as an example of why you check that the Methods and Results describe the same population.\n\nRecommendation: desk reject.","headline":"The Methods describe an HCC cohort while the Results report brain tumors; the internal inconsistency makes every performance claim uninterpretable.","tokens_in":15676,"tokens_out":1937,"would_cite":false,"duration_ms":20113,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims a multimodal MRI-plus-clinical model can predict early brain-tumor recurrence after resection, with XGBoost reaching a C-index of 0.782.","keywords":["brain tumor recurrence","radiomics","MRI","XGBoost","survival analysis","clinical biomarkers","risk stratification","multimodal machine learning"],"falsifier":"Pull the institutional cohort list behind Section 3.1: if the 186 patients underwent hepatic resection with liver MRI and AFP surveillance, the reported brain-tumor recurrence times and C-index cannot be produced from them. Short of that, a reader can test the out-of-sample claim by checking whether any patient was held out before model selection; the Methods only mention internal cross-validation.","tokens_in":14618,"feed_emoji":"🧠","tokens_out":4999,"duration_ms":56633,"temperature":0.7,"pith_summary":"The paper sets out to show that combining structural MRI radiomic features with clinical and molecular biomarkers improves early recurrence prediction for high-grade brain tumors after surgery. It trains four survival models on a multi-modal feature set and reports that XGBoost performs best, with a concordance index of 0.782 and 1- and 2-year AUCs of 0.804 and 0.767. It further reports that the model separates patients into high- and low-risk groups with median recurrence-free survival of 9.6 versus 21.2 months (log-rank p < 0.001). If these results hold, the model would be a directly usable risk-stratification aid for follow-up planning. The reader should note that the Methods and Results sections describe different tumor populations, an issue flagged in the inferences below.","feed_headline":"XGBoost hits 0.782 C-index in brain-tumor recurrence prediction","feed_subtitle":"Paper: fusing MRI radiomics with clinical biomarkers separates high- and low-risk groups by more than a year in median RFS.","key_machinery":"The load-bearing machinery is the multi-modal feature vector combined with survival-loss training: Cox partial likelihood for XGBoost and CoxBoost, log-rank splitting and cumulative-hazard averaging for RSF, and boosting for GBM. The paper also describes a temporal encoding module that applies positional encoding and self-attention to follow-up snapshots, intended to replace the static risk score with a dynamically learned one. Evaluation uses C-index, time-dependent AUC, calibration curves, Brier scores, and decision-curve net benefit.","core_discovery":"On its own terms, the paper's central claim is that a multi-modal feature vector—107 IBSI-compliant radiomic features extracted from preoperative structural MRI plus clinical and molecular variables such as MGMT methylation, IDH1/2 status, Ki-67, tumor size and resection type—carries enough signal to rank patients by recurrence risk. XGBoost trained with the Cox partial likelihood achieves the best discrimination; calibration curves and decision-curve analysis are claimed to favor it over RSF, CoxBoost and GBM. SHAP analysis names MGMT methylation, GLCM entropy and Ki-67 as the top contributors, and median-score splitting yields a statistically significant survival separation.","pith_inferences":["The strongest check is cohort identity: Section 3.1 describes patients who underwent hepatic resection for hepatocellular carcinoma, with liver MRI and alpha-fetoprotein surveillance, while Section 5.1 reports glioblastoma and anaplastic astrocytoma outcomes. If the Methods text describes the actual cohort, the brain-tumor results cannot be reproduced from it; if it is stale template text, the rep","The evaluation may be in-sample: the model section mentions optimizing hyperparameters by internal cross-validation but does not state a held-out test set; metrics computed on training data would overstate discrimination.","Section 7's 'immunological clustering' uses simulated immune enrichment scores, so the radiomic-intensity associations with immune clusters are illustrative rather than evidence-based.","The paper itself lists retrospective single-center design, moderate sample size, and lack of external validation as limitations; these would likely compress the reported C-index in a genuinely unseen cohort, making a multi-institutional test the natural next step."],"forward_implications":["If the XGBoost result generalizes, clinicians could use the risk score to schedule more intensive surveillance for high-risk patients (median RFS 9.6 months) and less frequent follow-up for low-risk patients.","The model targets the two-year window after surgery, which is the clinically urgent period for early recurrence, rather than only overall survival.","The reported feature rankings give a short list of routinely collected variables—MGMT methylation, IDH1 status, Ki-67, GLCM entropy—that could guide future data collection and model-building.","If confirmed, the performance comparison would position XGBoost as the default estimator among the four tested algorithms for this type of radiomic-plus-clinical fusion.","The reported calibration and net-benefit results, if valid, would support deployment as a decision-support tool in postoperative follow-up planning."],"supporting_citations":[{"why":"Prior work combining radiomic features with clinical variables for glioblastoma survival; the integrated approach the paper extends.","marker":"[94]"},{"why":"Transformer-based multimodal glioblastoma survival model with external validation; supplies the comparison point for multi-institutional generalizability.","marker":"[101]"},{"why":"IBSI guidelines that define the 107 radiomic features and standardize their extraction.","marker":"[105]"},{"why":"Cox proportional hazards model whose partial likelihood is used as the training loss for XGBoost and CoxBoost.","marker":"[108]"},{"why":"Dynamic survival modeling with longitudinal data; basis for the paper's proposed temporal encoding module.","marker":"[110]"},{"why":"Decision curve analysis method used to evaluate clinical utility via net benefit.","marker":"[111]"}],"fun_headline_variants":["Multimodal ML predicts early brain tumor recurrence with 0.782 C-index","Fusing MRI and biomarkers sharpens brain tumor recurrence prediction","XGBoost tops four ML models for postoperative brain tumor risk","0.782 C-index: multimodal ML for early brain tumor recurrence","MRI radiomics + biomarkers predict brain tumor recurrence early"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"That the 186 patients described in the Results (glioblastoma and anaplastic astrocytoma) are the same patients whose enrollment, imaging protocol, and follow-up are described in Section 3 (liver resection for hepatocellular carcinoma); the manuscript never reconciles these descriptions.","fun_headline_variants_meta":{"raw":{"variants":["Multimodal ML predicts early brain tumor recurrence with 0.782 C-index","Fusing MRI and biomarkers sharpens brain tumor recurrence prediction","XGBoost tops four ML models for postoperative brain tumor risk","0.782 C-index: multimodal ML for early brain tumor recurrence","MRI radiomics + biomarkers predict brain tumor recurrence early"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000688,"raw_usage":{"total_tokens":2898,"prompt_tokens":630,"completion_tokens":2268,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":374,"completion_tokens_details":{"reasoning_tokens":2179}},"tokens_in":374,"tokens_out":2268,"duration_ms":17699,"temperature":1.0,"reasoning_tokens":2179,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T12:49:21.661382+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Pull the institutional cohort list behind Section 3.1: if the 186 patients underwent hepatic resection with liver MRI and AFP surveillance, the reported brain-tumor recurrence times and C-index cannot be produced from them. Short of that, a reader can test the out-of-sample claim by checking whether any patient was held out before model selection; the Methods only mention internal cross-validation.","supporting_citations":[{"cited_title":"A deep learning-based radiomics model for prediction of survival in glioblastoma multiforme","cited_arxiv_id":null,"evidence_quote":"Prior work combining radiomic features with clinical variables for glioblastoma survival; the integrated approach the paper extends."},{"cited_title":"Comprehensive multimodal deep learning survival prediction enabled by a transformer architecture: A multicenter study in glioblastoma","cited_arxiv_id":null,"evidence_quote":"Transformer-based multimodal glioblastoma survival model with external validation; supplies the comparison point for multi-institutional generalizability."},{"cited_title":"The image biomarker standardization initiative: standard- ized quantitative radiomics for high-throughput image-based phenotyping.Radiology, 295(2):328–338, 2020","cited_arxiv_id":null,"evidence_quote":"IBSI guidelines that define the 107 radiomic features and standardize their extraction."},{"cited_title":"Regression models and life-tables","cited_arxiv_id":null,"evidence_quote":"Cox proportional hazards model whose partial likelihood is used as the training loss for XGBoost and CoxBoost."},{"cited_title":"Dynamic-deephit: A deep learning approach for dynamic survival analysis with competing risks based on longitudinal data","cited_arxiv_id":null,"evidence_quote":"Dynamic survival modeling with longitudinal data; basis for the paper's proposed temporal encoding module."},{"cited_title":"Decision curve analysis: a novel method for evaluating prediction models","cited_arxiv_id":null,"evidence_quote":"Decision curve analysis method used to evaluate clinical utility via net benefit."}],"review_version":1}