{"id":"14ce27ec-e7ca-4174-acc1-2bf5370fea28","arxiv_id":"2505.11041","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"high","formal_verification":"none","parameter_count":4,"one_line_summary":"The paper reports an MLP classifier on 25 CpG methylation features that distinguishes colorectal cancer from normal cfDNA with AUROC 0.89 on a hold-out set, but the validation is compromised by feature selection on the full dataset.","lead":"This study trains machine learning models on DNA methylation data from blood samples to distinguish colorectal cancer patients from healthy controls, reporting an AUROC of 0.89 for the best model. The work is a standard biomarker pipeline on a public dataset, but its validation strategy has leaks that likely inflate the reported performance.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The validation AUROC of 0.89 is not an independent performance estimate: Section 4.2.4 reports the best-performing model selected on the validation set, and Section 3.4 performs feature ranking before the split, so the held-out set influenced both feature and model selection.","rationale":"The reader's weakest assumption points to feature-selection leakage: Section 3.4 ranks and selects CpG sites before the 80:20 split in Section 3.6. I agree this is a real and load-bearing flaw. I also find an independent and more direct statement in Section 4.2.4: the paper reports the best-performing model over the validation dataset, meaning the 55-sample validation set was used to choose among many feature-selection and classifier combinations. The AUROC of 0.89 is therefore a selected maximum over a large model grid, not a single pre-specified model's performance on untouched data. This is a second violation of the same independence assumption, and it compounds the feature-leakage problem. The missing GSE149438 external validation and the inconsistent conclusion panel (different genes and a random forest model) strengthen the case that the headline result is not a reliable estimate of generalization, but they are secondary to the validation-set contamination. My agreement is partial because the reader emphasized feature-selection leakage while the explicit model-selection-on-validation statement is arguably the more decisive evidence. The reader's REJECT verdict should stand unchanged: the central claim cannot be accepted as an independent validation result in its current form.","tokens_in":61442,"tokens_out":5200,"duration_ms":50225,"concrete_test":"Re-run the pipeline with the 55 validation samples fully sequestered from all feature and model selection. Specifically: (1) compute univariate rankings and the top-100 candidate CpG set using only the 219 training samples; (2) for each feature-selection method and feature count, tune the MLP and other classifiers by nested 5-fold CV on the training set only; (3) pre-register the single final configuration; (4) apply it once to the held-out 55 samples and compute AUROC and MCC. If the resulting held-out AUROC is below 0.89 (e.g., below 0.80), the headline estimate is inflated by validation-set selection. As a secondary check, report the distribution of validation AUROC across all configurations in Table 4; if 0.89 is the maximum of the grid, selection bias is directly demonstrated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is the MLP/RFE-25 AUROC of 0.89 on the 55-sample validation set. For this to be a valid generalization estimate, the validation set must not influence any modeling choice. Two passages show it did. First, Section 4.2.4 states 'we have reported the best-performing model over the validation dataset', and Table 4 lists the maximum AUROC across a large grid of configurations (RFE/SFS/SVC-L1, 25/20/15/10/5 features, multiple classifiers). Reporting the maximum over this grid evaluated on the same 55 samples makes 0.89 a selected maximum, not an independent test result. Second, Section 3.4 selects the top 100 positive and negative CpG sites using univariate AUROC before the 80:20 split described in Section 3.6; RFE then operates on these full-data-selected candidates. Thus the validation samples contributed to choosing the 25 features. Either flaw alone invalidates the independence assumption; together they undermine the AUROC of 0.89 and MCC of 0.78 as unbiased estimates. The conclusion's discrepancy, describing a different 15-site RF panel, further obscures which model is actually being claimed.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript (arXiv:2505.11041) proposes a machine-learning pipeline for colorectal cancer detection from cell-free DNA methylation profiles. Using the GSE124600 cohort (142 CRC, 132 normal), the authors retain 97,863 CpG sites, perform a t-test to select 30,791 putatively significant sites, rank the top 100 positively and negatively correlated sites by univariate AUROC, and then apply RFE, SFS, and SVC-L1 feature selection followed by 11 classifiers plus a CNN. The central reported result is an MLP trained on 25 RFE-selected positive-correlation features with AUROC 0.89 and MCC 0.78 on a randomly held-out 55-sample validation set. The paper also provides functional annotation of the selected CpG sites and GEPIA2-based gene-expression validation for seven genes.","tokens_in":61719,"tokens_out":3835,"duration_ms":37738,"significance":"If the reported performance were a valid out-of-sample estimate, a cfDNA methylation-based CRC classifier would have clear clinical screening value. The study uses a well-known public dataset and provides extensive supplementary tables, which is commendable. However, the central performance claim is not a valid generalization estimate because the validation set influenced both feature selection and model selection, and because the manuscript contains a major inconsistency between the abstract/Table 4 model and the conclusion's panel. A correct re-analysis with proper nested validation and an external cohort would be needed for this result to be interpretable.","major_comments":[{"comment":"Feature selection is performed before the train/validation split: Section 3.4 ranks all 30,791 sites and selects the top 100 positively and negatively correlated sites using univariate AUROC computed on the full 274-sample dataset, and only Section 3.6 describes the 80:20 split that creates the 55-sample validation set. Because the validation samples contributed to choosing the candidate CpG sites that enter RFE/SFS/SVC-L1, the AUROC of 0.89 in Table 4 (RFE-positive, MLP-25) is not an independent out-of-sample estimate, and the abstract's claim of 'independent validation datasets' is not supported.","section":"Section 3.4 vs 3.6"},{"comment":"The manuscript states 'we have reported the best-performing model over the validation dataset' after evaluating RFE, SFS, and SVC-L1 with feature counts of 25, 20, 15, 10, and 5 across 11 classifiers. Selecting the maximum AUROC on the same 55 validation samples is a form of test-set selection ('winner's curse'); therefore the reported AUROC 0.89 and MCC 0.78 are selected maxima over a large grid of correlated estimates rather than the expected performance of a pre-specified model.","section":"Section 4.2.4 and Table 4"},{"comment":"The t-test over 97,863 CpG sites with a p-value threshold of 0.05 and no multiple-testing correction is reported as yielding 30,791 significant sites. Under the null hypothesis one would expect approximately 4,893 sites to pass this threshold by chance, so the claim of 30,791 'significantly altered' sites is likely dominated by false positives. This does not by itself invalidate the classifier, but it affects the biological interpretation and the feature-ranking step that starts from these sites.","section":"Section 3.3 and Section 4.1.1"},{"comment":"The conclusion describes a 'panel of 15 upregulated methylation sites on EVC, FRMD6, FRMD6-AS2, LHFPL6, LIFR, and ZFPM2' and a random forest model with specificity 86.21% and sensitivity 88.46%. Neither the panel composition nor the model matches the abstract's best-performing model (MLP with 25 RFE-selected positive-correlation sites, Table 4), and the genes FRMD6, LHFPL6, LIFR, and ZFPM2 do not appear in Table 6's list of the ten annotated genes for the best model. The manuscript therefore does not state a single, consistent best model.","section":"Section 3.5 (Conclusion) vs Abstract/Table 4"},{"comment":"The text promises that 'a cross-platform validation was also conducted using data from a different colorectal cancer study, GSE149438,' but no results for GSE149438 appear anywhere in the Results or Supplementary Tables. Either the cross-platform validation results must be reported and discussed, or the sentence describing this validation should be removed.","section":"Section 3.6"}],"minor_comments":[{"comment":"The title and abstract call the work an 'in silico tool,' but no software, web server, or code repository is described or made available; please clarify what is being delivered.","section":"Title and abstract"},{"comment":"There are two sections numbered '3.5' (one in the methodology and one after Section 4.2.6 labeled '3.5. Conclusion'), and the conclusion is placed after the results rather than in its own numbered position; please renumber and restructure.","section":"Section numbering"},{"comment":"The SVC-L1 row for negatively correlated sites in Table 4 shows 'NB 50' with 50 features, while the text states 'SVC-based ML model developed using 4 methylated sites extracted using SFS'; the row labels and the text are hard to reconcile and should be aligned.","section":"Table 4 and Section 4.2.4"},{"comment":"Supplementary Table 7.1 is empty except for headers; the SFS feature lists are embedded in the caption of Table 7.2. Please move the feature lists into Table 7.1 so that each supplementary table is self-contained.","section":"Supplementary Table 7.1"},{"comment":"The equations for sensitivity, specificity, accuracy, and MCC appear garbled in the manuscript source; please ensure they are typeset correctly with proper numerators, denominators, and parentheses.","section":"Equations (1)-(4)"},{"comment":"The abstract states the models were 'trained and tested using independent validation datasets,' but the validation set is a random 20% split of the same GSE124600 cohort (Section 3.6), not an independent cohort; please rephrase to avoid overclaiming.","section":"Abstract wording"}],"recommendation":"reject","confidential_remarks":"This manuscript's central performance claim (MLP/RFE-25, AUROC 0.89) is not a valid out-of-sample estimate because feature selection occurs before the data split and because the best model is selected on the validation set. The conclusion describes a different panel and model than the abstract and Table 4, and the promised GSE149438 validation is absent. These issues are load-bearing and would require re-running the entire analysis with a proper nested validation design and a genuinely external cohort; that is beyond the scope of a revision. I therefore recommend rejection, while acknowledging the substantial supplementary data and the potential clinical relevance of the underlying question."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe headline is simple: the central performance claim is not trustworthy as an out-of-sample result. The MLP with 25 RFE-selected features reaches AUROC 0.89 and MCC 0.78 on the 55-sample validation set, but two things break the independence assumption. First, Section 3.4 ranks all CpGs and selects the top 100 positive/negative sites before the 80:20 split in Section 3.6; the validation samples thus helped pick the candidate features that RFE later refined. Second, Section 4.2.4 reports the best-performing model over the validation dataset across a grid of feature selectors, feature counts, and classifiers, so 0.89 is the maximum of a large model search evaluated on the same 55 samples, not a single pre-specified model's score. Either issue alone is enough to make the number optimistic; together they mean the abstract's headline result is essentially a fitted value.\n\nWhat the paper does well: it uses a relevant public cfDNA methylation dataset (GSE124600), runs a broad sweep of standard classifiers and feature selectors, and provides large supplementary tables with performance for many configurations. The functional annotation and GEPIA2 cross-check on the selected genes is a reasonable attempt at biological plausibility.\n\nWhere it falls down, beyond the leakage: the promised cross-platform validation on GSE149438 is never shown. The conclusion section describes a completely different panel—15 sites on EVC, FRMD6, LIFR, etc. with a random forest—while the abstract and results push the MLP-25. No code, hyperparameters, or thresholds are given, so the 'tool' is not actually usable by others. Multiple testing on ~98k CpGs is not addressed.\n\nThis is a routine computational pipeline, not a new method, and the conclusions are not supported in the current form. But the flaws are fixable: redo feature selection inside the cross-validation loop, report the performance of a pre-specified model, and align the abstract with the conclusion. That would require a major revision, not a touch-up. I'd send it to peer review rather than desk reject, because a serious referee could force the re-analysis and the dataset is of interest. But I wouldn't trust the current numbers.","headline":"The reported AUROC of 0.89 is a selected maximum, not an independent estimate; the paper is a routine pipeline with a fixable but load-bearing validity problem.","tokens_in":62312,"tokens_out":3164,"would_cite":false,"duration_ms":32562,"reading_group":"maybe","serious_thinker":"no","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A 25-site blood DNA methylation panel detected colorectal cancer with AUROC 0.89.","keywords":["colorectal cancer","cell-free DNA","DNA methylation","liquid biopsy","machine learning","biomarker panel","CpG sites","noninvasive screening"],"falsifier":"Re-run the entire pipeline with all feature selection restricted to the 219 training samples, then evaluate the fixed model on the 55 held-out samples; if the AUROC falls from 0.89 toward the 0.71 achieved by the all-sites model, the claimed advantage of the compact 25-site signature would not be robust.","tokens_in":61266,"feed_emoji":"🧬","tokens_out":10624,"duration_ms":106417,"temperature":0.7,"pith_summary":"This paper claims that methylation marks in cell-free DNA from a blood sample can identify colorectal cancer without a tissue biopsy. The authors start with a public dataset of 274 blood samples, find about 31,000 methylation sites that differ between cancer and normal, and reduce them to a compact 25-site signature. Their best model, a multi-layer perceptron trained on those 25 sites, reached an AUROC of 0.89 and a Matthews correlation coefficient of 0.78 on a held-out set of 55 samples, with 93% sensitivity and 85% specificity. A convolutional neural network using a much larger site set reached 0.78 AUROC. If the result holds, a simple blood test could serve as a triage step before colonoscopy, potentially widening screening coverage.","feed_headline":"Blood DNA test flags colorectal cancer at AUROC 0.89","feed_subtitle":"A 25-site cell-free DNA methylation panel reaches 93% sensitivity and 85% specificity in a 55-sample holdout.","key_machinery":"The load-bearing object is a 25-CpG methylation signature selected by recursive feature elimination (RFE), an iterative procedure that trains a model, ranks features by importance, deletes the weakest, and repeats. RFE was applied to the top 100 positively correlated methylation sites chosen by univariate AUROC scoring. The resulting signature maps to genes including EVC, SLIT3, SDK1, and ADGRB1, and the paper shows that several corresponding genes are differentially expressed in colon and rectum adenocarcinomas. The multi-layer perceptron, a feed-forward neural network, carries the final classification, but the 25-site signature is the component that would be translated into a targeted clinical test.","core_discovery":"The paper's central claim is that DNA methylation profiles of circulating cell-free DNA carry enough cancer-specific information for a machine-learning classifier to separate colorectal cancer patients from healthy controls. On a discovery cohort of 142 cancer and 132 control blood samples profiled with a targeted methylated-CpG amplification sequencing assay, the authors report that a multi-layer perceptron trained on 25 CpG sites selected by recursive feature elimination achieves an AUROC of 0.89 and an MCC of 0.78 on the held-out 55-sample validation set, with 93.1% sensitivity and 84.6% specificity. A convolutional neural network trained on roughly 30,000 methylation sites reached an AUROC of 0.78 on the same validation set. The concluding summary also highlights a smaller 15-site panel run through a random forest, with 88.5% sensitivity and 86.2% specificity, which the authors describe as comparable to the approved methylated SEPT9 blood test.","pith_inferences":["A natural next experiment, not reported in this paper, is to freeze the 25-site panel and run it on an independent cohort collected with a different methylation assay; that would show whether the signature is portable across platforms.","The same pipeline could be applied to other cancer types where cell-free DNA methylation datasets exist, potentially yielding one blood-drawn panel that screens several cancers at once.","The 15-site random forest described in the conclusion and the 25-site MLP described in the abstract may be complementary; a consensus signature across both could be more stable than either panel alone."],"forward_implications":["A blood draw could replace or precede stool-based and invasive screening in settings where colonoscopy is not readily available, since the model needs methylation levels at only a small number of CpG sites.","The 25-site panel is small enough for targeted assays such as methylation-specific PCR, making clinical deployment cheaper than whole-methylome sequencing.","At 93.1% sensitivity and 84.6% specificity, the MLP panel sits in the same performance range as the methylated SEPT9 blood test cited in the paper, offering an alternative marker set for the same screening purpose.","Seven of the signature's ten annotated genes show differential expression in colon or rectum adenocarcinoma, giving the classifier a biological rationale beyond pattern matching."],"supporting_citations":[{"why":"Defines the clinical benchmark (methylated SEPT9, about 90% recall and 88% specificity) against which the paper measures its claimed diagnostic performance.","marker":"[25,26]"},{"why":"Underpins the central premise that aberrant methylation in circulating tumor or cell-free DNA can distinguish cancer from normal samples.","marker":"[29,30]"},{"why":"Provides the external gene-expression validation source used to claim that several signature genes are differentially expressed in colon and rectum adenocarcinoma.","marker":"[34]"},{"why":"Connects one signature gene (ADGRB1/BAI1) to reduced expression and metastasis in colorectal cancer, supporting the panel's biological plausibility.","marker":"[44]"},{"why":"Links another signature gene (HSD17B12) to colorectal cancer outcome, supporting the functional interpretation of the selected CpG sites.","marker":"[51]"}],"fun_headline_variants":["Blood DNA methylation test IDs colorectal cancer with AUROC 0.89","AI on cell-free DNA methylation flags colorectal cancer (AUROC 0.89)","25 blood CpG sites predict colorectal cancer: AUROC 0.89","cfDNA methylation panel detects colorectal cancer with 93% sensitivity","ML on blood DNA methylation spots colorectal cancer (AUROC 0.89)"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The 55 validation samples are treated as an untouched independent check, but they were used to rank and select the top CpG sites before the training/validation split was made, so the reported 0.89 AUROC is probably optimistic.","fun_headline_variants_meta":{"raw":{"variants":["Blood DNA methylation test IDs colorectal cancer with AUROC 0.89","AI on cell-free DNA methylation flags colorectal cancer (AUROC 0.89)","25 blood CpG sites predict colorectal cancer: AUROC 0.89","cfDNA methylation panel detects colorectal cancer with 93% sensitivity","ML on blood DNA methylation spots colorectal cancer (AUROC 0.89)"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000954,"raw_usage":{"total_tokens":4120,"prompt_tokens":1047,"completion_tokens":3073,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":663,"completion_tokens_details":{"reasoning_tokens":2975}},"tokens_in":663,"tokens_out":3073,"duration_ms":21433,"temperature":1.0,"reasoning_tokens":2975,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T20:58:33.748875+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the entire pipeline with all feature selection restricted to the 219 training samples, then evaluate the fixed model on the 55 held-out samples; if the AUROC falls from 0.89 toward the 0.71 achieved by the all-sites model, the claimed advantage of the compact 25-site signature would not be robust.","supporting_citations":[{"cited_title":"Fukushima, Y","cited_arxiv_id":null,"evidence_quote":"Connects one signature gene (ADGRB1/BAI1) to reduced expression and metastasis in colorectal cancer, supporting the panel's biological plausibility."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Links another signature gene (HSD17B12) to colorectal cancer outcome, supporting the functional interpretation of the selected CpG sites."}],"review_version":1}