{"id":"38aba14a-5c11-4912-a367-933f22e57ab5","arxiv_id":"1910.08157","paper_version":1,"verdict":"UNVERDICTED","confidence":"MODERATE","novelty_score":0.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"A narrative review concludes that machine learning in neuro-oncology diagnostics remains at an early stage, supported mostly by retrospective single-centre studies, and is not ready for clinical adoption.","lead":"This paper reviews recent machine learning studies for brain tumour diagnosis, prognosis, and treatment monitoring using MRI and PET. It concludes that the evidence is too weak and mostly retrospective and single-centre, so machine learning is not yet ready for routine clinical use.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The field-level conclusion rests on an unrepresentative sample; without a systematic search, the claim that the level of evidence is low cannot be established.","rationale":"The reader identified the same load-bearing concern: the conclusion generalizes from a small illustrative selection. I agree with that assessment. The paper is transparent that it is not a systematic review, and it even notes in Section 1.2 that it 'describes several illustrative research studies.' That honesty is a point in its favor, but it does not remove the gap between the evidence presented and the strength of the field-level claim in Section 5. The central assertion is about the whole evidence base ('the level of evidence is low'), while the method only supports a claim about the reviewed studies. A narrative review can legitimately offer a cautious expert synthesis, and the conclusion aligns with the cited examples and with general domain knowledge, so this is not a fatal flaw. However, because the manuscript is intended to inform clinical adoption decisions, the representativeness of the examples is the weakest link. My recommendation is unchanged: the verdict remains UNVERDICTED, reflecting that the paper is a useful but non-systematic review whose headline conclusion is not independently established by a reproducible evidence search.","tokens_in":6247,"tokens_out":3713,"duration_ms":40174,"concrete_test":"Perform a PRISMA-compliant systematic review of MEDLINE and Embase (inception to 2020) using a pre-registered protocol with terms for MRI, machine learning, brain tumour, and prospective or multicentre or external validation. Then tabulate how many studies satisfy each level of the OCEBM diagnostic/prognostic hierarchy, excluding the eight illustrative studies. If at least one prospective, multicentre study with independent external validation and low risk of bias is found for any of the four biomarker classes, the Section 5 conclusion must be qualified; if the search recovers only retrospective single-centre studies, the conclusion is supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 5 concludes that machine learning in neuro-oncology is not ready for the clinic because the level of evidence is low. The load-bearing support is the sample of eight illustrative studies described after Section 1.2, which is explicitly nonsystematic. The manuscript provides no search strategy, inclusion or exclusion criteria, date range, or risk-of-bias assessment, and does not attempt to identify unpublished, non-English, or grey-literature evidence. This matters because the conclusion is a universal claim about the entire evidence base, not just about the chosen examples. Even though the cited OCEBM levels [6] say most such studies are low-level, that citation does not by itself establish that most studies in the field are retrospective single-centre reports; only a systematic survey could. If, for any of the four biomarker types, prospective multicentre external-validation studies exist and were omitted, the central claim overstates the case. The paper honestly discloses its illustrative character, so this is a validity limitation rather than an internal inconsistency, but it is the weakest link in the argument.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper is a narrative update on machine learning in neuro-oncology diagnostics. It defines imaging biomarkers, distinguishes diagnostic, monitoring, and prognostic biomarkers, and presents seven illustrative studies using machine learning for IDH/1p19q prediction, glioma grading, pseudoprogression versus progression, and overall survival. The central claim is that most evidence is retrospective and single-centre, the level of evidence is low, and machine learning is not ready for clinical incorporation. The paper concludes with a call for larger, multicentre, prospective validation datasets.","tokens_in":6522,"tokens_out":4698,"duration_ms":42238,"significance":"The paper is a useful, clearly written introduction to the field for neuro-oncology and imaging audiences. It gives concrete examples of analytic pipelines and honestly highlights the limitations of each study, including the need for external validation. The conclusion is clinically important: if correct, it warns against premature adoption of ML-based imaging biomarkers. However, the paper's main limitation is that it is a non-systematic narrative review; the field-level conclusion goes beyond the evidence presented. The paper has no search strategy, inclusion criteria, or quality scoring, so its generalization is not independently verifiable. If the authors revise to either temper the conclusion or add a systematic approach, the paper could become a valuable educational perspective.","major_comments":[{"comment":"The conclusion in Section 5 that machine learning in neuro-oncology is 'not ready to be incorporated into the clinic as the level of evidence is low' is a field-level generalization, but the evidence base consists of seven illustrative studies selected without any documented search strategy, inclusion criteria, date range, or quality assessment. Section 1.2 explicitly states that the update 'describes several illustrative research studies,' and the paper does not establish that these studies are representative of the broader literature. The citation to the OCEBM levels of evidence [6] provides a grading scheme but does not by itself establish the distribution of study designs in this field. The authors should either (a) narrow the conclusion so it applies only to the described studies, or (b) add a systematic search and quality appraisal to support the claim about the overall level of evidence.","section":"Section 5 (also Abstract and Section 1.2)"},{"comment":"The manuscript repeatedly characterizes the evidence as 'low' and 'retrospective and single-centre,' yet two of the presented studies include prospective test datasets: Example 1 in Section 4 (Macyszyn et al.) used a prospective test dataset of 29 patients, and Example 1 in Section 3 (Booth et al.) used a prospective test dataset of 7 patients. The text does not explain how these designs affect the OCEBM grading or why the overall evidence remains low despite these prospective components. A per-study classification of each example against the OCEBM levels would make the summary claim transparent and easier to assess.","section":"Sections 3 and 4"}],"minor_comments":[{"comment":"There are typos: 'magetic resonance' should be 'magnetic resonance,' and '1H-magetic' should be '1H-magnetic.'","section":"Section 1.1"},{"comment":"The phrase 'area under the receiving-operator characteristic curve' should be 'receiver operating characteristic curve' (also in the following sentence).","section":"Section 2.2, Example 1"},{"comment":"'Afterall' should be 'After all.'","section":"Section 1.2"},{"comment":"Reference [3] appears incomplete: the author list ends with 'Galanis, E.,' and the remaining authors and publication details are missing.","section":"References"},{"comment":"The example numbers restart in each section (Example 1 in Diagnostics, Example 1 in Monitoring, Example 1 in Prognostic); consider numbering them uniquely (e.g., Example 1.1, 2.1) to avoid ambiguity for readers.","section":"General"},{"comment":"A summary table listing each example, the biomarker type, the ML method, the dataset size and design, and the reported performance would greatly increase the usability of the review as a reference.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"This is a single-author narrative review with a self-citation (Example 1 in Section 3) that is described critically; I do not see a citation-integrity problem. The main concern is scope: a non-systematic review making a field-level claim may be better framed as a clinical perspective or educational update. The journal should decide whether the evidence-of-level claim is meant to be a systematic review conclusion; if so, major work is needed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a straightforward narrative review, not original research. It does exactly what it says: it walks through eight illustrative studies using machine learning for diagnosis, monitoring, and prognosis in neuro-oncology, and concludes that the evidence is mostly retrospective, single-centre, and not yet ready for the clinic. If you know the field, there are no surprises; if you don't, it's a useful map.\n\nWhat's new: nothing in the way of data or analysis. The value is in the curation. The author picks studies that span the three biomarker types and gives a genuinely balanced summary, including weaknesses in each. Credit is due for the candid treatment of his own 2017 pseudoprogression paper: he reports the 0.86 accuracy in a prospective test set of 7 patients, then immediately notes the single-centre limitation and the need for multicentre validation. That is how self-citation should be handled.\n\nThe soft spots are real but, for a narrative review, not disqualifying. The biggest one is the one the stress test flags: Section 5 makes a universal claim about the field's evidence level on the basis of eight hand-picked examples, with no search strategy, inclusion criteria, or risk-of-bias assessment. The author does disclose that the studies are 'illustrative' in Section 1.2, so it's not a covert claim. But the conclusion as written—'the level of evidence is low'—is a stronger statement than the sample supports. A systematic review might reach the same conclusion, but this paper doesn't demonstrate it. That said, the paper cites OCEBM levels of evidence [6], and the author's own judgement as a clinical researcher in the area carries weight; this is a limitation in method, not a fatal error in reasoning.\n\nMinor issues: the reference list has small formatting gaps, and the claim in the abstract about low prevalence of brain tumours making de novo feature extraction challenging is asserted without a citation. Neither changes the takeaway.\n\nWho this is for: a clinical radiology or neuro-oncology reader who wants a compact, honest tour of where machine learning stands. An expert in ML will find it basic. The paper is a reasonable candidate for peer review in a journal that handles clinical reviews, provided the reviewers understand it is a narrative review and don't demand systematic-review methodology. In that context, it deserves a serious referee.","headline":"A well-written narrative review that honestly surveys eight illustrative ML neuro-oncology studies and concludes, with caveats, that the evidence is too low-level for clinical adoption; the main weakness is the non-systematic basis for that field-level conclusion.","tokens_in":6883,"tokens_out":2262,"would_cite":false,"duration_ms":21754,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Brain-tumour machine learning is not ready for the clinic.","keywords":["machine learning","neuro-oncology","imaging biomarkers","radiomics","clinical validation","magnetic resonance imaging","glioma","evidence-based medicine"],"falsifier":"A completed prospective multicentre trial in which a machine-learning imaging biomarker for brain tumours meets prespecified accuracy targets and demonstrably affects patient management would falsify the blanket 'not ready' conclusion, as would a systematic review with explicit inclusion criteria that located high-level prospective evidence.","tokens_in":6017,"feed_emoji":"🧠","tokens_out":6944,"duration_ms":66120,"temperature":0.7,"pith_summary":"This review establishes a status report for machine learning in brain-tumour imaging. It argues that machine-learning models can accurately classify tumours, predict molecular markers such as isocitrate dehydrogenase (IDH) status, and separate true progression from treatment effects, but nearly all of the supporting studies are retrospective and from single centres. The paper's conclusion is that the evidence base is too weak for these tools to enter routine clinical care. That conclusion matters because an unvalidated imaging biomarker could guide biopsy, resection, or treatment-response decisions on the strength of training-set accuracy alone.","feed_headline":"Brain-tumour machine learning is not ready for the clinic","feed_subtitle":"A review of eight illustrative studies finds most evidence is retrospective, single-centre, and unvalidated.","key_machinery":"The argument is carried by the standard image-analysis pipeline — pre-processing, feature estimation, feature selection, classification, and evaluation — together with two evaluative distinctions. The first distinguishes analytical validation (does the feature measure something accurately and reliably?) from clinical validation (does the biomarker work in a clinical trial?). The second is a levels-of-evidence scale on which retrospective single-centre studies count as low level. The review applies these tools to eight studies and finds the same pattern in each: promising reported accuracy, no prospective multicentre validation, which is what supports the overall verdict.","core_discovery":"On the paper's own terms, the central claim is a status report: machine learning has been tried at every stage of the neuro-oncology imaging pathway — diagnosis, prognosis, and treatment-response monitoring — and often separates classes accurately in its home institution, but it is not ready for the clinic. The review argues that the level of evidence is low across the board: the illustrative studies are retrospective, mostly single-centre, and none demonstrates prospective multicentre clinical validation. It adds that integrating demographic, clinical, and molecular data with imaging is the most plausible route to better biomarkers, and that large, well-annotated datasets assembled through multidisciplinary and multicentre collaborations are a necessary precondition.","pith_inferences":["Beyond the paper, the same single-centre retrospective pattern is likely to afflict published radiomics studies in other tumour types, so the review's verdict should probably be read as a general caution rather than a neuro-oncology-only result.","Beyond the paper, several studies' finding that age or simple features rival complex texture features suggests model complexity itself is a risk; an extension would compare simple and deep models on a common multicentre dataset.","A concrete test the review does not perform: re-run the eight studies' published pipelines on one shared dataset and measure how much accuracy falls from training to external validation, quantifying exactly the single-centre bias the paper warns about."],"forward_implications":["If the claim is right, hospitals should treat machine-learning tumour classifiers as research tools, not routine diagnostic tests, until prospective multicentre validation exists.","Validation studies should record and report simple clinical variables such as age and performance status, which the review notes can carry as much predictive weight as complex imaging features.","Combining imaging with demographic, clinical, and molecular data is the direction most likely to improve future biomarker accuracy.","The field needs shared, well-annotated datasets assembled through multidisciplinary and multicentre collaborations, since single-centre evidence cannot support clinical translation.","High accuracy on retrospective training sets should not be read as expected clinical performance; an external test set is the minimum check."],"supporting_citations":[{"why":"Supplies the analytical-versus-clinical validation distinction used to grade every study in the review.","marker":"[5]"},{"why":"Supplies the levels-of-evidence scale on which the low-evidence verdict rests.","marker":"[6]"},{"why":"Illustrative diagnostic study predicting IDH mutation status from routine MRI with age as the strongest feature; retrospective and single-centre.","marker":"[8]"},{"why":"Illustrative diagnostic study predicting molecular markers and grade; logistic-regression texture models beat random-forest combinations.","marker":"[9]"},{"why":"Illustrative diagnostic study using unsupervised clustering to grade gliomas; small retrospective series with limited prospective testing.","marker":"[10]"},{"why":"Illustrative monitoring study separating progression from pseudoprogression on T2-weighted MRI; retrospective training with a small prospective test set.","marker":"[11]"},{"why":"Illustrative monitoring study using PET textural features to separate progression from pseudoprogression; small retrospective series.","marker":"[12]"},{"why":"Illustrative prognostic study predicting glioblastoma survival; retrospective training plus prospective test set at one centre.","marker":"[13]"},{"why":"Illustrative prognostic study using convolutional-neural-network features; high training accuracy but low test accuracy, flagged as overfitting.","marker":"[14]"}],"fun_headline_variants":["Machine learning in brain tumours: promise, not proof","Brain-tumour ML: accurate at home, unproven in clinic","Neuro-oncology ML: retrospective, single-centre, unvalidated","Machine learning in neuro-oncology lacks clinical proof"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The blanket conclusion rests on the assumption that the eight illustrative studies represent the whole neuro-oncology machine-learning literature; if those examples were selected selectively, or if unpublished or non-English evidence already contains prospective validation, the low-evidence verdict could be wrong.","fun_headline_variants_meta":{"raw":{"variants":["Machine learning in brain tumours: promise, not proof","Brain-tumour ML: accurate at home, unproven in clinic","Neuro-oncology ML: retrospective, single-centre, unvalidated","Machine learning in neuro-oncology lacks clinical proof"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000637,"raw_usage":{"total_tokens":2858,"prompt_tokens":791,"completion_tokens":2067,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":407,"completion_tokens_details":{"reasoning_tokens":1996}},"tokens_in":407,"tokens_out":2067,"duration_ms":15478,"temperature":1.0,"reasoning_tokens":1996,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T14:14:07.702517+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A completed prospective multicentre trial in which a machine-learning imaging biomarker for brain tumours meets prespecified accuracy targets and demonstrably affects patient management would falsify the blanket 'not ready' conclusion, as would a systematic review with explicit inclusion criteria that located high-level prospective evidence.","supporting_citations":[],"review_version":1}