REVIEW 2 major objections 6 minor 3 references
An Update on Machine Learning in Neuro-oncology Diagnostics
T0 review · 2 major / 6 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read Brain-tumour machine learning is not ready for the clinic.
desk verdict A well-written narrative review that honestly surveys eight illustrative ML neuro-oncology studies and concludes, with caveats, that the evidence is too low-level for clinical adoption; the main weakness is the non-systematic basis for that field-level conclusion. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The argument is carried by the standard image-analysis pipeline — pre-processing, feature estimation, feature selection, classification, and evaluation — together with two evaluative distinctions. The first distinguishes analytical validation (does the feature measure something accurately and reliably?) from clinical validation (does the biomarker work in a clinical trial?). The second is a levels-of-evidence scale on which retrospective single-centre studies count as low level. The review applies these tools to eight studies and finds the same pattern in each: promising reported accuracy, no prospective multicentre validation, which is what supports the overall verdict.
What would settle it
A completed prospective multicentre trial in which a machine-learning imaging biomarker for brain tumours meets prespecified accuracy targets and demonstrably affects patient management would falsify the blanket 'not ready' conclusion, as would a systematic review with explicit inclusion criteria that located high-level prospective evidence.
Extended reading notes
Core claim
On the paper's own terms, the central claim is a status report: machine learning has been tried at every stage of the neuro-oncology imaging pathway — diagnosis, prognosis, and treatment-response monitoring — and often separates classes accurately in its home institution, but it is not ready for the clinic. The review argues that the level of evidence is low across the board: the illustrative studies are retrospective, mostly single-centre, and none demonstrates prospective multicentre clinical validation. It adds that integrating demographic, clinical, and molecular data with imaging is the most plausible route to better biomarkers, and that large, well-annotated datasets assembled through multidisciplinary and multicentre collaborations are a necessary precondition.
Load-bearing premise
The blanket conclusion rests on the assumption that the eight illustrative studies represent the whole neuro-oncology machine-learning literature; if those examples were selected selectively, or if unpublished or non-English evidence already contains prospective validation, the low-evidence verdict could be wrong.
Editorial extensions
If this is right
- If the claim is right, hospitals should treat machine-learning tumour classifiers as research tools, not routine diagnostic tests, until prospective multicentre validation exists.
- Validation studies should record and report simple clinical variables such as age and performance status, which the review notes can carry as much predictive weight as complex imaging features.
- Combining imaging with demographic, clinical, and molecular data is the direction most likely to improve future biomarker accuracy.
- The field needs shared, well-annotated datasets assembled through multidisciplinary and multicentre collaborations, since single-centre evidence cannot support clinical translation.
- High accuracy on retrospective training sets should not be read as expected clinical performance; an external test set is the minimum check.
Reading between the lines
- Beyond the paper, the same single-centre retrospective pattern is likely to afflict published radiomics studies in other tumour types, so the review's verdict should probably be read as a general caution rather than a neuro-oncology-only result.
- Beyond the paper, several studies' finding that age or simple features rival complex texture features suggests model complexity itself is a risk; an extension would compare simple and deep models on a common multicentre dataset.
- A concrete test the review does not perform: re-run the eight studies' published pipelines on one shared dataset and measure how much accuracy falls from training to external validation, quantifying exactly the single-centre bias the paper warns about.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper is a narrative update on machine learning in neuro-oncology diagnostics. It defines imaging biomarkers, distinguishes diagnostic, monitoring, and prognostic biomarkers, and presents seven illustrative studies using machine learning for IDH/1p19q prediction, glioma grading, pseudoprogression versus progression, and overall survival. The central claim is that most evidence is retrospective and single-centre, the level of evidence is low, and machine learning is not ready for clinical incorporation. The paper concludes with a call for larger, multicentre, prospective validation datasets.
Significance. The paper is a useful, clearly written introduction to the field for neuro-oncology and imaging audiences. It gives concrete examples of analytic pipelines and honestly highlights the limitations of each study, including the need for external validation. The conclusion is clinically important: if correct, it warns against premature adoption of ML-based imaging biomarkers. However, the paper's main limitation is that it is a non-systematic narrative review; the field-level conclusion goes beyond the evidence presented. The paper has no search strategy, inclusion criteria, or quality scoring, so its generalization is not independently verifiable. If the authors revise to either temper the conclusion or add a systematic approach, the paper could become a valuable educational perspective.
major comments (2)
- [Section 5 (also Abstract and Section 1.2)] The conclusion in Section 5 that machine learning in neuro-oncology is 'not ready to be incorporated into the clinic as the level of evidence is low' is a field-level generalization, but the evidence base consists of seven illustrative studies selected without any documented search strategy, inclusion criteria, date range, or quality assessment. Section 1.2 explicitly states that the update 'describes several illustrative research studies,' and the paper does not establish that these studies are representative of the broader literature. The citation to the OCEBM levels of evidence [6] provides a grading scheme but does not by itself establish the distribution of study designs in this field. The authors should either (a) narrow the conclusion so it applies only to the described studies, or (b) add a systematic search and quality appraisal to support the claim about the overall level of evidence.
- [Sections 3 and 4] The manuscript repeatedly characterizes the evidence as 'low' and 'retrospective and single-centre,' yet two of the presented studies include prospective test datasets: Example 1 in Section 4 (Macyszyn et al.) used a prospective test dataset of 29 patients, and Example 1 in Section 3 (Booth et al.) used a prospective test dataset of 7 patients. The text does not explain how these designs affect the OCEBM grading or why the overall evidence remains low despite these prospective components. A per-study classification of each example against the OCEBM levels would make the summary claim transparent and easier to assess.
minor comments (6)
- [Section 1.1] There are typos: 'magetic resonance' should be 'magnetic resonance,' and '1H-magetic' should be '1H-magnetic.'
- [Section 2.2, Example 1] The phrase 'area under the receiving-operator characteristic curve' should be 'receiver operating characteristic curve' (also in the following sentence).
- [Section 1.2] 'Afterall' should be 'After all.'
- [References] Reference [3] appears incomplete: the author list ends with 'Galanis, E.,' and the remaining authors and publication details are missing.
- [General] The example numbers restart in each section (Example 1 in Diagnostics, Example 1 in Monitoring, Example 1 in Prognostic); consider numbering them uniquely (e.g., Example 1.1, 2.1) to avoid ambiguity for readers.
- [General] A summary table listing each example, the biomarker type, the ML method, the dataset size and design, and the reported performance would greatly increase the usability of the review as a reference.
Circularity Check
No circularity: the paper is a narrative update assessing the evidence level of machine learning in neuro-oncology, with no fitted inputs or self-citation-dependent derivation.
full rationale
This manuscript is a narrative review and commentary, not a derivation. It does not introduce equations, fitted parameters, or a predictive model whose output is defined in terms of its inputs. Its central claim—that machine learning in neuro-oncology is not ready for clinical use because the level of evidence is low—is supported by reference to externally cited evidence-level frameworks and by descriptions of illustrative studies, each assessed on its own merits. The only self-citation (Section 3, Example 1, Booth et al. 2017) is presented critically, with the author explicitly stating that the biomarker 'requires clinical validation in a larger multicentre test dataset'; this self-citation is not load-bearing for the conclusion. The paper also discloses that its examples are illustrative rather than a systematic review, which is a limitation in generality but not a circularity. No self-definitional steps, fitted-input-as-prediction, uniqueness imports, ansatz smuggling, or renaming of known results are present. Therefore the circularity score is 0.
Assumptions & free parameters
assumptions (3)
- domain assumption Biomarker validity is best judged by analytical validation followed by clinical validation in prospective trials, following the FDA-NIH BEST framework.
- domain assumption The Oxford 2011 Levels of Evidence apply to imaging biomarker studies, so retrospective single-centre studies count as 'low level' evidence.
- ad hoc to paper The eight illustrative studies are representative of the broader literature on machine learning in neuro-oncology.
Cite this review
Pith. "Pith review of An Update on Machine Learning in Neuro-oncology Diagnostics." pith.science (2026). https://pith.science/paper/I3WFQAKU
@misc{pith2026191008157,
author = {Pith},
title = {Pith review of: An Update on Machine Learning in Neuro-oncology Diagnostics},
year = {2026},
howpublished = {\url{https://pith.science/paper/I3WFQAKU}},
note = {Machine review of arXiv:1910.08157}
}
read the original abstract
Imaging biomarkers in neuro-oncology are used for diagnosis, prognosis and treatment response monitoring. Magnetic resonance imaging is typically used throughout the patient pathway because routine structural imaging provides detailed anatomical and pathological information and advanced techniques provide additional physiological detail. Following image feature extraction, machine learning allows accurate classification in a variety of scenarios. Machine learning also enables image feature extraction de novo although the low prevalence of brain tumours makes such approaches challenging. Much research is applied to determining molecular profiles, histological tumour grade and prognosis at the time that patients first present with a brain tumour. Following treatment, differentiating a treatment response from a post-treatment related effect is clinically important and also an area of study. Most of the evidence is low level having been obtained retrospectively and in single centres.
Reference graph
Works this paper leans on
-
[2]
MacDonald, D., Cascino, T.L., Schold, S.C., Cairncross, J.G.: Response criteria for phase II studies of supratentorial malignant glioma. J Clin Oncol. 8, 1277–1280 (1990). https://doi.org/10.1200/JCO.1990.8.7. 1277
-
[3]
Wen, P.Y., Macdonald, D.R., Reardon, D.A., Cloughesy, T.F., Sorensen, A.G., Galanis, E.,: Updated response assessment criteria for high-grade gliomas: response assessment in 10 neuro-oncology working group. J Clin Oncol. 28, 1963–1972 (2010). https://doi.org/10.1200/JCO.2009.26.3541
-
[4]
Kassner, A., Thornhill, R.E.: Texture Analysis: A Review of Neurologic MR Imaging Ap-plications. Am J Neuroradiol. 31(5), 809-816 (2010) https://doi.org/10.3174/ajnr.A2061 5. Cagney, D.N., Sul, J., Huang, R.Y., Ligon, K.L., Wen, P.Y., Alexander B.M.: The FDA NIH biomarkers, endpoints, and other tools (BEST) resource in neuro-oncology. Neuro Oncol 20(9), 1...
arXiv 2010
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.