{"id":"eb44927a-40b7-46d3-a6e2-2721e760494f","arxiv_id":"2509.08012","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A fully automated deep learning tool predicted the 0-39 global cortical atrophy score from routine older-patient CT brain scans with mean absolute error 3.2 and moderate agreement (kappa 0.45) versus the human rater whose labels were used for training.","lead":"A new deep learning tool scores brain atrophy from routine CT scans on a 39-point global cortical atrophy scale without any human input. In a validation on 864 older patients, the tool matched or outperformed human-human agreement, but the ground truth was a single trained rater and no code or data was released.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Primary accuracy metrics may include training/optimisation scans; test-set-only MAE/kappa must be reported before the 'accurate' claim is supportable.","rationale":"The reader's weakest assumption was that a single rater's scores are a valid ground truth; that is an important external-validity concern. My concern is more direct and internal: the paper's primary numeric evidence may not be from the held-out test set at all. The Results section reports MAE 'for the dataset overall' after describing all 864 scans, and the only explicit test-set statistic appears later for the error histogram. If the headline MAE/kappa include training and optimisation scans, the agreement is partly in-sample and the central claim 'measured GCA score accurately' is not established, regardless of how reliable the reference rater is. The proposed check (partition-stratified recomputation) would settle this immediately. I therefore keep the reader's CONDITIONAL verdict: the paper should not be accepted as-is until this ambiguity is resolved, but the concern is addressable and does not by itself prove the tool is inaccurate. Agreement is partial because the reader flagged the rater-2 subset's split location but did not identify the possible whole-cohort contamination of the primary metrics.","tokens_in":12467,"tokens_out":6988,"duration_ms":79608,"concrete_test":"Ask the authors to recompute from saved predictions, separately for the training (n=518), optimisation (n=173) and testing (n=173) partitions: MAE, Cohen's weighted kappa, Bland-Altman limits of agreement, and three-class severity accuracy against rater-1, and to state which partition contains the 20 rater-2 scans. If the published 3.2 MAE / 0.45 kappa are not the testing-partition values, or the rater-2 subset lies outside the test set, the 'accurate' conclusion is not supported. The key comparison is published overall numbers vs test-set-only numbers.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that the DL tool 'measured GCA score accurately' is only meaningful if the reported agreement statistics are computed on scans not used to fit the model. Methods specify a 60/20/20 split and say 'tool testing using the testing set scored by rater-1', but the Results paragraph introducing the headline numbers reads: 'Among 864 patients ... the MAE ... was 3.2 for the dataset overall, 3.1 for ORCHARD-EPR, 3.3 for OCS, and 2.6 for the legacy scans.' It does not restrict these values to the 173-scan test set; the phrase 'dataset overall' and the n=864 framing suggest all scans are included. The kappa=0.45 for DL-tool vs rater-1 is likewise not tied to a partition. Only later is the test set explicitly invoked: 'For 88 (50%) of the testing set CT scans...' If the reported MAE and kappa include training/optimisation scans, they are in-sample fits to rater-1's labels and cannot support generalisable accuracy. The rater-2 comparison (n=20) also lacks a stated partition; if those 20 scans were in training/optimisation, the comparison is not a held-out test. This ambiguity is the most load-bearing unresolved point because it determines whether the central empirical claim is about generalisation or about memory.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper validates a fully automated deep-learning tool that predicts the Global Cortical Atrophy (GCA) score (0–39) from routine clinical CT brain scans of older patients. It uses 864 scans from an acute-medicine cohort, an acute-stroke cohort, and a legacy sample, with a stated 60/20/20 train/optimisation/test split. Visual GCA ratings by rater-1 on all 864 scans served as training labels; rater-2 rated a 20-scan subset. The reported agreement between the DL tool and rater-1 is MAE 3.2 and weighted kappa 0.45, with kappa 0.41 between tool and rater-2 and 0.28 between the two human raters. The abstract concludes that the tool 'measured GCA score accurately and without user input.'","tokens_in":12764,"tokens_out":4964,"duration_ms":60120,"significance":"Automated, fast (4 s/scan) GCA scoring on routine CT would be a genuinely useful tool for large-scale health-data research and potentially for clinical workflow, particularly because the GCA score is familiar to clinicians and the tool requires no user input. The study draws on real-world, multi-cohort CT data and provides a detailed operational protocol for the GCA scale, which is a useful contribution. However, the significance of the quantitative claims depends directly on whether the headline MAE/kappa values are computed on held-out test scans or include training/optimisation scans, and on whether a single rater's labels are a sufficiently reliable reference standard. The paper's own inter-rater data (rater-1 vs rater-2 kappa = 0.28 on 20 scans) make the latter concern concrete.","major_comments":[{"comment":"The headline metrics are not tied to the 173-scan test set. Methods state a 60/20/20 split, but Results introduce the MAE with 'Among 864 patients ... MAE ... was 3.2 for the dataset overall', and the kappa values are not labelled by partition. Only the percentile error sentence ('For 88 (50%) of the testing set CT scans...') explicitly invokes the test set. If these MAE and kappa values include training/optimisation scans, they are in-sample fits to rater-1's labels and cannot support the generalisability claim. The authors must report test-set-only MAE and kappa, with subgroup values, and state explicitly whether the rater-2 subset of 20 was part of training/optimisation/test.","section":"Results, first paragraph; Methods 'Prediction of GCA scores using a DL model'"},{"comment":"The reference standard is a single rater: 'GCA scores from rater-1 were used as the ground truth during the DL-tool training', and all 864 validation labels come from that same rater. The paper's own inter-rater agreement, kappa = 0.28 (fair) and MAE = 5.2 between rater-1 and rater-2 on 20 scans, shows that the visual reference is not stable across raters. Agreement of the DL tool with rater-1 therefore partly measures how well the model learned that rater's idiosyncrasies, not an independent ground truth. The manuscript acknowledges this in the Limitations, but the abstract's 'accurately' claim is stronger than the evidence supports. Add intra-rater reliability for rater-1 and a larger multi-rater held-out validation.","section":"Methods 'Application of the GCA scale'; Discussion, Limitations"},{"comment":"The text uses non-significant paired t-tests and ANOVA (e.g., t = -0.43, p = 0.66; F = 1.06, p = 0.35) to support 'no difference' and thus agreement. Absence of a statistically significant difference is not evidence of equivalence, particularly in the n = 20 rater-2 subset. The agreement metrics (MAE, Bland–Altman limits, kappa) are more appropriate, but they should be reported for the test set and with confidence intervals; the p-values should not be the basis for claiming interchangeability.","section":"Results, Bland–Altman and statistical tests"}],"minor_comments":[{"comment":"The abstract states 'Among 864 scans ... MAE ... was 3.2' without indicating that this is test-set-only. If the 3.2 MAE is in fact test-set-only, the text should say so explicitly; if not, it should be replaced with the test-set value.","section":"Abstract and Results"},{"comment":"'The accuracy of the DL-tool was 73% for mild atrophy and 70% for moderate and 70% for severe atrophy' is ambiguous. Is this per-class sensitivity, recall, or overall accuracy conditioned on true class? Please define and report cell counts, not just normalised percentages.","section":"Results, classification accuracy"},{"comment":"The repeated-measures one-way ANOVA comparing DL-tool, rater-1, and rater-2 is presumably on the 20 scans rated by rater-2, but the sample size is not stated in the Results. Please state n for each statistical comparison.","section":"Methods, Statistical Analysis"},{"comment":"Table 1 lists 'range=102-65 years' for ORCHARD-EPR; the order is reversed. Figure 1 caption refers to 'EC' but should identify rater-1 by the same label used elsewhere. Figure 6D legend has a typo: 'imapaired'.","section":"Table 1 and Figure 1"},{"comment":"The operationalisation protocol says 'the following scoring criteria has been developed' and provides a helpful manual, but the reliability results from the cited Hobden et al. study are not reproduced; since this protocol is central to the reference standard, a brief summary of those intra/inter-rater statistics would strengthen the paper.","section":"Supplementary Methods"}],"recommendation":"major_revision","confidential_remarks":"The most important request to the authors should be a strict separation of training/validation/test for every reported accuracy metric. If the headline MAE/kappa already are test-set-only, the paper is much closer to acceptable, but the reporting must remove all ambiguity. If the authors cannot provide test-set-only values, the central claim is unsupported and I would not be able to recommend acceptance. The single-rater ground truth is a substantive limitation; the rater-2 subset is useful evidence but too small and currently lacks a stated partition."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The short version: this is a practical, real-world validation of a CT-only deep learning tool that directly outputs the full 0-39 GCA score, which prior work hasn't done. The four-second runtime and the agreement numbers are attractive. But the paper as written does not let you verify the headline accuracy on the held-out test set, and the reference standard is one rater whose own inter-rater agreement is only fair. These are fixable, but they are load-bearing.\n\nWhat's genuinely new: ref 20 needed paired MRI and used a compressed 0-3 scale; this tool trains directly on CT with the full GCA scale on 864 real clinical scans from acute medicine and stroke. The supplement's operationalization of GCA on CT is a useful contribution on its own—it turns a vague visual scale into something reproducible. The tool-versus-human agreement (kappa 0.45 vs rater-1, 0.41 vs rater-2) compares favorably to human-human agreement (0.28). The Bland-Altman bias is near zero, and the correlations with age and cognition point the expected way. The authors are also upfront in the Discussion that training on a single rater means the model likely reflects that rater's systematic errors.\n\nThe soft spots, in proportion. The biggest one: the Results say \"Among 864 patients ... MAE ... was 3.2 for the dataset overall.\" That reads as computed on all 864, including training and optimisation scans. The Methods say testing was on the 173-scan test set, but the test-set MAE and kappa are never reported. The only test-set number is the error distribution (50% within ±2). Until the held-out MAE/kappa are reported, the \"measured accurately\" claim rests on in-sample fit. That is the difference between generalization and memory, and the stress-test note is right to put it front and center.\n\nSecond, the reference standard. Rater-1's labels are both the training target and the test reference. With kappa 0.28 between rater-1 and rater-2 on 20 scans, the human \"truth\" is not stable. The paper acknowledges this, but the abstract's \"accurately\" is too strong given that reference. The rater-2 comparison helps, but the paper doesn't say whether those 20 scans were in the training/optimisation split, which matters for interpreting kappa 0.41.\n\nThird, no code, architecture, or data release, so no independent replication. For a validation study that's a real constraint, though not a reason to reject.\n\nWho is this for? Researchers working on CT-based biomarkers or clinical AI validation. A careful reader will learn from the operationalization protocol and the cohort design, and the test-set ambiguity is a textbook example of why reporting standards matter. But as written, the central claim is not yet supported.\n\nMy recommendation: send to peer review—this deserves a serious referee. Ask the authors to report all metrics separately for training, optimisation, and test, and to clarify where the 20 rater-2 scans sit in the split. If they do that, the tool could be a useful contribution. I wouldn't cite it in its current form.","headline":"Useful proof-of-concept for automated CT-based GCA scoring, but the headline MAE/kappa appear to be computed on the full dataset rather than the held-out test set, and the single-rater reference standard has fair human-human agreement; both need to be fixed before the accuracy claim holds.","tokens_in":13307,"tokens_out":3658,"would_cite":false,"duration_ms":42003,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper validates a fully automated deep-learning tool that measures the Global Cortical Atrophy score on routine CT-brain scans in about four seconds, with error against trained human raters comparable to or smaller than the disagreemen","keywords":["global cortical atrophy","GCA scale","deep learning","CT brain imaging","automated atrophy measurement","visual rating scale validation","older patients","cognition"],"falsifier":"Have three or more independent experts rate a common set of, say, 100 CT scans and compare the tool's scores against their consensus; if the tool's weighted kappa against the consensus is no better than the inter-expert kappa, the claim of human-level measurement fails. A cheaper check: retest on scans with GCA below 3 or above 28, where the paper already reports systematic over- and under-estimation.","tokens_in":12350,"feed_emoji":"🧠","tokens_out":7515,"duration_ms":78413,"temperature":0.7,"pith_summary":"The paper asks whether a deep-learning tool can replace slow, subjective visual rating of brain atrophy on CT scans with a fast, automatic number that clinicians and researchers can trust. It validates a tool that predicts the Global Cortical Atrophy (GCA) score, a 0-39 visual scale, directly from routine CT images of patients over 65, without any user input. Across 864 real-world scans from acute medicine and stroke patients, the tool's mean absolute error against the primary rater was 3.2 points, its agreement was moderate (weighted kappa 0.45), and half of its predictions fell within two points of the rater. Crucially, the tool agreed with human raters at least as well as the two human raters agreed with each other (kappa 0.28), and it produced each score in four seconds versus about three minutes by hand. If these results hold, the tool would make standardised atrophy measurement feasible at scale on the world's most common brain imaging modality.","feed_headline":"Four-second CT tool scores brain atrophy as well as human raters","feed_subtitle":"Validated on 864 real-world scans from older patients, it could make routine atrophy measurement practical for dementia research.","key_machinery":"The central object is the GCA score, a visual rating that sums 0-3 severity grades across 13 brain regions (sulcal widening in the frontal, temporal, and parieto-occipital lobes of each hemisphere, plus dilatation of the frontal, occipital, and temporal horns and the third ventricle) into a total out of 39. The tool is a deep-learning regressor that maps pre-processed, skull-stripped, registration-normalised 3D CT volumes directly to this score, bypassing tissue segmentation. The authors operationalised the visual scale for CT with explicit written criteria and reference baseline scans, then trained on 518 scans, optimised on 173, and tested on 173 drawn from 864 routine clinical CTs, with r","core_discovery":"On its own terms, the central claim is that a deep-learning model trained on CT images labelled with one expert rater's visual GCA scores can reproduce those scores accurately enough for research and clinical use in older patient cohorts. The validation reports a mean absolute error of 3.2 between tool and rater-1, a mean signed difference of 0.18 with limits of agreement from -7.8 to 8.2, and Cohen's weighted kappa of 0.45 against rater-1 and 0.41 against a second rater, compared with 0.28 between the two raters. About half of the tool's predictions fell within two GCA points of rater-1. The tool also classified scans into no/mild, moderate, and severe atrophy with 70-73% accuracy against r","pith_inferences":["Because the model was trained on a single rater's scores as ground truth, its agreement with that same rater partly reflects imitation; a sharper test would compare the tool against consensus labels from several experts, using the same data.","The fact that tool-rater kappa (0.45, 0.41) exceeded human-human kappa (0.28) suggests the tool may be more consistent than individual raters; if replicated, it could serve as an objective bridge between raters and across cohorts.","The reported tendency to under-rate very high GCA scores and over-rate very low ones implies the tool is most dependable in the moderate range; enriching training with extreme scans is a testable extension the authors themselves flag.","A natural next validation is longitudinal: whether the tool can detect within-patient change in GCA over time, since clinical monitoring would depend on change rather than a single absolute score; the paper does not report test-retest or follow-up sensitivity."],"forward_implications":["Routine CT scans already in clinical archives could be automatically re-scored at four seconds per scan, making large-scale atrophy phenotyping practical for health-data research.","The numeric GCA output can be plugged into electronic-health-record prediction algorithms alongside clinical variables to flag patients at risk of dementia, delirium, falls, or functional decline.","At point of care, a clinically approved version could add a standardised atrophy score to every CT head report without adding radiologist workload.","Because the tool matches or exceeds human-human agreement, it could act as a consistent reference in studies where multiple expert raters are not available."],"supporting_citations":[{"why":"Defines the 13-region GCA scale and its 0-39 scoring, the target the tool predicts and the ground-truth basis for training.","marker":"[5]"},{"why":"Evaluates cranial cavity extraction tools on non-contrast CT, informing the skull-stripping preprocessing step.","marker":"[12]"},{"why":"Supplies the brain-extraction algorithm used to remove skull from CT images before registration and model input.","marker":"[13]"},{"why":"Provides the linear registration method used to align CT images to a common template.","marker":"[14]"},{"why":"Gives the Landis-Koch benchmarks used to interpret weighted kappa values as fair or moderate agreement.","marker":"[17]"},{"why":"Represents the volumetric CT measurement approach the paper contrasts with direct GCA score prediction.","marker":"[18]"},{"why":"Shows CT-based deep learning volumes track biomarkers of neurodegeneration, a comparison point for the tool's age and cognition correlations.","marker":"[19]"},{"why":"Previous CT-based GCA tool trained from paired MRI segmentation; the paper's direct CT training is positioned against it.","marker":"[20]"},{"why":"Describes the cognitive screen used to assess cognition in the stroke cohort.","marker":"[10]"},{"why":"Describes the abbreviated mental test used to measure cognitive impairment in the acute medicine cohort.","marker":"[16]"}],"fun_headline_variants":["AI CT tool scores brain atrophy on par with human raters","Automated CT tool matches human raters on brain atrophy","Deep learning CT tool validated for brain atrophy in elderly","4-second AI CT tool rates brain atrophy like expert raters","CT brain AI: atrophy scoring equals human raters, validated"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The load-bearing premise is that a single rater's visual GCA scores are a valid, reliable ground truth for brain atrophy; if that rater is idiosyncratic, the tool's low error against that rater does not prove it measures atrophy accurately.","fun_headline_variants_meta":{"raw":{"variants":["AI CT tool scores brain atrophy on par with human raters","Automated CT tool matches human raters on brain atrophy","Deep learning CT tool validated for brain atrophy in elderly","4-second AI CT tool rates brain atrophy like expert raters","CT brain AI: atrophy scoring equals human raters, validated"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000373,"raw_usage":{"total_tokens":1980,"prompt_tokens":1046,"completion_tokens":934,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":790,"completion_tokens_details":{"reasoning_tokens":851}},"tokens_in":790,"tokens_out":934,"duration_ms":9585,"temperature":1.0,"reasoning_tokens":851,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T22:40:11.319329+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have three or more independent experts rate a common set of, say, 100 CT scans and compare the tool's scores against their consensus; if the tool's weighted kappa against the consensus is no better than the inter-expert kappa, the claim of human-level measurement fails. A cheaper check: retest on scans with GCA below 3 or above 28, where the paper already reports systematic over- and under-estimation.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Describes the abbreviated mental test used to measure cognitive impairment in the acute medicine cohort."}],"review_version":1}