{"id":"60f47b79-27b2-44fb-8156-fa7ef50102fa","arxiv_id":"1908.02333","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Sex-specific random forest and CNN models trained on T1-weighted MRI volumes and region sizes failed to predict fluid intelligence significantly better than predicting the sample mean.","lead":"This paper tested whether brain scans alone can predict fluid intelligence in 9 to 10 year olds, using sex-specific machine learning models. The models barely beat guessing the average score, suggesting T1-weighted MRI structure alone is not enough to predict fluid intelligence.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The abstract's sex-specific benefit claim is untested: no per-sex zero-rule baseline, no pooled-model comparison, and the paper itself traces the female MSE advantage to fewer female outliers in the validation Gf distribution.","rationale":"The paper's core negative finding survives scrutiny: the random forest improves on the constant baseline by only 1.01 MSE with p=0.17 (Section 4.3), the CNNs are significantly worse than baseline, and Fig. 2 shows all models collapsing onto a narrow range of predictions. The claim that \"more information is necessary\" is well supported. The load-bearing weakness is the abstract's sex-specific claim, which is the paper's distinguishing contribution (it appears in the title). The only evidence is the raw per-sex MSE gap in Table 2, which the authors themselves attribute in Section 5 to the female validation subcohort having fewer outlier Gf scores. Because the baseline predictor is a constant, it would produce a per-sex MSE gap whenever per-sex validation Gf variances differ; per-sex baseline MSEs are never reported. Additionally, every model is trained sex-specifically, so \"may perform better when trained separately\" is never tested against a pooled model — the merged \"combined\" MSE is not that comparison. Finally, Section 2 reveals that ABCD residualized the Gf scores against brain volume and other demographics, so the Discussion's strong claim that T1W-identifiable structure does not explain Gf outruns what the residualized target can show. The reader's weakest assumption concerned the atlas registration/parcellation pipeline discarding signal; my concern is different and, I argue, more decisive: the sex-difference contrast is confounded by target distribution and lacks the minimal controls (per-sex baseline, pooled model, confidence intervals). I partially agree with the reader — we both view the sex-specific conclusion as overstated — but the missing per-sex baseline and pooled comparison are the place where the central claim is least secure, and they are checkable with data the authors already hold. This supports conditional acceptance: either remove or substantively soften the abstract's sex-specific sentence, or add the per-sex baseline and pooled-model analysis with bootstrap intervals.","tokens_in":7719,"tokens_out":12917,"duration_ms":120616,"concrete_test":"Recompute external-validation MSE for the zero-rule baseline (the training-set mean Gf) separately within the female and male validation subcohorts, using the same normalized targets as Table 2; and train a single pooled random-forest model on all 3739 training subjects, evaluating its per-sex external MSE. Bootstrap (10,000 resamples) 95% confidence intervals for the female-minus-male MSE gap in (a) the baseline, (b) the reported sex-specific models, and (c) the pooled model. If the baseline gap is approximately 20 points, or if the pooled model reproduces the same per-sex gap as the sex-specific models, then the abstract's sex-specific advantage is an artifact of the target distribution and the sentence should be removed or reworded.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The abstract's headline suggestion — that predicting Gf from volumetric T1W MRI features \"may perform better when trained separately on male and female data\" — is supported by a single raw contrast: external-validation MSE of 60.68 for the female random-forest model versus 80.74 for the male model (Table 2). That contrast is never actually tested.\n\nFirst, no per-sex zero-rule baseline is reported. The baseline (Section 3) predicts the training-set mean for every subject. Since the female validation subcohort has fewer outlier Gf scores — the paper's own Section 5 explanation for the gap — a constant predictor would also yield a lower female MSE. The entire ~20-point gap may be a property of the validation target distribution, not of sex-specific training.\n\nSecond, no pooled (both-sex) model is ever trained. All models are sex-specific (Sections 3.1 and 3.2), and the \"combined\" MSE of 70.83 merges sex-specific predictions. The claim that separate training helps is therefore never compared against a single jointly-trained model; it is an untested inference.\n\nThird, Section 4.3 reports no uncertainty on the sex contrast — no bootstrap confidence intervals, no paired test — although each validation subcohort has only about 200 subjects.\n\nCompounding this, Section 2 states that ABCD regressed brain volume (the best-established T1W-measurable correlate of intelligence), along with site, age, sex, and demographics, out of the Gf scores themselves. The Discussion's stronger statement that \"characteristics of physical brain structures as identifiable on T1W MRI do not explain Gf\" therefore overreaches: the analysis can only speak to residualized Gf after removing the dominant structural predictor. The negative result itself is honestly reported (RF vs baseline p=0.17, Section 4.3); the load-bearing weakness is the untested sex-specific claim in the abstract.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports an attempt to predict fluid intelligence (Gf) in 9-10-year-old children from T1-weighted MRIs and atlas-based volumetric features, using sex-specific random forest and 3D CNN models trained on ABCD Neurocognitive Prediction Challenge 2019 data and evaluated on an external validation set. The random forest models achieve a combined external-validation MSE of 70.83 versus a zero-rule baseline of 71.84 (a difference the authors report as not statistically significant), while CNN models perform worse than baseline. The female random forest model achieves a lower MSE (60.68) than the male model (80.74). The authors conclude that T1W MRI features alone do not provide compelling prediction of Gf and tentatively suggest that sex-specific training may improve performance.","tokens_in":8046,"tokens_out":2544,"duration_ms":28877,"significance":"If the central negative result holds, the paper provides a useful benchmark for the ABCD challenge and evidence that standard T1W volumetric features are weakly predictive of Gf in this age group after the challenge's preprocessing. The use of an external validation set and a zero-rule baseline are strengths, as is the explicit sex-specific modeling. However, the paper's more novel claim—that sex-specific training may improve prediction—is not supported by any statistical test, no pooled (both-sex) model is trained, and the paper's own discussion attributes the female advantage to the validation target distribution. The negative result is also confounded by the ABCD preprocessing that regressed brain volume and demographic variables out of the Gf scores. These issues materially affect the interpretation of the headline findings.","major_comments":[{"comment":"The claim that predictive models 'may perform better when trained separately on male and female data' is untested. No pooled (both-sex) model is ever trained; the 'combined' MSE in Table 2 is just the merged sex-specific predictions. Moreover, no per-sex zero-rule baseline is reported, so the raw female-vs-male MSE contrast (60.68 vs. 80.74) could be entirely explained by differences in the validation Gf distribution—indeed, Section 5 itself attributes the gap to fewer outlier Gf scores in the female validation subcohort. Please add a pooled model and per-sex baseline predictors, and report a formal test of the sex-by-model interaction or bootstrap confidence intervals for the contrast.","section":"Abstract and Section 5"},{"comment":"The Gf scores were normalized by regressing out brain volume, collection site, age, sex, and demographic factors. Because total brain volume is one of the most established T1W-measurable correlates of intelligence, this preprocessing removes a substantial part of the signal that volumetric T1W features could plausibly carry. The conclusion in Section 6 that 'structural MRI' is insufficient is therefore stronger than the evidence supports. If raw (unregressed) Gf scores are available, the analysis should be repeated on them; otherwise, the Discussion and Conclusion should explicitly acknowledge that the null result is partly a property of the target variable preprocessing rather than a pure property of T1W imaging.","section":"Section 2"},{"comment":"The statistical comparison to baseline is reported only as a '2-sided t-test' with p-values, with no description of the test unit (per-subject squared errors? per-subject predictions?), test statistic, or subject-level variance. No uncertainty is reported for the MSE values in Table 2, and the female-vs-male MSE difference is not tested at all. Please provide the full test details, confidence intervals for all reported MSEs, and a significance test for the sex contrast, given the small validation subcohorts (~200 subjects per sex).","section":"Section 4.3"},{"comment":"The feature-importance interpretation, such as the claim that hippocampus volumes are important for males but not females, is speculative when the model's overall performance is not significantly better than the zero-rule baseline and the predictions cluster near the mean (Fig. 2). Presenting these differences as evidence of sex-specific mechanisms goes beyond what the model performance supports; please label this as hypothesis-generating or remove the causal-sounding interpretation.","section":"Section 4.2 and Section 5"}],"minor_comments":[{"comment":"There are minor typographical and terminology issues: 'crystalized intelligence' should be 'crystallized intelligence'; 'NifTI' should be 'NIfTI'; 'cerebral fluid' should likely be 'cerebrospinal fluid'; and 'segmented in to' should be 'segmented into'.","section":"Throughout"},{"comment":"The table could be made more readable by clarifying the row and column structure; in particular, the 'Sex Specific' and 'Combined Population' column headers are not clearly separated from the MSE values, and the baseline row's dash for sex is ambiguous.","section":"Table 2"},{"comment":"The figures are not described in enough detail in the text; please add axis labels, units, and a brief description of what is shown so that readers can interpret the distributions without guessing.","section":"Figures 1 and 2"},{"comment":"The sentence 'levels of N-acetylaspartate and brain volume from MRI spectroscopy were shown to be associated with aspects of Gf, but not Gf itself' is confusing because brain volume is not measured by spectroscopy; please rephrase for accuracy.","section":"Section 5"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a challenge report, so the bar for novelty is appropriately lower, but the central sex-specific claim in the abstract is not supported by the analyses as reported. The negative result is also confounded by the ABCD Gf preprocessing. These issues are fixable within the scope of the paper by adding (or explicitly ruling out) a pooled model, per-sex baselines, and uncertainty quantification, and by rewriting the abstract and conclusion to match the evidence. I would not recommend rejection because the null result, appropriately hedged, is of value to the challenge community."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth a look, but the abstract overstates one thing. The work is a challenge paper from ABCD 2019, predicting fluid intelligence (Gf) from T1W MRIs in 9–10 year olds. The main empirical content is honest: all models hover near the zero-rule baseline, and the random forest MSE of 70.83 is not significantly better than baseline (p=0.17). That negative result is new and useful, since it suggests structural T1W alone is weak for this target. The paper also ships a clear comparison between CNNs and random forests, and the feature-importance table gives a concrete look at which volumetric regions matter, with hippocampus prominent in males only. Credit where due: the writing is straightforward, the methods are reproducible enough (standard ABCD preprocessing, scikit-learn, Keras), and the discussion explicitly acknowledges that the female advantage is likely driven by fewer female outliers. That last point shows the authors understand their own data.\n\nThe soft spots are real but addressable. The abstract's claim that sex-specific models \"may perform better\" is not tested. There is no per-sex zero-rule baseline, no pooled both-sex model for comparison, and no uncertainty on the sex contrast. The paper's own explanation—fewer female outliers in the validation distribution—would also lower a constant predictor's MSE for females, so the entire 20-point gap could be a property of the test set, not the models. The stress-test note lands. Also, the paper claims at the start that Gf \"has not been previously linked to T1W imaging,\" but then cites structural MRI studies of Gf (e.g., gray matter correlates, hippocampal volume). That is an overclaim, and the reader is right to flag it.\n\nA more subtle issue the stress-test note brings up: ABCD residualized the Gf scores against brain volume and demographics before the challenge. So the analysis can only speak to residualized Gf, not Gf itself. The Discussion's statement that \"characteristics of physical brain structures as identifiable on T1W MRI do not explain Gf\" overreaches, because the dominant structural predictor (total brain volume) had been regressed out. That's a genuine limitation worth stating plainly.\n\nOverall, the negative result holds up. The soft spots are fixable by rephrasing the abstract and adding a couple of baseline comparisons. This is not a breakthrough, but it is a solid negative result from a large dataset, and it deserves a serious referee. I'd send it to peer review if the authors tighten the claims and add the missing baseline. I probably wouldn't cite it in my own work, but it's the kind of paper that should be in the record so others don't chase T1W-only Gf prediction.\n\nVerdict: conditional, leaning accept after revisions.","headline":"A small, honestly reported negative result worth a look, but the abstract overstates an untested sex-specific benefit.","tokens_in":8662,"tokens_out":704,"would_cite":false,"duration_ms":9504,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that fluid intelligence cannot be predicted from T1-weighted structural MRI alone, and that sex-specific training shifts which brain regions matter but does not beat a mean-based baseline.","keywords":["fluid intelligence","T1-weighted MRI","sex differences","random forest","convolutional neural network","adolescent brain","neurocognitive prediction"],"falsifier":"Train the same sex-specific random forest and CNN on T1 images that have not been affinely aligned to an adult atlas, using raw intensities, cortical thickness, or local shape features, and evaluate on the same 415-subject external validation set; any model with MSE significantly below the baseline of 71.84 would overturn the claim that T1W MRI alone is insufficient.","tokens_in":7548,"feed_emoji":"🧠","tokens_out":6472,"duration_ms":67832,"temperature":0.7,"pith_summary":"Fluid intelligence (Gf) is the ability to reason about novel problems, and this paper asks whether it can be read off a standard anatomical T1-weighted brain scan in 9- to 10-year-olds. Using 3,739 training subjects from a large multi-site challenge cohort and 415 for external validation, the authors trained sex-specific 3D convolutional networks on whole-brain images and random forest regressors on volumes of 122 atlas-defined brain regions. The best model, a random forest, achieved an external-validation mean squared error of 70.83, essentially tied with the 71.84 of a baseline that always predicts the training-set mean. The authors conclude that volumetric T1-weighted features alone do not carry enough information to predict fluid intelligence accurately, while noting that models trained separately by sex shifted which regions mattered most and that female predictions were more accurate than male predictions (MSE 60.68 vs 80.74).","feed_headline":"T1 brain scans alone can't predict fluid intelligence","feed_subtitle":"Sex-specific models for 3,739 children barely beat guessing the average on 415 held-out scans.","key_machinery":"The central machinery is the comparison of sex-specific predictive models against a zero-rule baseline using mean squared error on a held-out cohort. The inputs are T1-weighted images preprocessed through skull-stripping, noise removal, field-inhomogeneity correction, affine alignment to a standard adult atlas, and parcellation into 122 regions; a 3D residual CNN consumes the voxels, while a random forest consumes the 122 regional volumes plus age. The baseline predictor, which always outputs the mean Gf of the training set, defines the threshold: any model that cannot beat it is judged to have found no predictive signal in T1W anatomy. Sex-specific training is the second mechanism, used to see whether male and female brains have different structural correlates of Gf.","core_discovery":"The paper's central claim is negative: anatomical information in T1-weighted MRI, whether represented as raw voxels or as atlas-region volumes, does not predict fluid intelligence in children beyond a trivial mean predictor. The authors show this by comparing sex-specific models against a zero-rule baseline: both ResNet variants were significantly worse than baseline, and the random forest's improvement was small (1.01 MSE) and not statistically significant (p = 0.17). They also report a sex difference in prediction accuracy—female MSE 60.68 vs male 80.74—and different top-ranking region volumes for each sex (pons white matter was most important for both; hippocampus volumes appeared only in the male model). Their stated conclusion is that fluid intelligence is not explained by the physical brain structures visible on T1W MRI, and that predicting it will require information beyond these images.","pith_inferences":["A testable extension is to run the same sex-specific random forest on features computed before affine alignment to an adult atlas, such as cortical thickness, surface area, gyrification, or non-parcellated voxel intensities, to determine whether the null result is a preprocessing artifact.","The female/male MSE gap can be disentangled from biology by resampling validation sets to match the score distributions; if the gap disappears, the region-importance differences still suggest sex-specific correlates worth studying.","If functional MRI and diffusion imaging do predict fluid intelligence, as the authors cite, the combined picture implies the predictive information is functional or connectivity-based rather than gross anatomy, so multimodal fusion rather than better T1 features is the promising route."],"forward_implications":["If the central claim holds, T1-weighted structural MRI alone will not rank children by fluid intelligence; in this cohort the best sex-specific model was statistically indistinguishable from predicting everyone's mean score.","The sex-specific result is at most a hint: female models showed lower MSE, but the paper attributes part of this to fewer outlier scores in the female validation set, so sex-separated training should not be assumed to work before re-testing on score-matched cohorts.","Restricting a CNN to caudate and putamen slices, regions previously linked to intelligence, made predictions worse than full-brain input, so localizing to candidate regions does not recover a T1-Gf signal.","Any successful T1-based predictor would need features beyond atlas-region volumes and voxel intensities, or a model substantially more powerful than the ones tested here."],"supporting_citations":[{"why":"Supplies the challenge cohort: 3,739 training and 415 validation subjects with T1W MRIs, regional volumes, age, sex, and normalized Gf scores.","marker":"[16]"},{"why":"Defines the image preprocessing pipeline (skull-stripping, noise removal, field-inhomogeneity correction, affine alignment) whose output the models consume.","marker":"[17]"},{"why":"Provides the standard adult atlas used for alignment and the 122-region parcellation that generates the random forest features.","marker":"[18]"},{"why":"Defines the cognitive test battery that produced the fluid intelligence scores used as prediction targets.","marker":"[4]"},{"why":"Cites the recommendation to consider biological sex in neuroscience, motivating the sex-specific model design.","marker":"[13]"},{"why":"Reports sex differences in structural brain development in children, the premise for expecting sex-specific MRI correlates.","marker":"[14, 15]"},{"why":"Supplies the residual 3D CNN architecture used for the voxel-based prediction experiments.","marker":"[11]"},{"why":"Provides the random forest regression implementation used for the volumetric-feature models.","marker":"[20]"},{"why":"Establishes that fMRI-based measures have been linked to fluid intelligence, the contrast for the claim that T1W lacks such signal.","marker":"[2, 8]"}],"fun_headline_variants":["T1 MRI fails to predict teen fluid intelligence","Brain scans can't gauge fluid intelligence in kids","Sex-specific models don't boost fluid intelligence prediction","No signal for fluid intelligence in T1 brain images","Fluid intelligence eludes prediction from T1 MRI"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The conclusion that T1-weighted MRI does not predict fluid intelligence assumes the preprocessing pipeline (skull-stripping, noise removal, field-inhomogeneity correction, affine alignment to an adult atlas, and 122-region parcellation) preserves whatever structural signal correlates with fluid intelligence.","fun_headline_variants_meta":{"raw":{"variants":["T1 MRI fails to predict teen fluid intelligence","Brain scans can't gauge fluid intelligence in kids","Sex-specific models don't boost fluid intelligence prediction","No signal for fluid intelligence in T1 brain images","Fluid intelligence eludes prediction from T1 MRI"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00034,"raw_usage":{"total_tokens":1904,"prompt_tokens":1000,"completion_tokens":904,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":616,"completion_tokens_details":{"reasoning_tokens":831}},"tokens_in":616,"tokens_out":904,"duration_ms":9186,"temperature":1.0,"reasoning_tokens":831,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T14:46:30.062173+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train the same sex-specific random forest and CNN on T1 images that have not been affinely aligned to an adult atlas, using raw intensities, cortical thickness, or local shape features, and evaluate on the same 415-subject external validation set; any model with MSE significantly below the baseline of 71.84 would overturn the claim that T1W MRI alone is insufficient.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the challenge cohort: 3,739 training and 415 validation subjects with T1W MRIs, regional volumes, age, sex, and normalized Gf scores."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the image preprocessing pipeline (skull-stripping, noise removal, field-inhomogeneity correction, affine alignment) whose output the models consume."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the standard adult atlas used for alignment and the 122-region parcellation that generates the random forest features."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Cites the recommendation to consider biological sex in neuroscience, motivating the sex-specific model design."},{"cited_title":"AMIA Annu","cited_arxiv_id":null,"evidence_quote":"Supplies the residual 3D CNN architecture used for the voxel-based prediction experiments."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the random forest regression implementation used for the volumetric-feature models."}],"review_version":1}