REVIEW 4 major objections 4 minor 26 references
Sex differences in predicting fluid intelligence of adolescent brain from T1-weighted MRIs
T0 review · 4 major / 4 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read This paper claims that fluid intelligence cannot be predicted from T1-weighted structural MRI alone, and that sex-specific training shifts which brain regions matter but does not beat a mean-based baseline.
desk verdict A small, honestly reported negative result worth a look, but the abstract overstates an untested sex-specific benefit. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central machinery is the comparison of sex-specific predictive models against a zero-rule baseline using mean squared error on a held-out cohort. The inputs are T1-weighted images preprocessed through skull-stripping, noise removal, field-inhomogeneity correction, affine alignment to a standard adult atlas, and parcellation into 122 regions; a 3D residual CNN consumes the voxels, while a random forest consumes the 122 regional volumes plus age. The baseline predictor, which always outputs the mean Gf of the training set, defines the threshold: any model that cannot beat it is judged to have found no predictive signal in T1W anatomy. Sex-specific training is the second mechanism, used to see whether male and female brains have different structural correlates of Gf.
What would settle it
Train the same sex-specific random forest and CNN on T1 images that have not been affinely aligned to an adult atlas, using raw intensities, cortical thickness, or local shape features, and evaluate on the same 415-subject external validation set; any model with MSE significantly below the baseline of 71.84 would overturn the claim that T1W MRI alone is insufficient.
Extended reading notes
Core claim
The paper's central claim is negative: anatomical information in T1-weighted MRI, whether represented as raw voxels or as atlas-region volumes, does not predict fluid intelligence in children beyond a trivial mean predictor. The authors show this by comparing sex-specific models against a zero-rule baseline: both ResNet variants were significantly worse than baseline, and the random forest's improvement was small (1.01 MSE) and not statistically significant (p = 0.17). They also report a sex difference in prediction accuracy—female MSE 60.68 vs male 80.74—and different top-ranking region volumes for each sex (pons white matter was most important for both; hippocampus volumes appeared only in the male model). Their stated conclusion is that fluid intelligence is not explained by the physical brain structures visible on T1W MRI, and that predicting it will require information beyond these images.
Load-bearing premise
The conclusion that T1-weighted MRI does not predict fluid intelligence assumes the preprocessing pipeline (skull-stripping, noise removal, field-inhomogeneity correction, affine alignment to an adult atlas, and 122-region parcellation) preserves whatever structural signal correlates with fluid intelligence.
Editorial extensions
If this is right
- If the central claim holds, T1-weighted structural MRI alone will not rank children by fluid intelligence; in this cohort the best sex-specific model was statistically indistinguishable from predicting everyone's mean score.
- The sex-specific result is at most a hint: female models showed lower MSE, but the paper attributes part of this to fewer outlier scores in the female validation set, so sex-separated training should not be assumed to work before re-testing on score-matched cohorts.
- Restricting a CNN to caudate and putamen slices, regions previously linked to intelligence, made predictions worse than full-brain input, so localizing to candidate regions does not recover a T1-Gf signal.
- Any successful T1-based predictor would need features beyond atlas-region volumes and voxel intensities, or a model substantially more powerful than the ones tested here.
Reading between the lines
- A testable extension is to run the same sex-specific random forest on features computed before affine alignment to an adult atlas, such as cortical thickness, surface area, gyrification, or non-parcellated voxel intensities, to determine whether the null result is a preprocessing artifact.
- The female/male MSE gap can be disentangled from biology by resampling validation sets to match the score distributions; if the gap disappears, the region-importance differences still suggest sex-specific correlates worth studying.
- If functional MRI and diffusion imaging do predict fluid intelligence, as the authors cite, the combined picture implies the predictive information is functional or connectivity-based rather than gross anatomy, so multimodal fusion rather than better T1 features is the promising route.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports an attempt to predict fluid intelligence (Gf) in 9-10-year-old children from T1-weighted MRIs and atlas-based volumetric features, using sex-specific random forest and 3D CNN models trained on ABCD Neurocognitive Prediction Challenge 2019 data and evaluated on an external validation set. The random forest models achieve a combined external-validation MSE of 70.83 versus a zero-rule baseline of 71.84 (a difference the authors report as not statistically significant), while CNN models perform worse than baseline. The female random forest model achieves a lower MSE (60.68) than the male model (80.74). The authors conclude that T1W MRI features alone do not provide compelling prediction of Gf and tentatively suggest that sex-specific training may improve performance.
Significance. If the central negative result holds, the paper provides a useful benchmark for the ABCD challenge and evidence that standard T1W volumetric features are weakly predictive of Gf in this age group after the challenge's preprocessing. The use of an external validation set and a zero-rule baseline are strengths, as is the explicit sex-specific modeling. However, the paper's more novel claim—that sex-specific training may improve prediction—is not supported by any statistical test, no pooled (both-sex) model is trained, and the paper's own discussion attributes the female advantage to the validation target distribution. The negative result is also confounded by the ABCD preprocessing that regressed brain volume and demographic variables out of the Gf scores. These issues materially affect the interpretation of the headline findings.
major comments (4)
- [Abstract and Section 5] The claim that predictive models 'may perform better when trained separately on male and female data' is untested. No pooled (both-sex) model is ever trained; the 'combined' MSE in Table 2 is just the merged sex-specific predictions. Moreover, no per-sex zero-rule baseline is reported, so the raw female-vs-male MSE contrast (60.68 vs. 80.74) could be entirely explained by differences in the validation Gf distribution—indeed, Section 5 itself attributes the gap to fewer outlier Gf scores in the female validation subcohort. Please add a pooled model and per-sex baseline predictors, and report a formal test of the sex-by-model interaction or bootstrap confidence intervals for the contrast.
- [Section 2] The Gf scores were normalized by regressing out brain volume, collection site, age, sex, and demographic factors. Because total brain volume is one of the most established T1W-measurable correlates of intelligence, this preprocessing removes a substantial part of the signal that volumetric T1W features could plausibly carry. The conclusion in Section 6 that 'structural MRI' is insufficient is therefore stronger than the evidence supports. If raw (unregressed) Gf scores are available, the analysis should be repeated on them; otherwise, the Discussion and Conclusion should explicitly acknowledge that the null result is partly a property of the target variable preprocessing rather than a pure property of T1W imaging.
- [Section 4.3] The statistical comparison to baseline is reported only as a '2-sided t-test' with p-values, with no description of the test unit (per-subject squared errors? per-subject predictions?), test statistic, or subject-level variance. No uncertainty is reported for the MSE values in Table 2, and the female-vs-male MSE difference is not tested at all. Please provide the full test details, confidence intervals for all reported MSEs, and a significance test for the sex contrast, given the small validation subcohorts (~200 subjects per sex).
- [Section 4.2 and Section 5] The feature-importance interpretation, such as the claim that hippocampus volumes are important for males but not females, is speculative when the model's overall performance is not significantly better than the zero-rule baseline and the predictions cluster near the mean (Fig. 2). Presenting these differences as evidence of sex-specific mechanisms goes beyond what the model performance supports; please label this as hypothesis-generating or remove the causal-sounding interpretation.
minor comments (4)
- [Throughout] There are minor typographical and terminology issues: 'crystalized intelligence' should be 'crystallized intelligence'; 'NifTI' should be 'NIfTI'; 'cerebral fluid' should likely be 'cerebrospinal fluid'; and 'segmented in to' should be 'segmented into'.
- [Table 2] The table could be made more readable by clarifying the row and column structure; in particular, the 'Sex Specific' and 'Combined Population' column headers are not clearly separated from the MSE values, and the baseline row's dash for sex is ambiguous.
- [Figures 1 and 2] The figures are not described in enough detail in the text; please add axis labels, units, and a brief description of what is shown so that readers can interpret the distributions without guessing.
- [Section 5] The sentence 'levels of N-acetylaspartate and brain volume from MRI spectroscopy were shown to be associated with aspects of Gf, but not Gf itself' is confusing because brain volume is not measured by spectroscopy; please rephrase for accuracy.
Circularity Check
No circularity: models are trained on imaging and demographic features and evaluated on a held-out external validation set; the target Gf scores are not derived from the features used for prediction.
full rationale
The paper's derivation chain is self-contained with respect to the data: convolutional and random forest models are trained on T1W MRI volumes and atlas-region volumes plus age, and predictions are evaluated on an external validation cohort that was not used for training or hyperparameter selection. Hyperparameter tuning was performed on an internal 80:20 split and early stopping used a validation-loss plateau, which does not make the external validation circular. The Gf scores were residualized by the ABCD consortium against demographic and brain-volume covariates, but the authors do not fit or reuse that residualization model as a predictor; the residualized scores are simply the target values. No feature is defined in terms of the target, and no equation in the paper reduces a predicted quantity to a fitted input. The sex-specific MSE comparison (60.68 vs. 80.74) is statistically fragile because no per-sex baseline, pooled model, or uncertainty interval is reported, and the authors themselves attribute the gap to fewer female outlier scores; however, this is a statistical-evidence concern, not circularity. The paper contains no load-bearing self-citation chain, no imported uniqueness theorem, and no ansatz smuggled in via citation. The central null conclusion — that T1W features do not compellingly predict Gf — follows from direct comparison against a zero-rule baseline on held-out data and is not equivalent to its inputs by construction.
Assumptions & free parameters
free parameters (5)
- Random forest max_depth =
Not reported (grid searched over 2-6)
- Random forest n_estimators =
Not reported (grid searched over 100-500, 750, 1000)
- CNN learning rate schedule =
Initial 0.001, reduced by factor 0.1 at validation loss plateau
- CNN batch size, epochs, stride, image resizing =
Batch 32, 100 epochs, stride 2 in initial conv, resized to 128x128x128
- Caudate-putamen slice range =
Slices 55-75
assumptions (3)
- domain assumption The ABCD-provided normalized Gf scores are a valid continuous target for regression.
- domain assumption Affine alignment to the SRI24 adult atlas and atlas-based parcellation preserve structural information relevant to Gf in 9-10 year old brains.
- domain assumption The training and validation splits are representative and the sex ratio imbalance (1.11 vs 1.02) does not materially bias comparisons.
Cite this review
Pith. "Pith review of Sex differences in predicting fluid intelligence of adolescent brain from T1-weighted MRIs." pith.science (2026). https://pith.science/paper/2QXKYHLN
@misc{pith2026190802333,
author = {Pith},
title = {Pith review of: Sex differences in predicting fluid intelligence of adolescent brain from T1-weighted MRIs},
year = {2026},
howpublished = {\url{https://pith.science/paper/2QXKYHLN}},
note = {Machine review of arXiv:1908.02333}
}
read the original abstract
Fluid intelligence (Gf) has been defined as the ability to reason and solve previously unseen problems. Links to Gf have been found in magnetic resonance imaging (MRI) sequences such as functional MRI and diffusion tensor imaging. As part of the Adolescent Brain Cognitive Development Neurocognitive Prediction Challenge 2019, we sought to predict Gf in children aged 9-10 from T1-weighted (T1W) MRIs. The data included atlas-aligned volumetric T1W images, atlas-defined segmented regions, age, and sex for 3739 subjects used for training and internal validation and 415 subjects used for external validation. We trained sex-specific convolutional neural net (CNN) and random forest models to predict Gf. For the convolutional model, skull-stripped volumetric T1W images aligned to the SRI24 brain atlas were used for training. Volumes of segmented atlas regions along with each subject's age were used to train the random forest regressor models. Performance was measured using the mean squared error (MSE) of the predictions. Random forest models achieved lower MSEs than CNNs. Further, the external validation data had a better MSE for females than males (60.68 vs. 80.74), with a combined MSE of 70.83. Our results suggest that predictive models of Gf from volumetric T1W MRI features alone may perform better when trained separately on male and female data. However, the performance of our models indicates that more information is necessary beyond the available data to make accurate predictions of Gf.
Reference graph
Works this paper leans on
-
[1]
Merrifield, P.R., Cattell, R.B.: Abilities: Their Structure, Growth, and Action, http://dx.doi.org/10.2307/1162752, (1975)
-
[2]
Colom, R., Haier, R.J., Head, K., Álvarez-Linera, J., Quiroga, M.Á., Shih, P.C., Jung, R.E.: Gray matter correlates of fluid, crystallized, and spatial intelligence: Testing the P -FIT model, http://dx.doi.org/10.1016/j.intell.2008.07.007, (2009)
-
[3]
Chamorro-Premuzic, T., Furnham, A.: Personality, intelligence and approaches to learning as predictors of academic performance, http://dx.doi.org/10.1016/j.paid.2008.01.003, (2008)
-
[4]
Akshoomoff, N., Beaumont, J.L., Bauer, P.J., Dikmen, S.S., Gershon, R.C., Mungas, D., Slotkin, J., Tulsky, D., Weintraub, S., Zelazo, P.D., Heaton, R.K.: VIII. NIH TOOLBOX COGNITION BATTERY (CB): COMPOSITE SCORES OF CRYSTALLIZED, FLUID, AND OVERALL COGNITION, http://dx.doi.org/10.1111/mono.12038, (2013)
-
[5]
Horn, J.L., Cattell, R.B.: Age differences in fluid and crystallized intelligence, http://dx.doi.org/10.1016/0001-6918(67)90011-x, (1967)
-
[6]
Kievit, R.A., Davis, S.W., Griffiths, J., Correia, M.M., Cam -Can, Henson, R.N.: A watershed model of individual differences in fluid intelligence. Neuropsychologia. 91, 186– 198 (2016)
work page 2016
-
[7]
Fry, A.F., Hale , S.: Relationships among processing speed, working memory, and fluid intelligence in children. Biol. Psychol. 54, 1–34 (2000)
work page 2000
-
[8]
Finn, E.S., Shen, X., Scheinost, D., Rosenberg, M.D., Huang, J., Chun, M.M., Papademetris, X., Constable, R.T.: Functional connectome fingerprinting: identifying individuals using patterns of brain connectivity. Nat. Neurosci. 18, 1664–1671 (2015)
work page 2015
Show all 26 references
-
[9]
Neuroimage
Paul, E.J., Larsen, R.J., Nikolaidis, A., Ward, N., Hillman, C.H., Cohen, N.J., Kramer, A.F., Barbey, A.K.: Dissociable brain biomarkers of fluid intelligence. Neuroimage. 137, 201 – 211 (2016)
2016
-
[10]
Cole, J.H., Poudel, R.P.K., Tsagkrasoulis, D., C aan, M.W.A., Steves, C., Spector, T.D., Montana, G.: Predicting brain age with deep learning from raw imaging data results in a reliable and heritable biomarker, http://dx.doi.org/10.1016/j.neuroimage.2017.07.059, (2017)
2017 doi
-
[11]
AMIA Annu
Yang, C., Rangarajan, A., Ranka, S.: Visual Explanations From Deep 3D Convolutional Neural Networks for Alzheimer’s Disease Classification. AMIA Annu. Symp. Proc. 2018, 1571–1580 (2018) 9
2018
-
[12]
IEEE J Biomed Health Inform
Shi, J., Zheng, X., Li, Y., Zhang, Q., Ying, S.: Multimodal Neuroimaging Feature Learning With Multimodal Stacked Deep Polynomial Networks for Diagnosis of Alzheimer’s Disease. IEEE J Biomed Health Inform. 22, 173–183 (2018)
2018
-
[13]
Shansky, R.M., Woolley, C.S.: Considering Sex as a Biological Variable Will Be Valuable for Neuroscience Research. J. Neurosci. 36, 11817–11822 (2016)
2016
-
[14]
Sowell, E.R., Trauner, D.A., Gamst, A., Jernigan, T.L.: Development of cortical and subcortical brain structures in childhood and adolescence: a structural MRI study. Dev. Med. Child Neurol. 44, 4–16 (2002)
2002
-
[15]
67, 728–734 (2010)
Giedd, J.N., Rapoport, J.L.: Structural MRI of pediatric brain development: what have we learned and where are we going? Neuron. 67, 728–734 (2010)
2010
-
[16]
Pfefferbaum, A., Kwon, D., Brumback, T., Thompson, W.K., Cummins, K., Tapert, S.F., Brown, S.A., Colrain, I.M., Baker, F.C., Prouty, D., De Bellis, M.D., Clark, D.B., Na gel, B.J., Chu, W., Park, S.H., Pohl, K.M., Sullivan, E.V.: Altered Brain Developmental Trajectories in Ado...
2018
-
[17]
Hagler, D.J., Hatton, S.N., Makowski, C., Daniela Cornejo, M., Fair, D.A., Dick, A.S., Sutherland, M.T., Casey, B.J., Barch, D.M., Harms, M.P., Watts, R., Bjork, J.M., Garavan, H.P., Hilmer, L., Pung, C.J., Sicat, C.S., Kuperman, J., Bartsch, H., Xue, F., Heitzeg, M.M., Laird,...
2018
-
[18]
Rohlfing, T., Zahr, N.M., Sullivan, E.V., Pfefferbaum, A.: The SRI24 multichannel atlas of normal adult human brain structure. Hum. Brain Mapp. 31, 798–819 (2010)
2010
-
[19]
Burgaleta, M., MacDonald, P.A., Martínez, K., Román, F.J., Álvarez -Linera, J., Ramos González, A., Karama, S., Colom, R.: Subcortical regional morphology correlates with fluid and spatial intelligence. Hum. Brain Mapp. 35, 1957–1968 (2014)
2014
-
[20]
Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., Others: Scikit -learn: Machine Learning in Python Journal of Machine Learning Research. (2011)
2011
-
[21]
McGraw -Hill Science, Engineering & Mathematics (2007) 10
Saladin, K.S.: Anatomy & Physiology: The Unity of Form and Function. McGraw -Hill Science, Engineering & Mathematics (2007) 10
2007
-
[22]
VanElzakker, M., Fevurly, R.D., Breindel, T., Spencer, R.L.: Environmental novelty is associated with a selective increase in Fos expression in the output elements of the hippocampal formation and the perirhinal cortex. Learn. Mem. 15, 899–908 (2008)
2008
-
[23]
PLoS One
Duarte, I.C., Ferreira, C., Marques, J., Castelo -Branco, M.: Anterior/posterior competitive deactivation/activation dichotomy in the human hippocampus as revealed by a 3D navigation task. PLoS One. 9, e86213 (2014)
2014
-
[24]
Maguire, E.A., Gadian, D.G., Johnsrude, I.S., Good, C.D., Ashburner, J., Frackowiak, R.S., Frith, C.D.: Navigation -related structural change in the hippocampi of taxi drivers. Proc. Natl. Acad. Sci. U. S. A. 97, 4398–4403 (2000)
2000
-
[25]
Houghton Mifflin Harcourt (HMH) (1971)
Cattell, R.B.: Abilities: Their Structure, Growth, and Action. Houghton Mifflin Harcourt (HMH) (1971)
1971
-
[26]
Raven, J.C., Court, J.H.: Manual for Raven’s Progressive Matrices and Vocabulary Scales: Advanced progressive matrices. (1998)
1998
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.