{"id":"1c3a07dd-033c-4cec-ad0a-53a3cfd63e23","arxiv_id":"1908.01555","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Constrained principal component models of functional connectivity predict brain age with mean absolute error around 12 years and transfer across three fMRI datasets, though no baseline comparison is provided.","lead":"This paper tests whether simple, interpretable models of brain connectivity can predict a person's age, comparing PCA with constrained variants. The constrained models give the lowest prediction errors on three fMRI datasets, but the absolute errors are high and no trivial baseline is reported.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Transfer generalization claim is not tested against trivial baselines: on HCP (ages 22–37) MHA's MAE of 12.71 yr is far worse than a constant mean-age predictor, and on ATR (ages 20–70) MHA's 12.18 yr is indistinguishable from the mean-age baseline, so Table 1 and Figures 8–9 do not support the…","rationale":"I agree with the reader's weakest-assumption analysis. The internal CamCAN comparison is a plausible methodological contribution: on CamCAN, MHA/MCF do outperform PCA and factor analysis, and a constant predictor would achieve an MAE of roughly 17.5 years on the 18–88 age range, so the within-dataset result is meaningful. However, the paper's central claim is not only that constrained models work on CamCAN; the abstract and conclusion explicitly claim generalization to HCP and ATR, with figure captions stating 'indicating good generalization.' The reported absolute MAEs are not interpretable without age-distribution baselines, and the paper itself notes the narrow HCP age range and the ATR range yet does not compute such baselines. The proposed concrete test would settle whether the cross-dataset claim lands. If it fails, the paper could still be acceptable as a conditional contribution focused on the CamCAN comparison, provided the generalization wording is softened or removed and the age-distribution issue is addressed. I also note a secondary technical issue: equation (12) as written subtracts a k×k matrix from a p×p matrix and has a dimensional inconsistency; the intended formula is presumably (tr(Σ)−tr(W^TΣW))/(p−k). This does not change the primary verdict but should be fixed. Therefore the appropriate action is CONDITIONAL rather than REJECT or UNCHANGED: the core method and internal results may survive, but the external-generalization claim requires substantive revision.","tokens_in":16403,"tokens_out":6764,"duration_ms":77915,"concrete_test":"Recompute the HCP and ATR rows of Table 1 with two additional columns: (i) MAE of a constant/null model that predicts the CamCAN training mean age for every subject, and (ii) Pearson/Spearman correlation and R^2 between predicted and chronological age. Then run the authors' own paired bootstrap procedure over the same 1000 subsets comparing MHA against the null model and against PCA. If HCP null MAE is below MHA's 12.71, or if ATR null MAE is within the reported standard deviation of MHA's 12.18, the cross-dataset generalization claim is unsupported and the 'generalizes well' wording must be removed or substantially qualified.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"Section 3.2.2 explicitly acknowledges that HCP ages span only 22–37 years and ATR ages span 20–70, yet no trivial baseline is computed. A constant predictor returning the CamCAN training mean would achieve an MAE of roughly 3–4 years on HCP (for a uniform 22–37 age range, mean absolute deviation is about 3.75 years), far below the reported MHA MAE of 12.71. On ATR, a constant predictor would achieve an MAE of roughly 12.5 years, which is within the error bars of the reported MHA value of 12.18. Therefore the absolute MAEs in Table 1 cannot establish that the learned networks encode age-related connectivity; they are also consistent with a model that simply outputs a dataset-specific offset or constant. This concern is load-bearing because the abstract, the introduction, and the figure captions (e.g., 'indicating good generalization') make cross-dataset generalization a central part of the contribution. If the models do not beat trivial baselines on HCP and ATR, the external-generalization support for the claim that constrained latent variable models improve predictive performance collapses, leaving only the internal CamCAN comparison.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a two-stage framework for brain-age prediction from resting-state fMRI functional connectivity. In the first stage, linear latent variable models (factor analysis, PCA, non-negative PCA, MCF, and MHA) are used to learn a shared loading matrix W whose columns are interpreted as functional connectivity networks; subject-specific network activities are then used as features in a linear regression to predict age. The framework is trained and validated on CamCAN, with the number of factors k selected by validation likelihood, and the fitted model is transferred without retraining to HCP and ATR Wide-Age-Range data. The central claim is that introducing non-negativity and orthonormality constraints improves both interpretability and predictive performance, with MHA and MCF reported as the best-performing methods on all three repositories.","tokens_in":16657,"tokens_out":4192,"duration_ms":45888,"significance":"If the central claims hold, the paper would contribute a transparent, interpretable approach to brain-age modeling from functional connectivity, with the inferred networks (e.g., default mode, salience, visual, somatomotor) offering neurobiological insight. The manuscript has several strengths: the method is clearly described; the choice of k using validation likelihood without using age labels is a sensible safeguard against circularity; the CamCAN internal comparison across latent-variable models is a useful empirical contribution; and the synthetic experiments, while favorable to the constrained models, are explicitly presented as a numerical validation of the estimation procedure. The cross-dataset transfer design (training only on CamCAN) is also commendable. However, the external generalization evidence is currently under-analyzed: the raw MAE values on HCP and ATR are not compared with trivial baselines, which is essential because the age ranges of those datasets are narrow or uneven.","major_comments":[{"comment":"The cross-dataset generalization claim is not evaluated against trivial baselines. On HCP, ages span only 22–37 years, so a constant predictor returning the mean age of the HCP sample would achieve an MAE equal to the mean absolute deviation of that age distribution, roughly 2–4 years (approximately 3.75 years even under a uniform assumption), whereas the reported MHA MAE is 12.71 years. On ATR, ages span 20–70 years, where a constant predictor would achieve an MAE of about 12.5 years under a uniform age distribution, essentially equal to the reported MHA value of 12.18 years. The raw MAEs therefore do not demonstrate that the models encode age-related connectivity; they are equally consistent with predictions that are nearly constant or driven by dataset-specific offsets. This undermines the abstract's and Section 4's claims that the models 'generalize well' to unseen repositories and that constraints improve predictive performance on those data.","section":"Section 3.2.2, Table 1"},{"comment":"The conclusion that MHA and MCF show the 'smallest drop' in performance and thereby generalize best is based on absolute MAE differences that are small relative to the age-range baseline. For example, on ATR the difference between MHA (12.18) and factor analysis (13.40) is about 1.2 years, while the constant-predictor baseline is about 12.5 years; on HCP the MAE values exceed the entire age range of the dataset. Without reporting the correlation between predicted and true age (or R²) on HCP and ATR, or explicitly comparing against a mean-age predictor, the relative ordering of methods could reflect calibration or dataset-specific offsets rather than age-related predictive signal. The authors should either add these baselines and calibration analyses to support the transfer claim, or substantially weaken the generalization claim and restrict the predictive-performance conclusion to the internal CamCAN comparison.","section":"Section 3.2.2, Figures 8–9"}],"minor_comments":[{"comment":"The word 'conncetivity' appears in the second paragraph of the introduction and should be corrected to 'connectivity'.","section":"Section 1"},{"comment":"The sentence 'The authors with to thank Steve Smith' should read 'wish to thank'.","section":"Acknowledgments"},{"comment":"The word 'othonormality' in the caption is a typo and should be 'orthonormality'.","section":"Figure 6 caption"},{"comment":"The hyper-parameter k is fixed to 5 for all models even though the authors note that the optimal k may differ across models; a sentence explaining why this fixed choice is appropriate for the cross-model comparison would be helpful.","section":"Section 2.3 and Section 3.2.1"},{"comment":"The caption states that bracketed values are standard deviations, but it would be clearer to specify that these are standard deviations across the 1000 random subsets of 30 subjects, not across the subjects within a dataset.","section":"Table 1 caption"},{"comment":"The sentence 'This is a phenomenon is also observed in the real data analysis' contains a grammatical error ('This is a phenomenon is') and should be rephrased.","section":"Section 3.1.1"}],"recommendation":"major_revision","confidential_remarks":"The manuscript presents an interesting and clearly described framework, and the internal CamCAN comparison is a useful contribution. However, the external validation section currently overclaims: the reported HCP and ATR MAEs are worse than or comparable to trivial constant predictors, so the generalization evidence needs substantial additional analysis (baselines, prediction-age correlations, calibration) before the paper's central claims can be accepted. I would encourage revision rather than rejection, as the methodological core appears sound and the authors can likely address the concern without changing the overall approach."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe paper is worth a referee slot, but not for the reason the authors emphasise. The useful part is the CamCAN comparison: MHA and MCF beat PCA, FA, and non-negative PCA for brain-age prediction from functional connectivity, and the inferred networks are more interpretable. That internal result looks solid enough to matter.\n\nThe synthetic experiments are a reasonable sanity check, but since the data is generated from the MHA model, they only confirm the algorithm recovers its own assumptions. The novelty claim is wrong: the paper says it is the first to interpret PCA as learning functional networks, yet cites Leonardi et al. (2013) doing just that. That should be corrected.\n\nThe transfer story is oversold. The stress-test claims MHA's HCP MAE of 12.71 is worse than a constant mean-age predictor. That's a miscalculation. A constant predictor returning the CamCAN training mean would produce an MAE around 25 years on HCP (ages 22-37), not 3.75. The 3.75 is the mean absolute deviation about the HCP mean, which is not a baseline the model could know. So the model does beat a reasonable null on HCP. On ATR the margin is much thinner: 12.18 versus about 14 for the CamCAN-mean null, and almost identical to the ATR-mean null (12.5) if you had target labels. The authors should report these null baselines explicitly; without them, 'good generalization' in the captions is hard to evaluate.\n\nOther soft spots: k is fixed at 5 for all models, which may disadvantage PCA/FA if their optimal k is different. No significance tests are given for the MAE differences, so we don't know if the ordering is reliable. And the train/test split for the unsupervised stage is ambiguous: is W estimated on the full CamCAN training set or on a subset? The authors should clarify.\n\nThis is a methods-comparison paper, not a breakthrough. With added baselines, significance testing, a corrected novelty claim, and softened transfer language, it deserves acceptance. I'd send it to a serious referee.","headline":"Useful internal comparison of constrained latent variable models for brain age, but the cross-dataset generalization is oversold and the novelty claim is wrong.","tokens_in":17239,"tokens_out":7167,"would_cite":false,"duration_ms":69950,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62H25","62J05","62P10"],"pacs":[],"model":"deepseek-v4-flash","headline":"For connectivity-based brain age prediction, constraining PCA-style latent variable models with non-negativity and orthonormality improves both interpretability and predictive accuracy.","keywords":["brain age","functional connectivity","resting-state fMRI","PCA","non-negative matrix factorization","interpretability","aging","linear latent variable models"],"falsifier":"Rerun the evaluation on HCP and ATR with a trivial baseline that always predicts the mean training age; if the MHA and MCF errors do not clearly beat this baseline, the transfer-generalization claim fails. A stronger test would restrict all three datasets to a common age range and re-compare the methods on that subset.","tokens_in":16187,"feed_emoji":"🧠","tokens_out":7442,"duration_ms":62655,"temperature":0.7,"pith_summary":"The paper argues that brain age can be predicted from resting-state fMRI functional connectivity with a two-step pipeline: first learn a small set of shared connectivity networks using a linear latent variable model, then regress age on each subject's activity within those networks. Its central thesis is that adding non-negativity and orthonormality constraints to the loading matrix yields networks that are easier to interpret and better predictors of age. On three open datasets (CamCAN, HCP, and ATR Wide-Age-Range), the constrained models MCF and MHA achieve the lowest mean absolute error, and the inferred networks recapitulate known ageing-sensitive systems such as the default mode network. If right, this offers a transparent alternative to black-box brain age predictors and turns the predictive model into a tool for probing how connectivity changes with age.","feed_headline":"Constrained network models outperform plain PCA for brain age","feed_subtitle":"Resting-state connectivity predicts age best when networks are forced into non-negative, non-overlapping components.","key_machinery":"The central object is the loading matrix $W \\in \\mathbb{R}^{p \\times k}$, whose columns define functional networks, constrained by non-negativity and orthonormality. MHA (Modular Hierarchical Analysis) is the probabilistic model that enforces both constraints and makes the decomposition identifiable; MCF (Modular Connectivity Factorization) is the related objective with the same constraints. Each subject's covariance is modelled as $\\Sigma^{(i)} = \\sum_j g^{(i)}_j W_j W_j^T + v^{(i)} I$, so the diagonal entries $g^{(i)}_j$ quantify the activity of network $j$ in subject $i$, and those activities become the features in the linear regression for age. The machinery converts high-dimensional connectivity matrices into a small set of non-negative, non-overlapping network scores while retaining predictive signal, which is what makes the subsequent model interpretation possible.","core_discovery":"The paper claims that functional connectivity features for brain age prediction are best extracted by decomposing each subject's covariance matrix as a low-rank sum of networks, $\\Sigma^{(i)} = \\sum_j g^{(i)}_j W_j W_j^T + v^{(i)} I$, with the loading matrix $W$ constrained to be non-negative ($W \\ge 0$) and orthonormal ($W^T W = I$). Under these constraints (the MHA and MCF models), each brain region belongs to exactly one network, the decomposition is identifiable up to trivial symmetries, and the associated network activities $g^{(i)}_j$ are more informative for linear regression than unconstrained PCA or factor analysis loadings. In experiments on 647 CamCAN participants, the constrained models produced spatially coherent, bilaterally symmetric networks matching the default mode, salience, visual, and somatomotor systems, and activity in most of these networks declined with age. Transferring the CamCAN-trained regression to HCP and ATR produced mean absolute errors of 12.71 and 12.18 years respectively, lower than all less-constrained baselines on the same held-out data.","pith_inferences":["The raw mean absolute error comparisons on HCP and ATR are fragile because those datasets have narrow or mid-range age distributions; a fairer evaluation would compare against an age-constant baseline, and the 'generalizes well' claim may weaken under that test.","The same constrained decomposition could be used to compute a per-network 'brain age gap', separating which network's decline most drives an individual's predicted age acceleration.","The identifiability of the constrained loading matrix could enable direct cross-cohort comparisons of network structure, turning brain age models into standardized connectomic atlases.","A testable extension is to apply the pipeline to task-based fMRI or to patient groups to see whether the same networks carry the age signal when cognition or pathology changes."],"forward_implications":["Connectivity-based brain age models can be made interpretable without sacrificing accuracy, provided the feature extraction step imposes non-negativity and orthonormality.","The inferred networks identify a concrete set of age-sensitive systems (default mode, salience, visual, and somatomotor) whose activity declines with age, giving targeted hypotheses for ageing research.","The two-step pipeline transfers across scanners and acquisition protocols, suggesting the learned networks capture reproducible between-subject variation rather than scanner-specific artefacts.","Because the regression is linear, each network's coefficient directly reads as the change in predicted age per unit change in network activity, enabling hypothesis-driven follow-up.","The performance figures set a baseline for healthy subjects only; extending the same pipeline to clinical groups would require re-validation before use as a biomarker."],"supporting_citations":[{"why":"Introduces MHA, the probabilistic model whose non-negativity and orthonormality constraints make the loading matrix identifiable and networks non-overlapping.","marker":"(Monti and Hyvärinen, 2018)"},{"why":"Introduces MCF, the alternative constrained objective that matches MHA's constraints and achieves the best generalization performance.","marker":"(Hirayama et al., 2016)"},{"why":"Provides the non-negative PCA baseline that relaxes orthonormality, used to isolate the contribution of the orthogonality constraint.","marker":"(Sigg and Buhmann, 2008)"},{"why":"Supplies the interpretation of PCA loadings as functional connectivity networks, which the two-step pipeline builds on.","marker":"(Leonardi et al., 2013)"},{"why":"Defines the 264-ROI parcellation whose time courses are the input to the connectivity analysis.","marker":"(Power et al., 2011)"},{"why":"Prior evidence that default mode network connectivity declines with age, used to validate the age trends of the inferred networks.","marker":"(Geerligs et al., 2015)"}],"fun_headline_variants":["Non-negative orthonormal networks beat PCA for brain age","Constrained latent networks improve interpretable brain age","Brain age prediction sharpened by network orthonormality","Identifiable functional networks excel at predicting age","Non-overlapping network components refine brain age estimates"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The transfer evaluation assumes raw mean absolute error is a meaningful measure of generalization on HCP and ATR, but on HCP (ages 22 to 37) a constant predictor would already give an error far below the reported 12.71 years, so the 'generalizes well' conclusion rests on a fragile premise.","fun_headline_variants_meta":{"raw":{"variants":["Non-negative orthonormal networks beat PCA for brain age","Constrained latent networks improve interpretable brain age","Brain age prediction sharpened by network orthonormality","Identifiable functional networks excel at predicting age","Non-overlapping network components refine brain age estimates"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000344,"raw_usage":{"total_tokens":1888,"prompt_tokens":945,"completion_tokens":943,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":561,"completion_tokens_details":{"reasoning_tokens":868}},"tokens_in":561,"tokens_out":943,"duration_ms":9800,"temperature":1.0,"reasoning_tokens":868,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T15:09:25.874633+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Rerun the evaluation on HCP and ATR with a trivial baseline that always predicts the mean training age; if the MHA and MCF errors do not clearly beat this baseline, the transfer-generalization claim fails. A stronger test would restrict all three datasets to a common age range and re-compare the methods on that subset.","supporting_citations":[],"review_version":1}