{"id":"f15b8d85-cb09-4771-8f55-85baf5489de5","arxiv_id":"2501.01510","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"VNN-based brain age gap analysis finds distinct anatomical patterns in Alzheimer's, frontotemporal dementia, and atypical Parkinsonian disorders, but not in Parkinson's disease.","lead":"The authors apply covariance neural networks, pretrained on healthy brains, to estimate brain age gap from cortical thickness in Alzheimer's, frontotemporal dementia, atypical Parkinsonian, and Parkinson's cohorts. They report elevated brain age gaps for the first three groups and distinct anatomical patterns tied to how the model uses covariance eigenvectors.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Covariance re-estimation per dataset changes the VNN's trained operator; resulting Δ-Age patterns may reflect covariance shift rather than neurodegeneration.","rationale":"After reading the manuscript, the central empirical claim is that Δ-Age is significantly elevated in AD, FTD, and APD, with disease-specific anatomical patterns and explainability via eigenvector inner products. The most load-bearing step is the inference-time substitution of the anatomical covariance matrix: the VNN's filters are polynomial functions of C, and the trained taps are only meaningful for the covariance used during training. Section 4's use of per-dataset HC covariance matrices (with different atlases/preprocessing, and for 4RTNI an HC group borrowed from NIFD) is not justified by any stability bound or sensitivity experiment. If this substitution is invalid, all reported Δ-Age values and eigenvector analyses are potentially artifacts. This concern is more fundamental than the absence of significance tests or the selection of one of ten models, because it undermines the validity of the model itself on every dataset. I agree with the reader's weakest_assumption. I do not change the verdict; the paper remains acceptable conditionally, pending the proposed covariance-robustness check.","tokens_in":9100,"tokens_out":4760,"duration_ms":48291,"concrete_test":"Run the full pipeline twice for ADNI, NIFD, and 4RTNI: once with the OASIS-3 training covariance matrix and once with the dataset-specific HC covariance matrix, keeping filter taps, bias-adjustment, and all analysis steps identical. Compare the Δ-Age distributions and the eigenvector inner-product group differences (Figs. 3 and 5). If the elevated Δ-Age or the reported eigenvector-specific effects (e.g., AD: eigenvectors 0,1,2,6; FTD: 0,1,4,5; APD: 8) are not reproduced with the training covariance matrix, the central claim fails. Also compute the relative perturbation ‖C_HC − C_train‖ / ‖C_train‖ and compare with the stability bounds in [15]; if the perturbation exceeds the bound, the inference is outside the theoretical guarantee.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper trains a two-layer VNN on OASIS-3 healthy controls with the anatomical covariance matrix C_train estimated from the training set (Section 3.2). At inference, Section 4 states: 'For each disease dataset, we used the anatomical covariance matrix estimated only from the respective HC group in the pre-trained VNN model.' Because the coVariance filter H(C) = Σ h_k C^k depends explicitly on C, replacing C_train with a dataset-specific C (estimated from 114–206 subjects, and with different preprocessing pipelines: FreeSurfer 5.1 for ADNI, CAT12 for NIFD/4RTNI/PPMI) changes the operator applied to every input. The learned filter taps H were optimized for C_train; VNN stability/transferability results [11,15,18] guarantee performance only under perturbations of the covariance matrix bounded in a specific norm. The paper neither verifies that the HC-specific covariance matrices lie within this bound nor reports any sensitivity analysis. Consequently, the elevated Δ-Age values and the eigenvector inner-product differences (Fig. 5) could be driven by the mismatch between C_train and C_test, or by the arbitrary choice of the HC-specific eigenbasis, rather than by neuropathology. Additionally, the explainability analysis is partly self-referential: the eigenvectors used to 'explain' Δ-Age are the same eigenvectors that define the re-estimated filter operator. This concern affects AD, FTD, and APD claims jointly, making it load-bearing.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper applies coVariance neural networks (VNNs) to predict brain age gap (Δ-Age) from cortical thickness features in Alzheimer's disease (AD), frontotemporal dementia (FTD), atypical Parkinsonian disorders (APD), and Parkinson's disease (PD), using a VNN pre-trained on healthy controls from OASIS-3. The authors report elevated Δ-Age in AD, FTD, and APD relative to healthy controls, with distinct anatomic patterns, and claim that these patterns are explainable through the inner products between VNN regional residuals and eigenvectors of the anatomical covariance matrix. The paper emphasizes the inherent interpretability of VNNs compared with black-box brain-age models.","tokens_in":9389,"tokens_out":2175,"duration_ms":21936,"significance":"If the central claims hold, the paper would contribute a transparent, anatomy-resolved brain-age-gap biomarker for multiple neurodegenerative conditions, extending prior VNN work beyond Alzheimer's disease. The paper's strengths are its use of publicly available datasets, a transparent architectural prior based on anatomical covariance, and a clear attempt to link model representations to biological structure. However, the load-bearing inferences currently rest on informal comparisons without proper statistical testing, on a single out of ten trained models, and on a covariance re-estimation step whose validity is not established. These issues must be resolved before the significance of the reported findings can be assessed.","major_comments":[{"comment":"The statement 'For each disease dataset, we used the anatomical covariance matrix estimated only from the respective HC group in the pre-trained VNN model' changes the operator applied at inference: the coVariance filter H(C) = Σ h_k C^k depends explicitly on C, and the filter taps were optimized for the OASIS-3 training covariance. The paper does not verify that the per-dataset HC covariance matrices fall within the stability/transferability bounds established for VNNs [11,15,18], nor does it report any sensitivity analysis with respect to this covariance shift. Since the datasets also differ in preprocessing (FreeSurfer 5.1 for ADNI vs. CAT12 for NIFD/4RTNI/PPMI), the elevated Δ-Age and eigenvector inner-product differences could reflect covariance mismatch or preprocessing artifacts rather than neuropathology. This affects the AD, FTD, and APD claims jointly and is load-bearing.","section":"Section 4"},{"comment":"The contributions claim 'significantly elevated Δ-Age' for AD, FTD, and APD, but the results only report means and standard deviations (e.g., '4.67±4.04 years' for AD vs. '0 ± 2.91 years' for HC) without p-values, confidence intervals, effect sizes, or explicit statistical tests for the Δ-Age group comparisons. The word 'significantly' is therefore unsupported in the current text. Formal hypothesis tests (e.g., two-sample tests with appropriate corrections) are needed for each disease-vs-HC comparison and for the PD null result.","section":"Sections 1.3 and 4"},{"comment":"The manuscript states 'The results reported in this paper are derived from one pre-trained VNN model among the 10 that were pre-trained using the above procedure.' Selecting one model post hoc can yield results that are not representative of the model family, especially given the modest age-prediction performance (MAE 7.25 ± 0.51 years, r = 0.44). The authors should either report results aggregated across all 10 models with appropriate uncertainty, or justify why a single model is sufficient for the disease-group comparisons.","section":"Section 3.2"},{"comment":"The explainability analysis computes inner products between regional residuals and eigenvectors of the anatomical covariance matrix, but this same covariance matrix defines the VNN filter applied to the data. The analysis is therefore at least partly self-referential: both the residuals and the eigenvectors are functions of the same estimated C. Moreover, the eigenvectors reported as significant (0,1,2,6 for AD; 0,1,4,5 for FTD; 8 for APD) are selected from 68 eigenvectors without controlling the false discovery rate across the 68 comparisons. The paper should report corrected p-values or FDR, and should discuss the interpretational limits of using the same eigenvectors that define the model.","section":"Section 4 and Fig. 5"}],"minor_comments":[{"comment":"The caption and text refer to 'AT P' in Figure 3, but the paper uses 'APD' elsewhere; this abbreviation should be made consistent.","section":"Figure 3 caption"},{"comment":"The terms 'interpretability' and 'explainability' are used somewhat interchangeably; the paper should clarify whether the anatomic characterization is an interpretation of the model's internal representation or an explanation of the Δ-Age prediction, as the two are conceptually distinct.","section":"Throughout"},{"comment":"There is a typo in the sentence 'they rely onlinear-shift-and-sum operators' — it should read 'they rely on linear-shift-and-sum operators.'","section":"Section 1.1"},{"comment":"The HC group for 4RTNI is taken from NIFD because of acquisition similarity, but this choice is not tested or discussed as a potential confound; a brief justification or sensitivity analysis would strengthen the report.","section":"Section 4"},{"comment":"The phrase 'brain age gap gap' appears in the introductory section; one 'gap' should be removed.","section":"Abstract and Section 1"}],"recommendation":"major_revision","confidential_remarks":"The stress-test concern about covariance re-estimation is well-founded and, in my reading, lands directly on the principal claims of the paper. The paper would need a stability analysis or a change in protocol (e.g., using the training covariance for all inference) plus proper statistical testing to support the 'significantly elevated' language. If the authors can supply those, the paper may be suitable; otherwise the central claims are not yet supported."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing to know: this is an incremental but useful application of the authors' own VNN brain-age pipeline to FTD, APD, and PD, with a clean null result for PD. The paper's real value is the attempt to explain Δ-Age anatomically via eigenvector inner products. The anatomy for AD replicates their earlier finding; FTD hotspots in frontal/temporal regions and the APD motor-cortex/occipital pattern are biologically plausible.\n\nWhere it stumbles is the gap between phrasing and evidence. The abstract says \"significantly elevated Δ-Age\" for AD, FTD, APD, but the results section gives only means and standard deviations—no p-values, CIs, or effect sizes for those comparisons. The eigenvector differences are reported for selected eigenvectors without correction for the number tested. And the headline results come from one of ten trained models, with no sensitivity analysis for the other nine.\n\nThe stress-test concern about covariance re-estimation is real. Section 3.2 trains with a covariance matrix from OASIS-3; Section 4 says they re-estimate it from each dataset's healthy controls before applying the pretrained model. Since the VNN filter is a polynomial in C, swapping in a new C changes the operator. The paper cites transferability results but never checks whether the dataset-specific covariances lie in the promised stability region. So the disease-vs-HC differences in Δ-Age and eigenvector inner products could partly reflect covariance shift or preprocessing differences (FreeSurfer vs CAT12) rather than neuropathology. That is load-bearing because the central claim is that the anatomical patterns are disease-specific.\n\nAlso, the explainability analysis is partly self-referential: the eigenvectors used to \"explain\" Δ-Age are the same eigenvectors that define the re-estimated filter. That doesn't make the result meaningless, but it weakens the claim that the explanation is independent of the model's construction.\n\nThat said, the paper is honest about its limitations: it notes the weak age-prediction performance (MAE 7.25, r=0.44), the heavy overlap of Δ-Age distributions across diseases, and it makes all data public. These are addressable issues. A revision that adds proper statistics, runs all ten models, and checks robustness to the covariance choice would make the findings much more convincing.\n\nWho is this for? Researchers interested in interpretable brain-age biomarkers and VNN methods. It deserves a serious referee, not a desk reject. My recommendation: send to peer review, but expect heavy revision.","headline":"Incremental but useful VNN brain-age extension to FTD/APD/PD with a null PD result; headline significance claims are under-statted and covariance re-estimation needs robustness checks.","tokens_in":9929,"tokens_out":2600,"would_cite":false,"duration_ms":24516,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A covariance neural network trained only on healthy aging produces elevated brain age gaps for Alzheimer's disease, frontotemporal dementia, and atypical parkinsonian disorders, with each disease's gap traceable to how the network uses…","keywords":["brain age gap","covariance neural networks","cortical thickness","neurodegeneration","explainable AI","Alzheimer's disease","frontotemporal dementia","atypical parkinsonian disorders"],"falsifier":"Re-run the analysis with the anatomical covariance matrix fixed to the training-set covariance for every disease cohort, or with a perturbed covariance; if the disease-specific eigenvector signatures and regional maps largely disappear or shift to different eigenvectors, the explainability is an artifact of per-cohort covariance re-estimation. An independent cohort with matched preprocessing should reproduce the same significant eigenvector indices (0, 1, 2, 6 for AD; 0, 1, 4, 5 for FTD; 8 for APD).","tokens_in":8870,"feed_emoji":"🧠","tokens_out":9846,"duration_ms":84579,"temperature":0.7,"pith_summary":"This paper argues that brain age gap (Δ-Age) can be made anatomically interpretable and mechanistically explainable by computing it with a covariance neural network (VNN), a network whose layers filter cortical-thickness measurements through powers of the anatomical covariance matrix. Using 68 cortical regions from a standard brain atlas, the authors show that a VNN trained only on healthy controls detects significantly elevated Δ-Age for Alzheimer's disease, frontotemporal dementia, and atypical parkinsonian disorders, but not for Parkinson's disease. The elevation is traced to disease-relevant regions and to specific eigenvectors of the anatomical covariance matrix that the network weighs differently for each disease. If correct, the result is a biomarker that localizes accelerated aging and says which anatomical modes carry it, rather than giving a single opaque number.","feed_headline":"Brain-age gaps show disease-specific brain maps in explainable model","feed_subtitle":"VNNs link elevated brain-age gap to specific cortical regions and covariance eigenmodes in three conditions.","key_machinery":"The coVariance filter $H(C)=\\sum_{k=0}^{K} h_k C^k$, with $C$ the $68\\times 68$ anatomical covariance matrix of cortical thickness across regions. Since a covariance filter is equivalent to a PCA transform, the VNN's representations are steered by the eigenvectors of $C$; the paper makes this explicit by computing inner products between final-layer regional residuals and those eigenvectors. The residual construction converts the network output into per-region contributions, so elevated Δ-Age can be assigned to specific anatomy and to specific eigenmodes.","core_discovery":"The central discovery is that a VNN pretrained on a healthy-aging population separates neurodegenerative conditions by Δ-Age and by how it uses the eigenspectrum of the anatomical covariance matrix. With the covariance matrix re-estimated from each dataset's own healthy controls, the model gives Δ-Age = 4.67 ± 4.04 years for Alzheimer's disease (healthy controls 0 ± 2.91), 6.17 ± 4.55 for frontotemporal dementia, and 2.49 ± 3.09 for atypical parkinsonian disorders, while Parkinson's disease shows no significant elevation. Regional residuals at the network's final layer map to disease-plausible cortex: medial temporal, entorhinal, and temporo-parietal regions for AD; frontal and temporal regions for FTD; and motor and occipital regions for APD. Inner products of these residuals with the eigenvectors of the anatomical covariance matrix differ significantly between each disease and its healthy controls — eigenvectors 0, 1, 2, and 6 for AD; 0, 1, 4, and 5 for FTD; and 8 for APD — indicating the distinct Δ-Age patterns arise from the network processing each disease along different covariance eigenmodes.","pith_inferences":["The paper re-estimates the anatomical covariance matrix per dataset; a direct stress test the paper does not report is to hold the covariance fixed to the training set and verify that the same eigenvectors and regions remain significant, separating neurodegeneration signal from covariance-estimation effects.","The significant eigenvector sets (0,1,2,6; 0,1,4,5; 8) are candidates for transdiagnostic signatures; an independent study could correlate these eigenvector loadings with clinical severity, cognitive decline, or longitudinal atrophy.","The same residual-to-eigenvector accounting could be applied to other morphometric features (volume, surface area, subcortical thickness) to test whether the disease-specific eigenmode signatures are anatomy-specific or generalize across modalities."],"forward_implications":["A VNN trained only on healthy aging can flag AD, FTD, and APD through elevated Δ-Age, so disease labels are not needed during training for the biomarker to work.","Parkinson's disease shows no significant Δ-Age elevation on cortical thickness, so the pipeline separates conditions with cortical accelerated aging from those without.","Because Δ-Age distributions overlap across diseases, the anatomical and eigenvector characterizations carry information that the scalar gap alone does not.","Differences in which eigenvectors are significant (several for AD and FTD, only one for APD) explain why Δ-Age elevation is smaller for APD than for the other disease groups."],"supporting_citations":[{"why":"Establishes the equivalence between covariance filters and PCA and provides the theoretical grounding for interpreting VNN outputs through covariance eigenvectors.","marker":"[11]"},{"why":"Introduces the VNN-based explainable brain age gap pipeline and the regional residual construction that this paper extends to multiple disease cohorts.","marker":"[10]"},{"why":"Provides transferability guarantees for VNNs across datasets and atlases, supporting the use of a model trained on one healthy population in other cohorts.","marker":"[15]"},{"why":"Supplies the healthy-aging cortical thickness dataset used to pretrain the VNN.","marker":"[23]"},{"why":"Provides the linear regression bias-adjustment used to compute brain age and Δ-Age from raw network estimates.","marker":"[25]"},{"why":"Provides the hyperparameter search used to select the VNN architecture and training configuration.","marker":"[24]"}],"fun_headline_variants":["VNN explains brain-age gaps with disease-specific cortical maps","Explainable network ties brain-age gap to distinct dementia patterns","Covariance eigenvectors reveal how brain-age gaps differ across diseases","Brain-age gaps in three diseases traced to distinct anatomical eigenmodes","Explainable VNN separates Alzheimer's, FTD, and APD by brain-age maps"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The anatomical covariance matrix that the network uses is re-estimated from each dataset's own healthy control group rather than kept fixed to the one used in training, so the disease-specific patterns could partly reflect differences in covariance estimation or preprocessing rather than neurodegeneration.","fun_headline_variants_meta":{"raw":{"variants":["VNN explains brain-age gaps with disease-specific cortical maps","Explainable network ties brain-age gap to distinct dementia patterns","Covariance eigenvectors reveal how brain-age gaps differ across diseases","Brain-age gaps in three diseases traced to distinct anatomical eigenmodes","Explainable VNN separates Alzheimer's, FTD, and APD by brain-age maps"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000245,"raw_usage":{"total_tokens":1575,"prompt_tokens":1021,"completion_tokens":554,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":637,"completion_tokens_details":{"reasoning_tokens":463}},"tokens_in":637,"tokens_out":554,"duration_ms":6126,"temperature":1.0,"reasoning_tokens":463,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T22:26:48.290533+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the analysis with the anatomical covariance matrix fixed to the training-set covariance for every disease cohort, or with a perturbed covariance; if the disease-specific eigenvector signatures and regional maps largely disappear or shift to different eigenvectors, the explainability is an artifact of per-cohort covariance re-estimation. An independent cohort with matched preprocessing should reproduce the same significant eigenvector indices (0, 1, 2, 6 for AD; 0, 1, 4, 5 for FTD; 8 for APD).","supporting_citations":[{"cited_title":"Ten years of brainage as a neuroimaging biomarker of brain aging: What insights have we gained?,","cited_arxiv_id":null,"evidence_quote":"Establishes the equivalence between covariance filters and PCA and provides the theoretical grounding for interpreting VNN outputs through covariance eigenvectors."},{"cited_title":"Predicting age using neu- roimaging: Innovative brain ageing biomarkers,","cited_arxiv_id":null,"evidence_quote":"Introduces the VNN-based explainable brain age gap pipeline and the regional residual construction that this paper extends to multiple disease cohorts."},{"cited_title":"A systematic review of multimodal brain age studies: Uncovering a divergence between model accuracy and utility,","cited_arxiv_id":null,"evidence_quote":"Provides transferability guarantees for VNNs across datasets and atlases, supporting the use of a model trained on one healthy population in other cohorts."},{"cited_title":"Stability properties of graph neural networks,","cited_arxiv_id":null,"evidence_quote":"Provides the linear regression bias-adjustment used to compute brain age and Δ-Age from raw network estimates."},{"cited_title":"Predicting brain age using transferable coVari- ance neural networks,","cited_arxiv_id":null,"evidence_quote":"Provides the hyperparameter search used to select the VNN architecture and training configuration."}],"review_version":1}