REVIEW 4 major objections 5 minor 15 references
petBrain: A New Pipeline for Amyloid, Tau Tangles and Neurodegeneration Quantification Using PET and MRI
T0 review · 4 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read petBrain, a fully automated web pipeline, produces A/T2/N Alzheimer's biomarker scores from PET and MRI that match established SPM and B-PIP pipelines, and it is the first to estimate all three biomarkers simultaneously.
desk verdict Solid integrated A/T2/N pipeline with strong aggregate ADNI validation; the cross-tracer calibration claims are broader than the per-tracer evidence. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is the combination of three components: (1) AssemblyNet, a large ensemble of convolutional neural networks that segments 132 brain structures from a T1-weighted MRI in under 15 minutes, yielding subject-specific regions of interest; (2) affine calibration equations, fitted on GAAIN datasets, that map the partial-volume-corrected SUVr from petBrain's subject-specific masks onto the universal Centiloid and CenTauRz scales; and (3) the HAVAs score, computed from the hippocampal, amygdala, and inferior lateral ventricle volumes against a pathological lifespan model. The subject-specific masks are the enabling object: they replace the fixed standardized masks of the Centiloid and CenTauR projects while the calibration equations preserve cross-tracer comparability, and the same MRI segmentation feeds the N biomarker, which is why a single pipeline can output A, T2, and N together.
What would settle it
Run petBrain on a cohort scanned with a tracer not in the calibration set (e.g., 18F-RO948 or 18F-GTP1) and compare its CTRz against a reference SPM-based CenTauRz; a systematic residual bias in the linear regression, or an ICC below the 0.92-0.98 range reported on ADNI, would falsify the claim that the affine calibration generalizes.
Extended reading notes
Core claim
The central discovery is that subject-specific anatomical masks produced by the AssemblyNet deep-learning segmentation can replace the fixed MNI-space masks of the Centiloid and CenTauR projects and still reproduce the same standardized values. After Level-1 calibration with the PiB dataset and Level-2 tracer-to-tracer regressions, petBrain's Centiloid values correlate at R2=0.94 with SPM-based Centiloid and R2=0.96 with B-PIP; CenTauRz values correlate at R2=0.84 with SPM and tau SUVr at R2=0.95 with B-PIP. The pipeline also assigns A+/A-, T2+/T2-, and N+/N- status using external thresholds (AMYPAD/CenTauR/HAVAs), and on ADNI these statuses and scores increase monotonically from amyloid-negative controls through amyloid-positive controls, MCI, and dementia, with combined A/T2/N explaining more variance in cognitive scores than any single biomarker. The authors therefore claim that a single unified, web-deployed pipeline can deliver biologically and clinically meaningful A/T2/N quantification.
Load-bearing premise
The calibration-generalization assumption: the linear equations that convert petBrain SUVr to Centiloid and CenTauRz scales were fitted on small GAAIN cohorts and are assumed to hold across scanners, tracers, and populations without residual bias.
Editorial extensions
If this is right
- Research groups without local imaging software or GPU infrastructure can obtain standardized CL, CTRz, and HAVAs values for a subject in about 20 minutes through the volBrain website.
- Because A and T2 are expressed on universal scales, results from different amyloid tracers (PiB, FBP, FBB, FTM, NAV) and tau tracers (FTP, MK, PI, and five more via conversion equations) can be pooled and compared across studies.
- The combined A/T2/N model predicts CDR-sb, MMSE, and MoCA better than amyloid load alone, supporting the three-biomarker framework for cognitive staging.
- On ADNI, the pipeline's A/T2/N staging separated amyloid-positive CN, MCI, and dementia groups for all three biomarkers with high significance.
- The paper claims petBrain is the first pipeline to estimate A, T2, and N simultaneously in one processing framework.
Reading between the lines
- The calibration-generalization claim is testable outside ADNI: if the same affine equations are fitted per cohort with a bias-correction term, the paper's approach could be extended to novel tracers and scanners, but the authors have not demonstrated that such corrections are needed.
- Because PVC changed results only marginally, a future MRI-free or PVC-free variant might preserve A/T2 quantification for research contexts where MRI is unavailable, at the cost of losing the N biomarker.
- The high concordance between subject-specific-mask pipelines (petBrain and B-PIP) suggests that deep-learning masks could substitute for FreeSurfer in other quantification pipelines, although the paper does not test this directly.
- A natural next step the authors do not take is longitudinal validation: showing that changes in petBrain CL, CTRz, and HAVAs track treatment response or disease progression would determine whether the pipeline is useful for monitoring disease-modifying therapies.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents petBrain, a web-based pipeline that combines AssemblyNet deep-learning segmentation of T1w MRI with PET processing to compute amyloid Centiloid (CL), tau CenTauRz (CTRz), and HAVAs neurodegeneration scores, and thereby to assign A/T2/N status. Calibration equations for CL and CTRz are derived on GAAIN cohorts and validated on ADNI against SPM and B-PIP pipelines. The authors report high concordance (R2 = 0.84 to 0.96, ICC = 0.92 to 0.98), expected associations with CSF and plasma biomarkers, and stronger combined A/T2/N associations with cognition than single-biomarker models.
Significance. If the calibration transfer is reliable, petBrain is a practically valuable contribution: it is freely accessible through the volBrain platform, processes a full A/T2/N workup in about 20 minutes per subject, uses subject-specific deep-learning segmentations, and is the first pipeline described as simultaneously estimating A, T2, and N. The ADNI validation is a genuine external test with strong aggregate concordance against two established pipelines, and the biological correlates (fluid biomarkers, clinical status, cognition) are directionally consistent with the AD literature. The manuscript also ships concrete calibration equations and a public implementation, which aids reproducibility. However, the paper's central cross-tracer standardization claim is broader than the evidence actually presented, and the absence of tracer-stratified validation and uncertainty quantification is load-bearing for that claim.
major comments (4)
- [Section 3.3, Table 1] The ADNI validation is pooled across tracers: 780 of the 821 tau scans are FTP, with only 23 MK and 18 PI scans, and the amyloid validation uses only FBB and FBP, leaving FTM and NAV unvalidated. The conversion equations for MK/PI in Table 3 are derived indirectly from the FTP Level-1 calibration plus published constants, not from petBrain-specific data. Because the central claim is that petBrain produces standardized CTRz values across tracers, the pooled R2=0.84/0.92 could conceal a tracer-specific bias that would directly affect T2 status. Please add tracer-stratified R2, ICC, and Bland-Altman limits of agreement for all tracers represented in the ADNI sample, or explicitly restrict the cross-tracer claim to the tracers that are actually validated.
- [Tables 2 and 3, Eqs. (1)-(2)] The calibration regressions are reported without standard errors, confidence intervals, or residual-bias diagnostics, despite being fitted on small GAAIN samples (for example, 79 PiB subjects for the Level-1 CL calibration and 46 FBP subjects for the FBP conversion). These affine parameters are load-bearing because they map SUVr to CL and CTRz and hence determine the A and T2 binary status. Please report confidence intervals for all slopes and intercepts, and ideally propagate the parameter uncertainty into CL/CTRz values or perform a sensitivity analysis across scanners, tracers, or cohorts.
- [Section 2.2, N status and HAVAs] The HAVAs score is the authors' own previously published model, and the manuscript states that ADNI subjects were excluded from its construction but provides no verifiable documentation or code to support this claim. If any ADNI participants were part of the HAVAs training data, the validation correlations in Tables 4 and 5 for the N biomarker would be partly circular. Please provide the exclusion procedure, a list of overlapping subjects, or a clear description of the HAVAs training cohort separately from the ADNI validation cohort.
- [Section 2.2, Centiloid subject-specific mask] The subject-specific Centiloid mask is selected on the GAAIN-PiB dataset with a Cohen's d > 5 threshold, and the same dataset is then used for the Level-1 CL calibration in Eq. (1). This double use of the calibration cohort may overstate the reported calibration fit in Supplementary Fig. 3. The ADNI validation is independent, but the stability of the mask-selection procedure itself is not assessed. Please report a bootstrap or resampling-based stability analysis of the selected mask, or demonstrate that the d>5 threshold does not materially change CL values.
minor comments (5)
- [Table 1] The table header lists 'ADNI N = 831', while the text in Section 2.1 states that the study used 821 subjects; the row sums in the table equal 821, so this appears to be a typographical inconsistency that should be corrected.
- [Figure 5] The caption states that 'Amyloid status was established with the B-PIP PET pipeline and ADNI thresholds,' which conflicts with the surrounding text's presentation of petBrain-derived staging; please clarify which pipeline defines the A-group labels used in the ANOVA analyses in Section 3.4.
- [Throughout] There are inconsistent spellings of 'CenTauR': 'CenTaur' and 'Centaur' appear in Figure 2 and in the main text; the standardized spelling should be used consistently.
- [Legends and Supplementary Material] The legends section lists 'Figure 2' twice (once for the global overview and once for the comparison with SPM), and the supplementary material contains two different items numbered 'Supplementary Figure 5'; renumbering is needed.
- [Section 2.1] The text states that 'we used 499 amyloid-PET images' while Table 1 reports GAAIN N=375; the relationship between the number of images, the number of participants, and the tracer-specific counts should be stated more clearly.
Circularity Check
Core validation is independent (ADNI, external SPM/B-PIP); only the GAAIN calibration checks partly reduce to their own fits.
-
self definitional
[Section 3.1, Eq. (1) and Supplementary Fig. 1 (PiB Level-1 CL check)]
"CL = 100 x (PiBSUVrpetBrain - 0.9659) / (1.8972 - 0.9659) Eq. (1) ... we compared the CL value obtained by petBrain with the CL values published by the Centiloid Project7 using the official SPM8-based pipeline ... our calibration yielded to very high correlation between CL PiB published by the Centiloid Project7 and CL PiB obtained with petBrain ... (y=0.99x + 0.57; R2=0.99)"
Level-1 Centiloid calibration defines the scale by setting the yCN A- group mean to 0 and the AD A+ group mean to 100 on the same GAAIN-PiB dataset used for the comparison. The constants 0.9659 and 1.8972 in Eq. (1) are the petBrain SUVr anchor values (yCN A- and AD A+ group means) produced by that Level-1 procedure, and the published SPM8 Centiloid values were defined with the same anchors on the same PiB subjects. The two methods therefore agree at 0 and 100 by construction, and the PiB sample contains only these two extreme groups, so the reported R2=0.99 largely reflects the shared anchor definition rather than an independent prediction. The independent evidence for equivalence is the ADNI comparison, not this calibration check.
-
fitted input called prediction
[Section 3.2, Eq. (2) and Supplementary Fig. 3 (FTP MetaTemporal calibration check)]
"FTPSUVrCTR = (FTPSUVrpetBrain - 0.2222) / 0.7646 Eq. (2) ... Third, once calibrated, we compared the CTRz values obtained by petBrain with the CTRz values published by the CenTauR Project12 using their SPM8 pipeline. Supplementary Figure 3 presents the results of this comparison using their Meta-temporal mask (y=0.9804x +0.096; R2=0.9803)."
Eq. (2) is the inverse of the linear regression fitted on the same 100-subject FTP GAAIN dataset between the published SPM8 MetaTemporal FTP SUVr/CTRz and petBrain FTP SUVr. Reporting R2=0.9803 after applying that fitted equation to the same training data is a restatement of the fit, not a validation of the calibration. The MetaTemporal mask was also selected on this same dataset as the mask with the highest R2 among the five CenTauR masks, further making the subsequent comparison a selected-on-training-data result. This does not affect the independent ADNI validation, but it should not be read as evidence of cross-tracer generalization.
full rationale
The paper's principal claims are tested on ADNI data that were not used to fit the CL/CTRz calibration equations: the petBrain CL and CTRz are compared with external SPM-based pipelines and B-PIP on 821 ADNI subjects, and the A/T2/N staging is correlated with fluid biomarkers, clinical status, and cognition on the same held-out cohort. The calibration equations in Tables 2-3 are fitted on GAAIN reference samples as they should be, and the independent ADNI agreement (R2 = 0.94-0.96 for CL, R2 = 0.84-0.95 for tau) is the load-bearing evidence. The two flagged items are confined to the GAAIN development dataset, where the paper presents training-set agreement (CL anchors and the FTP regression R2) as calibration confirmation; these are partly by construction but are not the central validation. Self-citations to HAVAs, AssemblyNet, and the lifespan models are prior externally published methods, and the paper asserts that ADNI subjects used to construct HAVAs were excluded from this study; that assertion is not verifiable from the manuscript but does not make the ADNI validation circular by construction. Overall, the derivation chain is not circular: the main predictions are genuinely out-of-sample.
Assumptions & free parameters
free parameters (5)
- Centiloid mask structure selection threshold (Cohen's d > 5) =
5
- Centiloid Level-1 anchors (PET SUVr means) =
0.9659 (yCN), 1.8972 (dementia A+)
- Level-2 Centiloid conversion slopes and intercepts for FBP, FBB, FTM, NAV =
Table 2: e.g., FBP slope 194.8721, intercept -191.8315
- CenTauR Level-1 calibration slope and intercept =
slope 0.7646, intercept 0.2222 (Eq. 2)
- CenTauR conversion constants for RO, MK, GTP, PBB3, PI =
Table 3 slopes/intercepts, e.g., MK slope 12.2417, intercept -12.7801
assumptions (5)
- domain assumption AssemblyNet segmentation correctly identifies the 132 Neuromorphometrics structures and provides accurate volumes and subject-specific ROIs for PET quantification.
- domain assumption The linear regression relationships between petBrain SUVr and published SPM-based CL/CTRz values are sufficient to harmonize PET quantification across pipelines and tracers.
- domain assumption The HAVAs score and its N+/N- threshold (probability > 0.5) are valid for the ADNI cohort, with ADNI subjects excluded from the model construction.
- standard math Whole cerebellum (amyloid) and cerebellar gray matter (tau) are valid reference regions for SUVr normalization.
- domain assumption The partial volume correction method (Manjon et al. 2010) improves or at least does not degrade PET quantification when combined with large target masks.
Cite this review
Pith. "Pith review of petBrain: A New Pipeline for Amyloid, Tau Tangles and Neurodegeneration Quantification Using PET and MRI." pith.science (2026). https://pith.science/paper/ELDODYMY
@misc{pith2026250603217,
author = {Pith},
title = {Pith review of: petBrain: A New Pipeline for Amyloid, Tau Tangles and Neurodegeneration Quantification Using PET and MRI},
year = {2026},
howpublished = {\url{https://pith.science/paper/ELDODYMY}},
note = {Machine review of arXiv:2506.03217}
}
read the original abstract
INTRODUCTION: Quantification of amyloid plaques (A), neurofibrillary tangles (T2), and neurodegeneration (N) using PET and MRI is critical for Alzheimer's disease (AD) diagnosis and prognosis. Existing pipelines face limitations regarding processing time, variability in tracer types, and challenges in multimodal integration. METHODS: We developed petBrain, a novel end-to-end processing pipeline for amyloid-PET, tau-PET, and structural MRI. It leverages deep learning-based segmentation, standardized biomarker quantification (Centiloid, CenTauR, HAVAs), and simultaneous estimation of A, T2, and N biomarkers. The pipeline is implemented as a web-based platform, requiring no local computational infrastructure or specialized software knowledge. RESULTS: petBrain provides reliable and rapid biomarker quantification, with results comparable to existing pipelines for A and T2. It shows strong concordance with data processed in ADNI databases. The staging and quantification of A/T2/N by petBrain demonstrated good agreement with CSF/plasma biomarkers, clinical status, and cognitive performance. DISCUSSION: petBrain represents a powerful and openly accessible platform for standardized AD biomarker analysis, facilitating applications in clinical research.
Figures
Reference graph
Works this paper leans on
-
[1]
Background Alzheimer’s disease (AD) is pathologically characterized by amyloid-β (Aβ) plaques, tau neurofibrillary tangles, and neurodegeneration. Recent recommendations from the Alzheimer's Association (AA) establish amyloid-PET as the gold standard for identifying brain amyloidosis (A) in vivo, and tau PET for quantifying and staging tauopathy (T2). Neu...
-
[2]
Methods 2.1. Participants GAAIN Datasets This study used datasets collected from the publicly available Global Alzheimer’s Association Interactive Network (GAAIN) repository (https://www.gaain.org). First, we used 499 amyloid-PET images from the Centiloid project obtained with five different amyloid-PET tracers: 11C-PiB (PiB), 18F-Florbetapir (FBP), 18F-F...
work page 2015
-
[3]
Results 3.1. Centiloid Calibration PiB calibration First, the GAAIN-PiB dataset was used to convert the original SPM8 Centiloid pipeline and petBrain. To this end, we performed the Level-1 calibration procedure7 and we obtained the following equation: CL = 100 x (PiBSUVrpetBrain - 0.9659) / (1.8972 - 0.9659) Eq. (1) 13 Therefore, an individual SUVr PiB va...
-
[4]
Discussion In this study, we introduced petBrain, a novel accurate and efficient processing pipeline for amyloid-PET, tau-PET, and structural MRI dedicated to AD research purposes. The petBrain pipeline enables standardized and accessible quantification of amyloid and tau burden using the CL and CTRz scales, ensuring cross-tracer comparability. Additional...
work page 2024
-
[5]
Jovalekic, A. et al. Validation of quantitative assessment of florbetaben PET scans as an adjunct to the visual assessment across 15 software methods. Eur J Nucl Med Mol Imaging 50, 3276–3289 (2023). 6. Lee, J. et al. Development of Amyloid PET Analysis Pipeline Using Deep Learning-Based Brain MRI Segmentation—A Comparative Validation Study. Diagnostics 1...
work page 2023
-
[6]
CenTauR FTP calibration Supplementary Figure 5: petBrain validation for CenTauRz (CTRz) measure: calibration with FTP tracer. The used FTP dataset was composed of 50 young A- cognitively normal subject and 50 old A+ patients with AD dementia. -5051015202530petBrain CTRz FTP -5 0 5 10 15 20 25 30Published SPM8 MetaTemporal CTRz FTP Level-1 CTRz analyses Al...
-
[7]
Validation of PET measurements without Partial Volume Correction (PVC) In this section, we evaluated the influence of the partial volume correction (PVC) step on the outcomes of our processing pipeline. As illustrated in Supplementary Figure 4, the correlation between petBrain and B-PIP remained highly consistent whether PVC was applied (see Figure 3) or ...
-
[8]
GAAIN Datasets description In this section, we provide additional information on the GAAIN data used in our study. • PiB dataset (Klunk et al. 2015): This dataset consists of 79 paired T1-w MRI and PiB PET scans acquired 50–70 minutes post-injection. It includes data from 34 young cognitively normal (yCN) controls and 45 patients with AD. The 34 yCN were ...
work page 2015
Show all 15 references
-
[9]
First, we selected all the structures with a Cohen’s d score > 5
Centiloid subject-specific mask To establish the list of structures to include into the petBrain Centiloid mask, we estimated Cohen’s d scores for each structure between the yCN A- and Dementia A+ of the PiB dataset. First, we selected all the structures with a Cohen’s d score...
-
[10]
The used PiB dataset was composed of 34 young A- cognitively normal subject and 45 A+ patients with dementia
PiB calibration Supplementary Figure 3: petBrain validation for Centiloid (CL) measure: calibration with the PiB tracer. The used PiB dataset was composed of 34 young A- cognitively normal subject and 45 A+ patients with dementia. -20020406080100120140160petBrain CL PIB-20 0 2...
-
[11]
0.811.21.41.61.822.22.4SUVr PIB0.911.11.21.31.41.51.61.71.81.9SUVr FBP petBrain: PiB vs
Other Amyloid tracers’ calibration Supplementary Figure 4: Calibration of the FBP, FBB, FTM and NAV amyloid tracers using the corresponding Centiloid Project datasets. 0.811.21.41.61.822.22.4SUVr PIB0.911.11.21.31.41.51.61.71.81.9SUVr FBP petBrain: PiB vs. FBP Young ControlsAl...
-
[12]
Supplementary Table 3: Cohen’s d scores for each structure on the young CN A- and old AD A+ of the FTP dataset
Centiloid subject-specific mask For the CenTauR mask, we used the following list of structures – entorhinal area, amygdala, parahippocampal gyrus, fusiform gyrus, inferior and middle temporal gyrus, and temporal pole. Supplementary Table 3: Cohen’s d scores for each structure ...
-
[15]
Automatically generated report by the web-based VolBrain platform Supplementary Figure 6: Example of PDF report produced by petBrain about an amyloid-positive patient with dementia in ADNI version 1.0 release 01-May-2025 Subject: ADNI_022_S_6796_MR_Accelerated_Sagittal_MPRAGE_...
2025
-
[17]
FreeSurfer
Fischl, B. FreeSurfer. Neuroimage 62, 774–781 (2012). 18. Collij, L. E. et al. Centiloid recommendations for clinical context-of-use from the AMYPAD consortium. Alzheimer’s & Dementia 20, 9037–9048 (2024). 19. Coupé, P. et al. AssemblyNet: A large ensemble of CNNs for 3D w...
- [29]
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.