REVIEW 2 major objections 5 minor 22 references
Trajectory shape, captured by topological entropy, predicts conversion from mild cognitive impairment to Alzheimer's disease, with per-patient uncertainty sets.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 17:56 UTC pith:6G2L6LYV
load-bearing objection A transparent TDA-AD conversion paper with a genuinely useful leakage audit and a plausible survival signal, but the 'formally valid' conformal guarantee does not actually hold for the OOF procedure used, and the headline biomarker is partly confounded by visit count. the 2 major comments →
Calibrated Alzheimer's Conversion Risk in Mild Cognitive Impairment: Persistent Homology of Clinical Trajectories with Conformal Guarantees
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central claim is that H0 persistence entropy, computed from Vietoris-Rips persistent homology on each subject's visits as a point cloud in biomarker space, is a biologically motivated topological biomarker of cognitive decline. Treating only pre-conversion visits under a uniform four-year follow-up cap, the corrected pipeline reaches an internal AUC between 0.840 and 0.866 and an external AUC of 0.879 on a temporally separated zero-overlap cohort. The paper further claims that a cross-conformal procedure provides the first distribution-free finite-sample guarantee of 90% coverage for individual conversion-risk predictions, and that survival models with TDA features improve concordance (C
What carries the argument
The load-bearing object is H0 persistence entropy, H = - sum (l_i/L) ln(l_i/L), computed from the zero-dimensional persistence diagram of a Vietoris-Rips filtration built on each patient's longitudinal visits in twelve-dimensional biomarker space. It is paired with a class-conditional split-conformal wrapper that emits per-patient prediction sets {0}, {1}, or ambiguous {0,1} at a target miscoverage of 0.10. The entropy is described as measuring trajectory complexity, and is interpreted as lower in APOE4 carriers, implying more stereotyped decline.
Load-bearing premise
The formal coverage guarantee rests on out-of-fold cross-validation scores being exchangeable with test scores; cross-conformal predictors are only approximately exchangeable, so the 90% coverage may not hold exactly in finite samples.
What would settle it
Compute the empirical coverage of the described cross-conformal procedure on a synthetic dataset where labels are independent of all features; if coverage falls systematically below 90% across thousands of replications, the formal guarantee fails. In the MCI cohort itself, replace out-of-fold scores with a strict split-conformal calibration set and check whether coverage deviates materially from 90%.
If this is right
- If H0 entropy is a true biomarker, it provides a computationally cheap topological feature usable in clinical batch processing pipelines.
- The conformal prediction sets give individual-level uncertainty for each MCI patient, enabling trial enrichment that excludes ambiguous cases or flags them for closer monitoring.
- The five-item leakage audit can be applied to other longitudinal electronic health record prediction studies, potentially reducing inflated AUC claims across the field.
- TDA's predictive value is most clearly expressed in time-to-event models, so future studies should evaluate topological features in survival settings rather than only fixed-horizon classification.
- External calibration is overconfident (slope 2.06), so any clinical deployment requires local recalibration before probabilities are used.
Where Pith is reading between the lines
- The exchangeability assumption for out-of-fold cross-conformal scores is not exactly finite-sample valid; the stated 'formally valid' 90% guarantee should be treated as approximate, since a strict split-conformal calibration set would be exact but uses less data.
- The APOE4-entropy association attenuates after adjusting for visit count and follow-up length, so the 'genetic validation' may partly reflect observation density rather than pure biology; replication on a cohort with matched visit counts would settle this.
- The same trajectory-entropy feature could be studied in other progressive diseases, such as Parkinson's or frontotemporal dementia, as a general marker of stereotyped decline.
- Combining topological features with more powerful time-to-event architectures than penalized Cox could yield larger survival concordance gains than the +0.045 reported here.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a leak-audited machine-learning pipeline for predicting MCI-to-AD conversion from longitudinal clinical trajectories, adding persistent-homology features (notably H0 persistence entropy) to conventional slopes and engineered features. It reports an internal nested AUC of 0.840, external AUC of 0.879 on a temporally separated ADNI cohort, improved concordance in survival models with TDA features (Cox +0.045, RSF +0.014), an association between H0 entropy and APOE4 dosage, and a conformal prediction procedure claimed to provide the first formally valid individual-level risk guarantee for this task, with internal coverage 90.4% and external coverage 96.9%.
Significance. The study is unusually transparent in several respects: it quantifies five leakage sources step by step (Table 2), implements a nested feature-selection protocol to estimate optimism (+0.026 AUC), performs temporal external validation on a zero-overlap cohort, reports calibration and decision-curve analyses, and includes a subgroup fairness audit. If the results are taken at face value, the leakage-audit framework and the proposal of H0 persistence entropy as a computationally cheap topological biomarker are useful contributions. However, the paper's central conformal-guarantee claim is not supported as stated, and the APOE4 association conclusion overstates the evidence. These issues are load-bearing because C3 and part of C2 are headline contributions, but they are fixable within the manuscript's scope.
major comments (2)
- [§3.5, Eq. (4), Conclusion] The claim that the cross-conformal procedure yields a 'formally guaranteed' 90.4% coverage is not supported. Out-of-fold calibration scores are produced by models trained without the subject's fold, while a test score is produced by models trained on all calibration data (or an aggregate of fold models); the score vector is not exchangeable, so the finite-sample lemma behind Eq. (4) does not apply. The paper itself contains contradictory statements in §3.5: first 'We emphasise that this is an empirical observation, not a formal guarantee,' then 'satisfying the exchangeability requirement within the training set ... a formally guaranteed result.' Because C3 is a headline contribution, the guarantee must either be re-derived with a split-conformal calibration set disjoint from training or be re-labeled empirical. The abstract's 'split-conformal' wording is also inaccurate for the implement
- [§3.6 vs. Conclusion] The conclusion states that the APOE4–H0 entropy association 'persists after controlling for disease stage and visit count,' but §3.6 reports that after full adjustment including visit count, r=0.057, p=0.12, i.e., not significant. The paper's own text acknowledges 'partial confounding by observation density.' Thus the conclusion overstates the evidence. The pre-visit-count-adjusted association (r=-0.111, p=0.003 after controlling for conversion status and MCI subtype) is still worth reporting as hypothesis-generating, but the 'persists' language must be removed or qualified.
minor comments (5)
- [Abstract, §2.5] The method is cross-conformal using out-of-fold scores, but the abstract and introduction call it 'split-conformal.' Please align terminology throughout.
- [§3.5] The contradictory statements about formal vs. empirical coverage should be resolved; see major comment 1.
- [Figure 1, §3.1] It is unclear whether the internal 5-fold CV is performed on the full 741-subject cohort or on the ADNI-1 subset (n=330) used as the temporal training set. Table 2 reports N=741 while Figure 1 labels 'Internal — ADNI-1 n=330'. Please specify which cohort the internal results refer to and ensure consistency of all reported internal metrics.
- [§2.3] There is a LaTeX artifact in the text: 'extttclass_weight' should be rendered as class_weight='balanced'.
- [§3.6] The within-fixed-visit-count-strata null results are important; consider placing them more prominently so the reader does not miss that the APOE4 association is partly driven by observation density.
Circularity Check
Mild, disclosed feature-selection circularity is quantified by the paper; the central TDA biomarker claim is not definitionally circular. The conformal 'formal guarantee' is unsupported because OOF scores are not exactly exchangeable, which is a correctness issue rather than a definitional circularity.
specific steps
-
other
[Section 2.3 (Models and Evaluation); Section 4.5 (Limitations)]
"To quantify this circularity, we implemented a nested selection protocol: in each outer fold, a 60% inner split was used to compute SHAP values and select the top-20 features, and the final model was trained on the full outer training set using those features. The nested selection AUC was 0.840 versus 0.866 for same-fold selection, yielding an optimism estimate of +0.026 AUC points."
The top-20 feature set is selected using SHAP values derived from the same 5-fold CV subjects on which the final Stack+Top20 model is evaluated, so the 0.866 same-fold AUC is an optimistically selected quantity rather than an unbiased out-of-sample prediction. The paper explicitly labels this circularity and reports the nested-selection 0.840 as the primary unbiased estimate, making the circular step disclosed and bounded rather than the hidden driver of the conclusions.
-
other
[Section 3.5 (Survival Analysis and Conformal Guarantees); Methods 2.5]
"Internal (formally valid): We implemented a cross-conformal procedure [15] in which each subject is scored by a model trained without it (5-fold CV), satisfying the exchangeability requirement within the training set. This yields 90.4% ± 2.2% coverage against the 90% target—a formally guaranteed result."
The claimed formal finite-sample guarantee is derived from an exchangeability premise that the described OOF scoring procedure does not satisfy: each calibration score is produced by a model trained without that subject's fold, while a test score is produced by a model trained on all calibration subjects (or by averaging fold models). The vector of calibration scores plus the test score is therefore not exchangeable, so the quantile lemma behind Eq. (4) does not apply. The 90.4% figure is an empirical coverage observation relabeled as a theorem; this is primarily an unsupported correctness claim load-bearing for contribution C3, not a definitional circularity.
full rationale
The paper's central TDA claim is not circular by construction: H0 persistence entropy is measured from raw clinical trajectory point clouds via a Vietoris-Rips filtration and this does not use the conversion label as an input. The leakage audit and external temporal validation on a zero-overlap ADNI-2/GO/3 cohort provide independent checks. The main circularity in the paper is the SHAP-based feature selection performed on the same CV folds used for evaluation, which the authors transparently acknowledge and quantify as +0.026 AUC via a nested selection protocol; since they report the nested estimate (0.840) as the primary unbiased AUC, this is a partial, disclosed circularity rather than a forced result. No load-bearing self-citation chain is present: reference [15] is an external textbook, and the 'uniqueness' of the conformal guarantee is not imported from the authors' prior work. The conformal-coverage issue flagged above is important but is a correctness/mathematical-support problem, not a definitional equivalence between inputs and outputs. Overall, the derivation chain is largely self-contained, with one admitted bounded circularity and one unsupported formal-guarantee claim, justifying a moderate rather than high circularity score.
Axiom & Free-Parameter Ledger
free parameters (5)
- Top-20 feature selection cutoff =
20
- Delay embedding parameters =
m=2, tau=1
- Conformal alpha =
0.10
- Uniform follow-up cap =
4 years
- RF class_weight='balanced' =
balanced
axioms (4)
- domain assumption Vietoris-Rips persistent homology and Takens embedding are appropriate for these short, discrete clinical trajectories
- domain assumption ADNI diagnoses are not algorithmically derived from CDRSB, so using CDRSB is not label leakage
- ad hoc to paper Out-of-fold cross-validation scores satisfy exchangeability for split-conformal guarantees
- standard math Standard probability and statistical inference assumptions (DeLong test, bootstrap, log-rank)
invented entities (1)
-
H0 persistence entropy as a topological biomarker
no independent evidence
read the original abstract
Background. Predicting conversion from mild cognitive impairment (MCI) to Alzheimer's disease (AD) is central to trial enrichment and care planning, yet existing models provide no individual-level uncertainty estimates and rarely include transparent leakage audits. We introduce the first application of persistent homology to longitudinal clinical trajectory point clouds for this task, and the first split-conformal individual risk guarantee for any AD-conversion model. Methods. We analysed 741 MCI subjects (240 converters, 32.4%) from ADNI with a uniform 4-year follow-up cap. Five leakage sources were corrected; without them a naive pipeline achieved AUC=0.934, inflated by +0.075. Vietoris-Rips persistent homology and sublevel-set proxies were combined with trajectory slopes and engineered features (76 total) in a stacking ensemble evaluated by 5-fold cross-validation. Results. Cox and Random Survival Forest models with TDA features achieved concordance C=0.799 and C=0.826 versus C=0.753 and C=0.812 without (+0.045 and +0.014). The primary nested AUC is 0.840 (same-fold bound 0.866); external AUC was 0.879 on a zero-overlap ADNI-2/GO/3 cohort. H0 persistence entropy was the top SHAP feature and significantly associated with APOE4 dosage (Spearman r=-0.191, p<0.0001, Bonferroni-corrected). Cross-conformal coverage was 90.4%+-2.2% (target 90%); empirical external coverage 96.9%. Maximum fairness gap in false-negative rate across seven subgroups was 0.092. Conclusions. We propose H0 persistence entropy as a topological biomarker of cognitive decline and demonstrate that a leakage-audited, conformally calibrated pipeline reaches competitive accuracy with individual-level uncertainty quantification not previously available for this task.
Figures
Reference graph
Works this paper leans on
-
[1]
Mild cognitive impairment: clinical characteri- zation and outcome.Arch Neurol
Petersen RC, Smith GE, Waring SC, et al. Mild cognitive impairment: clinical characteri- zation and outcome.Arch Neurol. 1999;56(3):303–308
1999
-
[2]
Sparse learning and stability selection for predicting MCI to AD conversion.BMC Neurol
Ye J, Wu T, Li J, Chen K. Sparse learning and stability selection for predicting MCI to AD conversion.BMC Neurol. 2012;12:46
2012
-
[3]
Hierarchical fully convolutional network for joint atro- phy localization and Alzheimer’s disease diagnosis.IEEE Trans Pattern Anal Mach Intell
Lian C, Liu M, Zhang J, Shen D. Hierarchical fully convolutional network for joint atro- phy localization and Alzheimer’s disease diagnosis.IEEE Trans Pattern Anal Mach Intell. 2020;42(4):880–893
2020
-
[4]
Leakage and the reproducibility crisis in machine-learning-based science.Patterns
Kapoor S, Narayanan A. Leakage and the reproducibility crisis in machine-learning-based science.Patterns. 2023;4(9):100804
2023
-
[5]
An introduction to topological data analysis: fundamental and practical aspects for data scientists.Front Artif Intell
Chazal F, Michel B. An introduction to topological data analysis: fundamental and practical aspects for data scientists.Front Artif Intell. 2021;4:667963
2021
-
[6]
White matter brain network research in Alzheimer’s disease using persistent features.Molecules
Kuang L, Han X, Chen K, Caselli RJ, Thompson PM, Wang Y. White matter brain network research in Alzheimer’s disease using persistent features.Molecules. 2020;25(11):2472
2020
-
[7]
A spatiotemporal brain network analysis of Alzheimer’s disease based on persistent homology.Front Aging Neurosci
Zhao X, Wu C, Cheng L, et al. A spatiotemporal brain network analysis of Alzheimer’s disease based on persistent homology.Front Aging Neurosci. 2022;13:779508
2022
-
[8]
American Mathemat- ical Society; 2010
Edelsbrunner H, Harer J.Computational Topology: An Introduction. American Mathemat- ical Society; 2010
2010
-
[9]
The Alzheimer’s Disease Neuroimaging Initiative 3: continuedinnovationforclinicaltrialimprovement.Alzheimers Dement.2017;13(5):561–571
Weiner MW, Veitch DP, Aisen PS, et al. The Alzheimer’s Disease Neuroimaging Initiative 3: continuedinnovationforclinicaltrialimprovement.Alzheimers Dement.2017;13(5):561–571
2017
-
[10]
Ripser: efficient computation of Vietoris-Rips persistence barcodes.J Appl Comput Topol
Bauer U. Ripser: efficient computation of Vietoris-Rips persistence barcodes.J Appl Comput Topol. 2021;5(3):391–423
2021
-
[11]
Detecting strange attractors in turbulence
Takens F. Detecting strange attractors in turbulence. In:Dynamical Systems and Turbu- lence, Lecture Notes in Mathematics vol898. Springer; 1981:366–381
1981
-
[12]
Scikit-learn: machine learning in Python.J Mach Learn Res
Pedregosa F, Varoquaux G, Gramfort A, et al. Scikit-learn: machine learning in Python.J Mach Learn Res. 2011;12:2825–2830
2011
-
[13]
XGBoost: a scalable tree boosting system
Chen T, Guestrin C. XGBoost: a scalable tree boosting system. In:Proc. 22nd ACM SIGKDD. 2016:785–794
2016
-
[14]
A unified approach to interpreting model predictions.Adv Neural Inf Process Syst
Lundberg SM, Lee SI. A unified approach to interpreting model predictions.Adv Neural Inf Process Syst. 2017;30
2017
-
[15]
Vovk V, Gammerman A, Shafer G.Algorithmic Learning in a Random World. 2nd ed. Springer International Publishing; 2022. doi:10.1007/978-3-031-06649-8
-
[16]
Comparing the areas under two or more correlated receiver operating characteristic curves.Biometrics
DeLong ER, DeLong DM, Clarke-Pearson DL. Comparing the areas under two or more correlated receiver operating characteristic curves.Biometrics. 1988;44(3):837–845
1988
-
[17]
Hosmer DW, Lemeshow S, Sturdivant RX.Applied Logistic Regression. 3rd ed. Wiley; 2013
2013
-
[18]
Evaluating the added predic- tive ability of a new marker: from area under the ROC curve to reclassification and beyond
Pencina MJ, D’Agostino RB Sr, D’Agostino RB Jr, Vasan RS. Evaluating the added predic- tive ability of a new marker: from area under the ROC curve to reclassification and beyond. Stat Med. 2008;27(2):157–172
2008
-
[19]
Decision curve analysis: a novel method for evaluating prediction models.Med Decis Making
Vickers AJ, Elkin EB. Decision curve analysis: a novel method for evaluating prediction models.Med Decis Making. 2006;26(6):565–574
2006
-
[20]
Lyu Q, Hudson J, Kawas M, Jiang Y, You C, Whitlow CT. PROMISE-AD: progression- aware multi-horizon survival estimation for Alzheimer’s disease progression and dynamic 16 tracking.arXiv. 2026;arXiv:2604.28055
Pith/arXiv arXiv 2026
-
[21]
An empirical characterization of fair machine learning for clinical risk prediction.J Biomed Inform
Pfohl SR, Foryciarz A, Shah NH. An empirical characterization of fair machine learning for clinical risk prediction.J Biomed Inform. 2021;113:103621. doi:10.1016/j.jbi.2020.103621
arXiv 2021
-
[22]
TRIPOD+AI statement: updated guidance for reporting clinical prediction models.BMJ
Collins GS, Moons KGM, Dhiman P, et al. TRIPOD+AI statement: updated guidance for reporting clinical prediction models.BMJ. 2024;385:e078378. 17
2024
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.