Pith. sign in

REVIEW 2 major objections 5 minor 22 references

Trajectory shape, captured by topological entropy, predicts conversion from mild cognitive impairment to Alzheimer's disease, with per-patient uncertainty sets.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 17:56 UTC pith:6G2L6LYV

load-bearing objection A transparent TDA-AD conversion paper with a genuinely useful leakage audit and a plausible survival signal, but the 'formally valid' conformal guarantee does not actually hold for the OOF procedure used, and the headline biomarker is partly confounded by visit count. the 2 major comments →

arxiv 2607.17442 v1 pith:6G2L6LYV submitted 2026-07-20 cs.LG stat.ML

Calibrated Alzheimer's Conversion Risk in Mild Cognitive Impairment: Persistent Homology of Clinical Trajectories with Conformal Guarantees

classification cs.LG stat.ML
keywords mild cognitive impairmentAlzheimer's disease conversiontopological data analysispersistent homologyconformal predictiondata leakagesurvival analysisbiomarker
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper argues that the shape of a patient's longitudinal clinical trajectory, summarized by zero-dimensional persistent homology, carries genuine signal for progression from mild cognitive impairment to Alzheimer's disease. It presents a leakage audit that accounts for five sources of inflated performance, introduces a topological biomarker called H0 persistence entropy that ranks as the top predictor and associates with a genetic risk factor, and provides the first individual-level conformal risk guarantees for this prediction task. The authors claim that TDA features add meaningful concordance in survival models (Cox +0.045, survival forest +0.014) even though the fixed-horizon AUC gain is small, and they emphasize methodological transparency over state-of-the-art discrimination.

Core claim

The central claim is that H0 persistence entropy, computed from Vietoris-Rips persistent homology on each subject's visits as a point cloud in biomarker space, is a biologically motivated topological biomarker of cognitive decline. Treating only pre-conversion visits under a uniform four-year follow-up cap, the corrected pipeline reaches an internal AUC between 0.840 and 0.866 and an external AUC of 0.879 on a temporally separated zero-overlap cohort. The paper further claims that a cross-conformal procedure provides the first distribution-free finite-sample guarantee of 90% coverage for individual conversion-risk predictions, and that survival models with TDA features improve concordance (C

What carries the argument

The load-bearing object is H0 persistence entropy, H = - sum (l_i/L) ln(l_i/L), computed from the zero-dimensional persistence diagram of a Vietoris-Rips filtration built on each patient's longitudinal visits in twelve-dimensional biomarker space. It is paired with a class-conditional split-conformal wrapper that emits per-patient prediction sets {0}, {1}, or ambiguous {0,1} at a target miscoverage of 0.10. The entropy is described as measuring trajectory complexity, and is interpreted as lower in APOE4 carriers, implying more stereotyped decline.

Load-bearing premise

The formal coverage guarantee rests on out-of-fold cross-validation scores being exchangeable with test scores; cross-conformal predictors are only approximately exchangeable, so the 90% coverage may not hold exactly in finite samples.

What would settle it

Compute the empirical coverage of the described cross-conformal procedure on a synthetic dataset where labels are independent of all features; if coverage falls systematically below 90% across thousands of replications, the formal guarantee fails. In the MCI cohort itself, replace out-of-fold scores with a strict split-conformal calibration set and check whether coverage deviates materially from 90%.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • If H0 entropy is a true biomarker, it provides a computationally cheap topological feature usable in clinical batch processing pipelines.
  • The conformal prediction sets give individual-level uncertainty for each MCI patient, enabling trial enrichment that excludes ambiguous cases or flags them for closer monitoring.
  • The five-item leakage audit can be applied to other longitudinal electronic health record prediction studies, potentially reducing inflated AUC claims across the field.
  • TDA's predictive value is most clearly expressed in time-to-event models, so future studies should evaluate topological features in survival settings rather than only fixed-horizon classification.
  • External calibration is overconfident (slope 2.06), so any clinical deployment requires local recalibration before probabilities are used.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The exchangeability assumption for out-of-fold cross-conformal scores is not exactly finite-sample valid; the stated 'formally valid' 90% guarantee should be treated as approximate, since a strict split-conformal calibration set would be exact but uses less data.
  • The APOE4-entropy association attenuates after adjusting for visit count and follow-up length, so the 'genetic validation' may partly reflect observation density rather than pure biology; replication on a cohort with matched visit counts would settle this.
  • The same trajectory-entropy feature could be studied in other progressive diseases, such as Parkinson's or frontotemporal dementia, as a general marker of stereotyped decline.
  • Combining topological features with more powerful time-to-event architectures than penalized Cox could yield larger survival concordance gains than the +0.045 reported here.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper proposes a leak-audited machine-learning pipeline for predicting MCI-to-AD conversion from longitudinal clinical trajectories, adding persistent-homology features (notably H0 persistence entropy) to conventional slopes and engineered features. It reports an internal nested AUC of 0.840, external AUC of 0.879 on a temporally separated ADNI cohort, improved concordance in survival models with TDA features (Cox +0.045, RSF +0.014), an association between H0 entropy and APOE4 dosage, and a conformal prediction procedure claimed to provide the first formally valid individual-level risk guarantee for this task, with internal coverage 90.4% and external coverage 96.9%.

Significance. The study is unusually transparent in several respects: it quantifies five leakage sources step by step (Table 2), implements a nested feature-selection protocol to estimate optimism (+0.026 AUC), performs temporal external validation on a zero-overlap cohort, reports calibration and decision-curve analyses, and includes a subgroup fairness audit. If the results are taken at face value, the leakage-audit framework and the proposal of H0 persistence entropy as a computationally cheap topological biomarker are useful contributions. However, the paper's central conformal-guarantee claim is not supported as stated, and the APOE4 association conclusion overstates the evidence. These issues are load-bearing because C3 and part of C2 are headline contributions, but they are fixable within the manuscript's scope.

major comments (2)
  1. [§3.5, Eq. (4), Conclusion] The claim that the cross-conformal procedure yields a 'formally guaranteed' 90.4% coverage is not supported. Out-of-fold calibration scores are produced by models trained without the subject's fold, while a test score is produced by models trained on all calibration data (or an aggregate of fold models); the score vector is not exchangeable, so the finite-sample lemma behind Eq. (4) does not apply. The paper itself contains contradictory statements in §3.5: first 'We emphasise that this is an empirical observation, not a formal guarantee,' then 'satisfying the exchangeability requirement within the training set ... a formally guaranteed result.' Because C3 is a headline contribution, the guarantee must either be re-derived with a split-conformal calibration set disjoint from training or be re-labeled empirical. The abstract's 'split-conformal' wording is also inaccurate for the implement
  2. [§3.6 vs. Conclusion] The conclusion states that the APOE4–H0 entropy association 'persists after controlling for disease stage and visit count,' but §3.6 reports that after full adjustment including visit count, r=0.057, p=0.12, i.e., not significant. The paper's own text acknowledges 'partial confounding by observation density.' Thus the conclusion overstates the evidence. The pre-visit-count-adjusted association (r=-0.111, p=0.003 after controlling for conversion status and MCI subtype) is still worth reporting as hypothesis-generating, but the 'persists' language must be removed or qualified.
minor comments (5)
  1. [Abstract, §2.5] The method is cross-conformal using out-of-fold scores, but the abstract and introduction call it 'split-conformal.' Please align terminology throughout.
  2. [§3.5] The contradictory statements about formal vs. empirical coverage should be resolved; see major comment 1.
  3. [Figure 1, §3.1] It is unclear whether the internal 5-fold CV is performed on the full 741-subject cohort or on the ADNI-1 subset (n=330) used as the temporal training set. Table 2 reports N=741 while Figure 1 labels 'Internal — ADNI-1 n=330'. Please specify which cohort the internal results refer to and ensure consistency of all reported internal metrics.
  4. [§2.3] There is a LaTeX artifact in the text: 'extttclass_weight' should be rendered as class_weight='balanced'.
  5. [§3.6] The within-fixed-visit-count-strata null results are important; consider placing them more prominently so the reader does not miss that the APOE4 association is partly driven by observation density.

Circularity Check

2 steps flagged

Mild, disclosed feature-selection circularity is quantified by the paper; the central TDA biomarker claim is not definitionally circular. The conformal 'formal guarantee' is unsupported because OOF scores are not exactly exchangeable, which is a correctness issue rather than a definitional circularity.

specific steps
  1. other [Section 2.3 (Models and Evaluation); Section 4.5 (Limitations)]
    "To quantify this circularity, we implemented a nested selection protocol: in each outer fold, a 60% inner split was used to compute SHAP values and select the top-20 features, and the final model was trained on the full outer training set using those features. The nested selection AUC was 0.840 versus 0.866 for same-fold selection, yielding an optimism estimate of +0.026 AUC points."

    The top-20 feature set is selected using SHAP values derived from the same 5-fold CV subjects on which the final Stack+Top20 model is evaluated, so the 0.866 same-fold AUC is an optimistically selected quantity rather than an unbiased out-of-sample prediction. The paper explicitly labels this circularity and reports the nested-selection 0.840 as the primary unbiased estimate, making the circular step disclosed and bounded rather than the hidden driver of the conclusions.

  2. other [Section 3.5 (Survival Analysis and Conformal Guarantees); Methods 2.5]
    "Internal (formally valid): We implemented a cross-conformal procedure [15] in which each subject is scored by a model trained without it (5-fold CV), satisfying the exchangeability requirement within the training set. This yields 90.4% ± 2.2% coverage against the 90% target—a formally guaranteed result."

    The claimed formal finite-sample guarantee is derived from an exchangeability premise that the described OOF scoring procedure does not satisfy: each calibration score is produced by a model trained without that subject's fold, while a test score is produced by a model trained on all calibration subjects (or by averaging fold models). The vector of calibration scores plus the test score is therefore not exchangeable, so the quantile lemma behind Eq. (4) does not apply. The 90.4% figure is an empirical coverage observation relabeled as a theorem; this is primarily an unsupported correctness claim load-bearing for contribution C3, not a definitional circularity.

full rationale

The paper's central TDA claim is not circular by construction: H0 persistence entropy is measured from raw clinical trajectory point clouds via a Vietoris-Rips filtration and this does not use the conversion label as an input. The leakage audit and external temporal validation on a zero-overlap ADNI-2/GO/3 cohort provide independent checks. The main circularity in the paper is the SHAP-based feature selection performed on the same CV folds used for evaluation, which the authors transparently acknowledge and quantify as +0.026 AUC via a nested selection protocol; since they report the nested estimate (0.840) as the primary unbiased AUC, this is a partial, disclosed circularity rather than a forced result. No load-bearing self-citation chain is present: reference [15] is an external textbook, and the 'uniqueness' of the conformal guarantee is not imported from the authors' prior work. The conformal-coverage issue flagged above is important but is a correctness/mathematical-support problem, not a definitional equivalence between inputs and outputs. Overall, the derivation chain is largely self-contained, with one admitted bounded circularity and one unsupported formal-guarantee claim, justifying a moderate rather than high circularity score.

Axiom & Free-Parameter Ledger

5 free parameters · 4 axioms · 1 invented entities

The central claims rest on the choice of TDA summary statistics, the conformal exchangeability assumption, and the clinical assumption that ADNI diagnoses are not derived from CDRSB. The paper quantifies the optimism from feature selection but does not ship code or independently validate the biomarker.

free parameters (5)
  • Top-20 feature selection cutoff = 20
    Feature selection is performed on the training folds via SHAP; the top-20 cutoff is a modeling choice that introduces optimism (+0.026 AUC), quantified via nested selection.
  • Delay embedding parameters = m=2, tau=1
    Takens embedding dimension m=2 and delay tau=1 visit are chosen by hand due to sparse 2–10 visits; sensitivity analysis shows H1 features are near-chance.
  • Conformal alpha = 0.10
    Target miscoverage of 10% chosen by convention; not fitted to the data.
  • Uniform follow-up cap = 4 years
    Applied identically to both groups to fix informative censoring; window-sensitivity analysis across 2–5 years is reported, so the specific cap is a design choice.
  • RF class_weight='balanced' = balanced
    Random forest uses class balancing to handle 32.4% conversion prevalence; fixed a priori, not tuned in CV.
axioms (4)
  • domain assumption Vietoris-Rips persistent homology and Takens embedding are appropriate for these short, discrete clinical trajectories
    The paper applies TDA to 2–10 visits per subject, which the authors acknowledge is sparse; H1 loop detection is near-chance. Section 2.2.
  • domain assumption ADNI diagnoses are not algorithmically derived from CDRSB, so using CDRSB is not label leakage
    Section 2.2: 'CDRSB was retained because ADNI diagnoses are assigned by site-investigator clinical consensus, not derived from CDRSB algorithmically.'
  • ad hoc to paper Out-of-fold cross-validation scores satisfy exchangeability for split-conformal guarantees
    Section 3.5: 'we implemented a cross-conformal procedure ... satisfying the exchangeability requirement within the training set.' This is load-bearing for the 'formally valid' claim but is subtle for cross-conformal methods.
  • standard math Standard probability and statistical inference assumptions (DeLong test, bootstrap, log-rank)
    Used throughout for significance testing; assumed standard.
invented entities (1)
  • H0 persistence entropy as a topological biomarker no independent evidence
    purpose: Proposed as a measure of trajectory complexity; top SHAP feature and associated with APOE4 dosage
    The only evidence is from the same ADNI cohort and the same feature set; the authors call it hypothesis-generating and recommend external validation on NACC. The association partly attenuates after adjusting for visit count, so no independent falsifiable handle outside the paper.

pith-pipeline@v1.3.0-alltime-deepseek · 14084 in / 14604 out tokens · 128390 ms · 2026-08-01T17:56:15.883108+00:00 · methodology

0 comments
read the original abstract

Background. Predicting conversion from mild cognitive impairment (MCI) to Alzheimer's disease (AD) is central to trial enrichment and care planning, yet existing models provide no individual-level uncertainty estimates and rarely include transparent leakage audits. We introduce the first application of persistent homology to longitudinal clinical trajectory point clouds for this task, and the first split-conformal individual risk guarantee for any AD-conversion model. Methods. We analysed 741 MCI subjects (240 converters, 32.4%) from ADNI with a uniform 4-year follow-up cap. Five leakage sources were corrected; without them a naive pipeline achieved AUC=0.934, inflated by +0.075. Vietoris-Rips persistent homology and sublevel-set proxies were combined with trajectory slopes and engineered features (76 total) in a stacking ensemble evaluated by 5-fold cross-validation. Results. Cox and Random Survival Forest models with TDA features achieved concordance C=0.799 and C=0.826 versus C=0.753 and C=0.812 without (+0.045 and +0.014). The primary nested AUC is 0.840 (same-fold bound 0.866); external AUC was 0.879 on a zero-overlap ADNI-2/GO/3 cohort. H0 persistence entropy was the top SHAP feature and significantly associated with APOE4 dosage (Spearman r=-0.191, p<0.0001, Bonferroni-corrected). Cross-conformal coverage was 90.4%+-2.2% (target 90%); empirical external coverage 96.9%. Maximum fairness gap in false-negative rate across seven subgroups was 0.092. Conclusions. We propose H0 persistence entropy as a topological biomarker of cognitive decline and demonstrate that a leakage-audited, conformally calibrated pipeline reaches competitive accuracy with individual-level uncertainty quantification not previously available for this task.

Figures

Figures reproduced from arXiv: 2607.17442 by Navin Bondade.

Figure 1
Figure 1. Figure 1: TRIPOD+AI study flow. Of 878 ADNI MCI subjects screened, 741 met inclusion criteria. Training (ADNI-1, 2004–2005) and external test (ADNI-2/GO/3, 2010–2021) sets share zero subjects, constituting temporal external validation. 6 [PITH_FULL_IMAGE:figures/full_fig_p006_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Prediction performance. (a) Feature ablation from baseline-only to the full TDA pipeline; AUC increases incrementally with each feature group added. (b) Cross-validated AUC with 95 % boot￾strap CI (1,000 resamples); RF (AUC 0.859) recommended for deployment on calibration grounds; Stack+Top20 achieves highest discrimination (AUC 0.866). Stack+Top20 achieved AUC = 0.866 (95 % CI 0.840–0.890); RF reached AUC… view at source ↗
Figure 3
Figure 3. Figure 3: External validity and clinical utility. (a) [PITH_FULL_IMAGE:figures/full_fig_p009_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Survival and uncertainty quantification. (a) [PITH_FULL_IMAGE:figures/full_fig_p010_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Topological biomarker and biological validation. (a) [PITH_FULL_IMAGE:figures/full_fig_p012_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Clinical trajectory shapes and their persistence barcodes. (a,b) [PITH_FULL_IMAGE:figures/full_fig_p013_6.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

22 extracted references · 1 linked inside Pith

  1. [1]

    Mild cognitive impairment: clinical characteri- zation and outcome.Arch Neurol

    Petersen RC, Smith GE, Waring SC, et al. Mild cognitive impairment: clinical characteri- zation and outcome.Arch Neurol. 1999;56(3):303–308

  2. [2]

    Sparse learning and stability selection for predicting MCI to AD conversion.BMC Neurol

    Ye J, Wu T, Li J, Chen K. Sparse learning and stability selection for predicting MCI to AD conversion.BMC Neurol. 2012;12:46

  3. [3]

    Hierarchical fully convolutional network for joint atro- phy localization and Alzheimer’s disease diagnosis.IEEE Trans Pattern Anal Mach Intell

    Lian C, Liu M, Zhang J, Shen D. Hierarchical fully convolutional network for joint atro- phy localization and Alzheimer’s disease diagnosis.IEEE Trans Pattern Anal Mach Intell. 2020;42(4):880–893

  4. [4]

    Leakage and the reproducibility crisis in machine-learning-based science.Patterns

    Kapoor S, Narayanan A. Leakage and the reproducibility crisis in machine-learning-based science.Patterns. 2023;4(9):100804

  5. [5]

    An introduction to topological data analysis: fundamental and practical aspects for data scientists.Front Artif Intell

    Chazal F, Michel B. An introduction to topological data analysis: fundamental and practical aspects for data scientists.Front Artif Intell. 2021;4:667963

  6. [6]

    White matter brain network research in Alzheimer’s disease using persistent features.Molecules

    Kuang L, Han X, Chen K, Caselli RJ, Thompson PM, Wang Y. White matter brain network research in Alzheimer’s disease using persistent features.Molecules. 2020;25(11):2472

  7. [7]

    A spatiotemporal brain network analysis of Alzheimer’s disease based on persistent homology.Front Aging Neurosci

    Zhao X, Wu C, Cheng L, et al. A spatiotemporal brain network analysis of Alzheimer’s disease based on persistent homology.Front Aging Neurosci. 2022;13:779508

  8. [8]

    American Mathemat- ical Society; 2010

    Edelsbrunner H, Harer J.Computational Topology: An Introduction. American Mathemat- ical Society; 2010

  9. [9]

    The Alzheimer’s Disease Neuroimaging Initiative 3: continuedinnovationforclinicaltrialimprovement.Alzheimers Dement.2017;13(5):561–571

    Weiner MW, Veitch DP, Aisen PS, et al. The Alzheimer’s Disease Neuroimaging Initiative 3: continuedinnovationforclinicaltrialimprovement.Alzheimers Dement.2017;13(5):561–571

  10. [10]

    Ripser: efficient computation of Vietoris-Rips persistence barcodes.J Appl Comput Topol

    Bauer U. Ripser: efficient computation of Vietoris-Rips persistence barcodes.J Appl Comput Topol. 2021;5(3):391–423

  11. [11]

    Detecting strange attractors in turbulence

    Takens F. Detecting strange attractors in turbulence. In:Dynamical Systems and Turbu- lence, Lecture Notes in Mathematics vol898. Springer; 1981:366–381

  12. [12]

    Scikit-learn: machine learning in Python.J Mach Learn Res

    Pedregosa F, Varoquaux G, Gramfort A, et al. Scikit-learn: machine learning in Python.J Mach Learn Res. 2011;12:2825–2830

  13. [13]

    XGBoost: a scalable tree boosting system

    Chen T, Guestrin C. XGBoost: a scalable tree boosting system. In:Proc. 22nd ACM SIGKDD. 2016:785–794

  14. [14]

    A unified approach to interpreting model predictions.Adv Neural Inf Process Syst

    Lundberg SM, Lee SI. A unified approach to interpreting model predictions.Adv Neural Inf Process Syst. 2017;30

  15. [15]

    Vovk V, Gammerman A, Shafer G.Algorithmic Learning in a Random World. 2nd ed. Springer International Publishing; 2022. doi:10.1007/978-3-031-06649-8

  16. [16]

    Comparing the areas under two or more correlated receiver operating characteristic curves.Biometrics

    DeLong ER, DeLong DM, Clarke-Pearson DL. Comparing the areas under two or more correlated receiver operating characteristic curves.Biometrics. 1988;44(3):837–845

  17. [17]

    Hosmer DW, Lemeshow S, Sturdivant RX.Applied Logistic Regression. 3rd ed. Wiley; 2013

  18. [18]

    Evaluating the added predic- tive ability of a new marker: from area under the ROC curve to reclassification and beyond

    Pencina MJ, D’Agostino RB Sr, D’Agostino RB Jr, Vasan RS. Evaluating the added predic- tive ability of a new marker: from area under the ROC curve to reclassification and beyond. Stat Med. 2008;27(2):157–172

  19. [19]

    Decision curve analysis: a novel method for evaluating prediction models.Med Decis Making

    Vickers AJ, Elkin EB. Decision curve analysis: a novel method for evaluating prediction models.Med Decis Making. 2006;26(6):565–574

  20. [20]

    PROMISE-AD: progression- aware multi-horizon survival estimation for Alzheimer’s disease progression and dynamic 16 tracking.arXiv

    Lyu Q, Hudson J, Kawas M, Jiang Y, You C, Whitlow CT. PROMISE-AD: progression- aware multi-horizon survival estimation for Alzheimer’s disease progression and dynamic 16 tracking.arXiv. 2026;arXiv:2604.28055

  21. [21]

    An empirical characterization of fair machine learning for clinical risk prediction.J Biomed Inform

    Pfohl SR, Foryciarz A, Shah NH. An empirical characterization of fair machine learning for clinical risk prediction.J Biomed Inform. 2021;113:103621. doi:10.1016/j.jbi.2020.103621

  22. [22]

    TRIPOD+AI statement: updated guidance for reporting clinical prediction models.BMJ

    Collins GS, Moons KGM, Dhiman P, et al. TRIPOD+AI statement: updated guidance for reporting clinical prediction models.BMJ. 2024;385:e078378. 17