REVIEW 2 major objections 3 minor 41 references
When Is a Conformal Guarantee Fair? Auditing Silent Subgroup Under-Coverage in Alzheimer's Disease Longitudinal Prediction
T0 review · 2 major / 3 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read A marginal conformal guarantee can look valid on average while silently under-covering the high-risk subgroups it should protect; this paper finds the pattern in 57 of 68 audited combinations and traces it to rarity and tail-heaviness.
desk verdict Solid empirical audit with a real soft spot: the analysis counts visits, not patients, and the unexamined visit-level exchangeability assumption deserves a subject-level robustness check. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is a pair of coverage identities plus one repair rule. The small-cell coverage law gives the finite-sample coverage of a Mondrian (per-subgroup) conformal cell of size n as k/(n+1), where k = min(ceil((n+1)(1-$\alpha$)), n); this is the quantitative statement of the rarity mechanism and the reason a cell of n < ceil(1/$\alpha$)-1 is starved. The tail-driven marginal deficit gives the population-limit coverage of subgroup a under the marginal band as F_a(q_marg), with q_marg = $F^{{-1}}$(1-$\alpha$); the subgroup under-covers exactly when its own quantile q_a exceeds q_marg, and the shortfall is independent of calibration size. The repair is the three-step recipe applied in order: condition on the sensitive attribute to remove the tail deficit, pool out-of-fold scores across subject-disjoint folds (cross-conformal pooling) to enlarge each stratum cell, and apply the marginal floor q*_a = max(qhat_a, qhat_marg) so that no thin cell falls below pooled marginal coverage. The order matters because conditioning creates the rarity that pooling then removes.
What would settle it
Recompute every coverage statistic at the subject level, taking one visit per subject at each horizon, and check whether the 57-of-68 under-coverage count survives; if the deficit mostly vanishes, the reported silent under-coverage is an artifact of pooling multiple visits per patient rather than a property of the conformal method.
Extended reading notes
Core claim
The central claim is that a conformal band can satisfy the nominal marginal coverage guarantee while systematically failing the high-risk subgroups it is meant to protect. The paper's specific discovery is that this silent under-coverage has two distinct and characterizable causes. Rarity (Proposition 1) says a group-conditional cell calibrated on n exchangeable scores covers a fresh point with probability k/(n+1) with k = min(ceil((n+1)(1-$\alpha$)), n), so small cells fall below nominal. Tail-heaviness (Proposition 2) says a population-wide quantile q_marg = $F^{{-1}}$(1-$\alpha$) covers subgroup a exactly with probability F_a(q_marg), which is below 1-$\alpha$ precisely when the subgroup's own quantile exceeds the marginal one; this deficit is distributional and persists at any sample size. The two mechanisms worsen at opposite tolerance levels, and the paper's mechanism-matched repair — per-subgroup conditioning, leakage-safe cross-conformal pooling, and a one-sided marginal floor — restores high-risk subgroup coverage to at least nominal across both cohorts and both forecasters, except for cells so rare that fewer than about 1/$\alpha$ calibration patients exist.
Load-bearing premise
The audit and repair treat each patient-visit as an exchangeable unit and pool coverage counts over horizons, so the finite-sample guarantees and the measured under-coverage rates could change if visits from the same patient are correlated rather than exchangeable.
Editorial extensions
If this is right
- Any pipeline that reports only marginal conformal coverage cannot certify coverage for any named high-risk subgroup; the audit shows the marginal rate can be at nominal while the subgroups a clinician actually faces are under-covered.
- The rarity and tail mechanisms move in opposite directions with the tolerance level, so no single alpha adjustment fixes both; a repair must combine conditioning with pooling.
- The repair is a post-hoc quantile computation that leaves the base forecaster untouched, preserves marginal coverage by construction, and transfers across very different forecasters and cohorts.
- For subgroup intersections that are simultaneously extremely rare and heavy-tailed (on the order of a dozen calibration patients), no post-hoc quantile method can certify 1-alpha coverage, and the paper treats this as a data-limited frontier.
- Across the audited axes the recipe raises subgroup coverage by a mean 4.3 percentage points (6.7 on clinical-risk axes) relative to the marginal band, and it does so by widening bands only where the tail demands it rather than uniformly.
Reading between the lines
- A subject-level re-audit (one visit per subject per horizon) would separate patient-level risk from visit-multiplicity effects; if the 57-of-68 under-coverage count changes sharply, the headline number is partly an artifact of the paper's visit-level pooling.
- The same rarity/tail dichotomy should appear wherever marginal conformal bands are applied to populations with identifiable risk strata, such as rare-disease screening or portfolio tail-risk forecasting, so the audit protocol transfers beyond Alzheimer's disease.
- The marginal floor acts as a no-regret safety net relative to pooled split conformal, which suggests the recipe can be run online, updating pooled quantiles and per-subgroup quantiles as new visits arrive while keeping the floor.
- The two propositions could be stress-tested in fully synthetic exchangeable data with known heavy tails and small cells, which would tell a reader whether the residual-based Figure 2 confirms the law or is driven by cohort artifacts.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper develops a mechanism-driven framework for auditing and repairing subgroup under-coverage in conformal prediction bands for longitudinal Alzheimer's disease biomarker forecasts. Using two cohorts (ADNI, OASIS-3), two base forecasters, and nine a priori defined attributes, the authors report that standard population-level conformal bands under-cover high-risk subgroups in 57 of 68 audited (subgroup, model, alpha) combinations despite nominal marginal coverage. They attribute the failures to two mechanisms: rarity (small-cell coverage bounded by k/(n+1)) and tail-heaviness (population quantile too narrow for heavy-tailed subgroups), and propose a repair combining per-subgroup conditioning, cross-conformal pooling, and a marginal floor. The paper reports that the repair restores target coverage for nearly all high-risk subgroups and compares favorably with a conditional-conformal baseline.
Significance. If the empirical findings hold, the paper makes a useful practical contribution by providing a simple, model-agnostic diagnostic and repair for a clinically important failure mode of conformal prediction. The study benefits from pre-specified high-risk subgroups, subject-disjoint cross-validation, external validation on a second cohort, and a baseline comparison with a state-of-the-art conditional-conformal method. The theoretical propositions are standard but clearly stated, and the mechanisms are validated empirically. The main reservation concerns the statistical unit of analysis, which is load-bearing for the headline claims.
major comments (2)
- [Section 3 and Section 6] The audit and repair treat each patient-visit as an exchangeable conformal unit, with coverage count-pooled over horizons h in {12,24,48} (Section 3) and cross-conformal pooling described as lifting per-(stratum, horizon) cells to about 78 joint visits (Section 6). Because a subject can contribute up to three scores that are likely correlated (e.g., a fast progressor has large residuals at all horizons), the multiset of visit scores is not exchangeable under arbitrary permutations, and the finite-sample guarantees cited (Propositions 1 and 2; the 1-2α-O(K^{-1}) bound in Section 6) may not hold at the visit level. The paper never justifies visit-level exchangeability or reports a subject-level analysis. This matters because the central claims (57/68 under-covered combinations, mean +4.3 pp and +6.7 pp repair gains) are computed on patient-visit units; if high-risk subgroups have different numbers of follow-up visits, count-pooling can over-represent certain patients, so the measured under-coverage could be an artifact of pooling rather than a per-patient phenomenon. Please provide a subject-level analysis (e.g., one randomly selected horizon per subject, or cluster-robust inference) to confirm the findings, or explicitly reframe the claims as per-visit coverage with appropriate caveats.
- [Section 4 and Table 1] The abstract and Section 4 state that 'population-level bands under-cover high-risk subgroups in 57 of 68 audited combinations,' but Table 1 (JointFlow) shows M-CP alone under-covers only 22 of 34 combinations when counting any deficit relative to nominal; the 57/68 figure appears to be the union of M-CP and Mondrian under-coverage. Moreover, the table's underline convention uses >1 pp below nominal, while the 57/68 count seems to use any deficit. Please clarify which standard recipe the count refers to, align the abstract with the reported numbers, and state the threshold used for 'under-covered' consistently.
minor comments (3)
- [Section 4] The phrase 'standard recipe' is used in the main text without a precise definition; it could mean M-CP alone, Mondrian alone, or any of the standard methods. Please define it explicitly and use consistent terminology.
- [Section 5, Proposition 2] The supplement-derived bound that the tail deficit is at most α(1-π_a)/π_a for a subgroup with population share π_a is mentioned only in passing; stating it in the main text would strengthen the reader's intuition about why large subgroups cannot exhibit large tail deficits.
- [Figure 2 and Section 5] Figure 2B reports a correlation of r=-0.83 over 90 cells, but the text does not state how many cells come from each cohort or base model, nor whether the correlation is weighted by cell size. Adding this information would help assess the robustness of the tail-heaviness mechanism.
Circularity Check
No significant circularity: audit subgroups are fixed a priori, folds are subject-disjoint, and repair parameters are not tuned on test outcomes.
full rationale
The central audit and repair chain is not circular. High-risk subgroups are fixed before coverage is examined ('For each of the eight attributes with an a-priori risk direction we fix the highest-risk subgroup before looking at coverage', Section 3), and the calibration/test split is subject-disjoint with no subject appearing in two splits. The two mechanisms are formalized through Proposition 1, which explicitly restates the standard conformal coverage identity, and Proposition 2, which is a direct CDF quantile identity with a proof relegated to the supplement. Neither proposition fits a parameter to the test outcomes it is later used to explain. The repair ingredients (Mondrian conditioning, cross-conformal pooling, and the marginal floor q̂*_a = max(q̂_a, q̂_marg)) are post-hoc quantile operations with no tuning on test coverage, and the cross-conformal guarantee is cited from external, parameter-free results (Vovk 2015; Barber et al. 2021b). Empirical claims such as '57 of 68 audited combinations' and the aggregate +4.3 pp improvement are evaluated on held-out test folds against fixed rules, so they are not forced by construction. The paper's unexamined visit-level exchangeability assumption under count-pooling over horizons is a validity or robustness concern, not a circularity, because no predicted quantity is definitionally equal to a fitted input. No self-citation is load-bearing, and no known result is merely renamed as a new contribution.
Assumptions & free parameters
free parameters (2)
- kappa (score-scale floor) =
0.25
- under-coverage threshold =
1 percentage point
assumptions (4)
- domain assumption Visit-level exchangeability: each patient-visit score is exchangeable with all other visit scores in the calibration set for the split and cross-conformal guarantees
- domain assumption A priori high-risk subgroup labels (APOE4+, e4/e4, AD, CDR-SB>4, MMSE<24, age>=80, <=12y education, Non-White; Male by convention)
- standard math Regularity of subgroup upper-tail density f_a non-increasing on [q_marg, q_a] in Proposition 2
- domain assumption The base model's predictive interval provides a valid per-channel scale for residual normalization
Cite this review
Pith. "Pith review of When Is a Conformal Guarantee Fair? Auditing Silent Subgroup Under-Coverage in Alzheimer's Disease Longitudinal Prediction." pith.science (2026). https://pith.science/paper/3PBTULBD
@misc{pith2026260804254,
author = {Pith},
title = {Pith review of: When Is a Conformal Guarantee Fair? Auditing Silent Subgroup Under-Coverage in Alzheimer's Disease Longitudinal Prediction},
year = {2026},
howpublished = {\url{https://pith.science/paper/3PBTULBD}},
note = {Machine review of arXiv:2608.04254}
}
abstract
Longitudinal prediction of Alzheimer's disease biomarkers increasingly informs clinical decisions, and a forecast is only useful if it also reports how much to trust it. Conformal prediction supplies this by wrapping any forecaster in a prediction band with a finite-sample coverage guarantee under exchangeability. However, standard population-level conformal prediction guarantees only marginal coverage and may mask substantial under-coverage within clinically important subgroups. We introduce a general mechanism-driven framework for auditing and repairing such subgroup under-coverage. Across two cohorts (ADNI, OASIS-3), two base forecasters, and nine attributes spanning genetic risk, demographics, and clinical severity, we find that population-level bands under-cover high-risk subgroups in 57 of 68 audited combinations, despite achieving nominal marginal coverage. We trace these failures to two mechanisms: (A) \emph{rarity}, where a group-conditional band calibrated on only $n$ patients covers at most $k/(n+1)$; and (B) \emph{tail-heaviness}, where a population-wide band is too narrow for a heavy-tailed subgroup and additional data cannot close the gap. Under-coverage falls disproportionately on patients with high genetic risk and disease severity (6.1 pp mean deficit, 95\% CI [3.3, 8.9]), while demographic groups remain at the target level on average (0.0 pp, CI [$-1.9$, 1.7]). We pair each mechanism with a corresponding conformal correction: cross-conformal pooling for rarity, per-subgroup calibration for tail-heaviness, and a coverage-safe marginal floor when both arise. Together, these corrections restore target coverage for nearly every high-risk subgroup across both cohorts and forecasters.
Figures
Figures from the paper (1 more)
Reference graph
Works this paper leans on
-
[1]
2005 , publisher =
Algorithmic Learning in a Random World , author =. 2005 , publisher =
2005
-
[2]
Journal of the American Statistical Association , volume =
Distribution-free predictive inference for regression , author =. Journal of the American Statistical Association , volume =
-
[3]
arXiv preprint arXiv:2107.07511 , year =
A Gentle Introduction to Conformal Prediction and Distribution-Free Uncertainty Quantification , author =. arXiv preprint arXiv:2107.07511 , year =
-
[4]
Advances in Neural Information Processing Systems (NeurIPS) , year =
Conformalized quantile regression , author =. Advances in Neural Information Processing Systems (NeurIPS) , year =
-
[5]
Advances in Neural Information Processing Systems (NeurIPS) , volume =
Adaptive Conformal Inference Under Distribution Shift , author =. Advances in Neural Information Processing Systems (NeurIPS) , volume =
-
[6]
Advances in Neural Information Processing Systems (NeurIPS) , year =
Uncertainty-Calibrated Prediction of Randomly-Timed Biomarker Trajectories with Conformal Bands , author =. Advances in Neural Information Processing Systems (NeurIPS) , year =
-
[7]
and Duvenaud, David , booktitle =
Rubanova, Yulia and Chen, Ricky T.Q. and Duvenaud, David , booktitle =. Latent
-
[8]
Pattern Recognition , volume =
Copula-based conformal prediction for multi-target regression , author =. Pattern Recognition , volume =. 2021 , doi =
work page 2021
Show all 41 references
-
[9]
International Conference on Learning Representations (ICLR) , year =
Copula Conformal prediction for multi-step time series prediction , author =. International Conference on Learning Representations (ICLR) , year =
-
[10]
Jack, Clifford R and Bennett, David A and Blennow, Kaj and Carrillo, Maria C and Dunn, Billy and Haeberlein, Samantha Budd and Holtzman, David M and Jagust, William and others , journal =
-
[11]
Learning Spatio-Temporal Model of Disease Progression with
Lachinov, Dmitrii and Chakravarty, Arunava and Grechenig, Christoph and Schmidt-Erfurth, Ursula and Bogunovi. Learning Spatio-Temporal Model of Disease Progression with. IEEE Transactions on Medical Imaging , volume =. 2024 , doi =
2024
-
[12]
Deep Kernel Learning with Temporal Gaussian Processes for Clinical Variable Prediction in
Tassopoulou, Vasiliki and Yu, Fanyang and Davatzikos, Christos , booktitle =. Deep Kernel Learning with Temporal Gaussian Processes for Clinical Variable Prediction in
-
[13]
Jung, Wooseok and Park, Joonhyuk and Kim, Won Hwa , booktitle =
-
[14]
Wen, Zheyu and Biros, George , booktitle =
-
[15]
Advances in Neural Information Processing Systems (NeurIPS) , year =
Park, Wonjung and Ahn, Suhyun and Vald. Advances in Neural Information Processing Systems (NeurIPS) , year =
-
[16]
Inducing Clusters Deep Kernel
Liang, Junjie and Ren, Weijieying and Sahar, Hanifi and Honavar, Vasant , booktitle =. Inducing Clusters Deep Kernel. 2024 , doi =
2024
-
[17]
AAAI Conference on Artificial Intelligence (AAAI) , volume =
Probabilistic Forecasting of Irregularly Sampled Time Series with Missing Values via Conditional Normalizing Flows , author =. AAAI Conference on Artificial Intelligence (AAAI) , volume =. 2025 , doi =
2025
-
[18]
International Conference on Machine Learning (ICML) , volume =
A Unified Comparative Study with Generalized Conformity Scores for Multi-Output Conformal Regression , author =. International Conference on Machine Learning (ICML) , volume =. 2025 , note =
2025
-
[19]
Transactions on Machine Learning Research , year =
Minimax Multi-Target Conformal Prediction with Applications to Imaging Inverse Problems , author =. Transactions on Machine Learning Research , year =
-
[20]
International Conference on Learning Representations (ICLR) , year =
Batch Multivalid Conformal Prediction , author =. International Conference on Learning Representations (ICLR) , year =
-
[21]
Advances in Neural Information Processing Systems (NeurIPS) , year =
Class-Conditional Conformal Prediction with Many Classes , author =. Advances in Neural Information Processing Systems (NeurIPS) , year =
-
[22]
International Conference on Machine Learning (ICML) , volume =
Kandinsky Conformal Prediction: Beyond Class- and Covariate-Conditional Coverage , author =. International Conference on Machine Learning (ICML) , volume =
-
[23]
arXiv preprint arXiv:2508.13288 , year =
Hierarchical Conformal Classification , author =. arXiv preprint arXiv:2508.13288 , year =
-
[24]
Density Estimation Using
Dinh, Laurent and Sohl-Dickstein, Jascha and Bengio, Samy , booktitle =. Density Estimation Using
-
[25]
2003 , url=
Mondrian Confidence Machine , author=. 2003 , url=
2003
-
[26]
Annals of Mathematics and Artificial Intelligence , volume=
Cross-conformal predictors , author=. Annals of Mathematics and Artificial Intelligence , volume=. 2015 , doi=
2015
-
[27]
The Annals of Statistics , volume=
Predictive inference with the jackknife+ , author=. The Annals of Statistics , volume=
-
[28]
and Benzinger, Tammie L.S
LaMontagne, Pamela J. and Benzinger, Tammie L.S. and Morris, John C. and Keefe, Sarah and Hornbeck, Russ and Xiong, Chengjie and Grant, Elizabeth and Hassenstab, Jason and Moulder, Krista and Vlassenko, Andrei G. and others , journal =. 2019 , doi =
2019
-
[29]
Biometrika , volume =
Localized conformal prediction: a generalized inference framework for conformal prediction , author =. Biometrika , volume =. 2023 , doi =
2023
-
[30]
Journal of the Royal Statistical Society Series B: Statistical Methodology , volume =
Conformal prediction with local weights: randomization enables robust guarantees , author =. Journal of the Royal Statistical Society Series B: Statistical Methodology , volume =. 2025 , doi =
2025
-
[31]
and Nowozin, Sebastian and Dillon, Joshua V
Ovadia, Yaniv and Fertig, Emily and Ren, Jie and Nado, Zachary and Sculley, D. and Nowozin, Sebastian and Dillon, Joshua V. and Lakshminarayanan, Balaji and Snoek, Jasper , booktitle =. Can You Trust Your Model's Uncertainty?
-
[32]
Biometrics , volume =
Random-Effects Models for Longitudinal Data , author =. Biometrics , volume =. 1982 , doi =
1982
-
[33]
Asian Conference on Machine Learning (ACML) , pages=
Conditional validity of inductive conformal predictors , author=. Asian Conference on Machine Learning (ACML) , pages=
-
[34]
Information and Inference: A Journal of the IMA , volume=
The limits of distribution-free conditional predictive inference , author=. Information and Inference: A Journal of the IMA , volume=
-
[35]
Harvard Data Science Review , volume=
With malice toward none: Assessing uncertainty via equalized coverage , author=. Harvard Data Science Review , volume=
-
[36]
Advances in Neural Information Processing Systems (NeurIPS) , volume=
Practical adversarial multivalid conformal prediction , author=. Advances in Neural Information Processing Systems (NeurIPS) , volume=
-
[37]
Journal of the Royal Statistical Society Series B: Statistical Methodology , volume=
Conformal prediction with conditional guarantees , author=. Journal of the Royal Statistical Society Series B: Statistical Methodology , volume=. 2025 , publisher=
2025
-
[38]
Alzheimer's & Dementia , volume =
Estimating long-term multivariate progression from short-term data , author =. Alzheimer's & Dementia , volume =
-
[39]
Nature Reviews Neuroscience , volume =
Data-driven modelling of neurodegenerative disease progression: thinking outside the black box , author =. Nature Reviews Neuroscience , volume =. 2024 , doi =
2024
-
[40]
International Conference on Machine Learning (ICML) , pages =
On Calibration of Modern Neural Networks , author =. International Conference on Machine Learning (ICML) , pages =
-
[41]
Alzheimer's Disease Neuroimaging Initiative (
Petersen, Ronald C and Aisen, Paul S and Beckett, Laurel A and Donohue, Michael C and Gamst, Anthony C and Harvey, Danielle J and Jack, Clifford R and Jagust, William J and Shaw, Leslie M and Toga, Arthur W and Trojanowski, John Q and Weiner, Michael W , journal =. Alzheimer's...
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.