{"id":"07d68180-c2bc-4ce6-92d5-ae4604980b94","arxiv_id":"2608.04254","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"low","formal_verification":"none","parameter_count":2,"one_line_summary":"Marginal conformal bands under-cover high-risk Alzheimer's subgroups even at nominal average coverage, and rarity and tail-heaviness explain and repair the deficit.","lead":"Population-level conformal prediction bands for Alzheimer's biomarker forecasts look reliable on average but silently under-cover high-risk patients, failing in 57 of 68 audited group-model-tolerance combinations. The paper traces the failure to small subgroup samples and heavy-tailed subgroup risk, and shows a three-step conformal repair that restores coverage for nearly all high-risk groups in two cohorts.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Visit-level exchangeability is assumed but not justified; count-pooling over horizons may drive the reported under-coverage and repair results rather than subject-level performance.","rationale":"The reader's weakest assumption identifies the same core issue. I agree: the paper's audit and repair are built on visit-level counting without establishing exchangeability. This is load-bearing because both the descriptive finding (57/68) and the repair's success (mean +4.3 pp, restoration to nominal) are measured on this pooled metric. The paper's own guarantees hinge on exchangeability of the scores, which is not established at the visit level. A subject-level analysis is the natural and inexpensive check. I also note a related validity concern: the cross-conformal pooling described in Section 6 may be asymmetric—the test fold's score is produced by a model trained on all other folds, while calibration scores from fold j are produced by models that include the test fold in their training—so the cited CV+/jackknife+ bound may not apply as stated. But that concern affects the theoretical guarantee, whereas the subject-level check directly tests the central empirical claim. If the subject-level analysis confirms the pattern, the paper is strengthened; if not, the headline finding is an artifact. Hence I do not change the verdict: the paper remains conditionally acceptable pending this analysis.","tokens_in":14693,"tokens_out":23458,"duration_ms":203162,"concrete_test":"Recompute the audit (Table 1 deficits, 57/68 count) and the repair (recipe coverage gains, Tables 2 and 4) using a subject-level test unit: for each test subject, either (a) randomly select one of the available horizons, or (b) define subject coverage as requiring all retained horizons covered, keeping the same 10 folds and calibration pooling. If the under-coverage counts or the repair gains change materially (e.g., the 57/68 drops or the mean +4.3 pp moves outside its CI), the central claim depends on visit-level pooling rather than subject-level performance.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's empirical claims—57/68 under-covered combinations and the repair's improvements (mean +4.3 pp, +6.7 pp for clinical axes)—are computed on patient-visit units: coverage is count-pooled over h∈{12,24,48} (Section 3), and the cross-conformal calibration pool in Section 6 is described as lifting per-(stratum, horizon) cells from ~17 to ~78 'joint visits.' A subject can contribute up to three scores to both calibration and test. The conformal guarantees cited (Propositions 1 and 2, and the 1−2α−O(K^{−1}) bound in Section 6) require exchangeability of the scores. Visit-level scores from the same subject are generally correlated (e.g., a fast progressor has large residuals at all horizons), so the multiset of all visit scores is not exchangeable under arbitrary permutations. The paper never justifies visit-level exchangeability or reports a subject-level analysis. If high-risk subgroups have different numbers of follow-up visits, count-pooling can over-represent certain patients, so the measured under-coverage could be an artifact of pooling rather than a per-patient phenomenon. This matters because the central claim is about silent under-coverage for patients in high-risk subgroups. A subject-level analysis is needed to confirm that the audit's 57/68 and the repair's gains are not an artifact of visit counting.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops a mechanism-driven framework for auditing and repairing subgroup under-coverage in conformal prediction bands for longitudinal Alzheimer's disease biomarker forecasts. Using two cohorts (ADNI, OASIS-3), two base forecasters, and nine a priori defined attributes, the authors report that standard population-level conformal bands under-cover high-risk subgroups in 57 of 68 audited (subgroup, model, alpha) combinations despite nominal marginal coverage. They attribute the failures to two mechanisms: rarity (small-cell coverage bounded by k/(n+1)) and tail-heaviness (population quantile too narrow for heavy-tailed subgroups), and propose a repair combining per-subgroup conditioning, cross-conformal pooling, and a marginal floor. The paper reports that the repair restores target coverage for nearly all high-risk subgroups and compares favorably with a conditional-conformal baseline.","tokens_in":14921,"tokens_out":12176,"duration_ms":98667,"significance":"If the empirical findings hold, the paper makes a useful practical contribution by providing a simple, model-agnostic diagnostic and repair for a clinically important failure mode of conformal prediction. The study benefits from pre-specified high-risk subgroups, subject-disjoint cross-validation, external validation on a second cohort, and a baseline comparison with a state-of-the-art conditional-conformal method. The theoretical propositions are standard but clearly stated, and the mechanisms are validated empirically. The main reservation concerns the statistical unit of analysis, which is load-bearing for the headline claims.","major_comments":[{"comment":"The audit and repair treat each patient-visit as an exchangeable conformal unit, with coverage count-pooled over horizons h in {12,24,48} (Section 3) and cross-conformal pooling described as lifting per-(stratum, horizon) cells to about 78 joint visits (Section 6). Because a subject can contribute up to three scores that are likely correlated (e.g., a fast progressor has large residuals at all horizons), the multiset of visit scores is not exchangeable under arbitrary permutations, and the finite-sample guarantees cited (Propositions 1 and 2; the 1-2α-O(K^{-1}) bound in Section 6) may not hold at the visit level. The paper never justifies visit-level exchangeability or reports a subject-level analysis. This matters because the central claims (57/68 under-covered combinations, mean +4.3 pp and +6.7 pp repair gains) are computed on patient-visit units; if high-risk subgroups have different numbers of follow-up visits, count-pooling can over-represent certain patients, so the measured under-coverage could be an artifact of pooling rather than a per-patient phenomenon. Please provide a subject-level analysis (e.g., one randomly selected horizon per subject, or cluster-robust inference) to confirm the findings, or explicitly reframe the claims as per-visit coverage with appropriate caveats.","section":"Section 3 and Section 6"},{"comment":"The abstract and Section 4 state that 'population-level bands under-cover high-risk subgroups in 57 of 68 audited combinations,' but Table 1 (JointFlow) shows M-CP alone under-covers only 22 of 34 combinations when counting any deficit relative to nominal; the 57/68 figure appears to be the union of M-CP and Mondrian under-coverage. Moreover, the table's underline convention uses >1 pp below nominal, while the 57/68 count seems to use any deficit. Please clarify which standard recipe the count refers to, align the abstract with the reported numbers, and state the threshold used for 'under-covered' consistently.","section":"Section 4 and Table 1"}],"minor_comments":[{"comment":"The phrase 'standard recipe' is used in the main text without a precise definition; it could mean M-CP alone, Mondrian alone, or any of the standard methods. Please define it explicitly and use consistent terminology.","section":"Section 4"},{"comment":"The supplement-derived bound that the tail deficit is at most α(1-π_a)/π_a for a subgroup with population share π_a is mentioned only in passing; stating it in the main text would strengthen the reader's intuition about why large subgroups cannot exhibit large tail deficits.","section":"Section 5, Proposition 2"},{"comment":"Figure 2B reports a correlation of r=-0.83 over 90 cells, but the text does not state how many cells come from each cohort or base model, nor whether the correlation is weighted by cell size. Adding this information would help assess the robustness of the tail-heaviness mechanism.","section":"Figure 2 and Section 5"}],"recommendation":"major_revision","confidential_remarks":"This is a well-written paper with a clinically relevant contribution. The main technical concern is the unexamined visit-level exchangeability assumption, which affects the validity of the theoretical guarantees and the interpretation of the empirical counts. The 57/68 headline should also be reconciled with Table 1. These issues are addressable with additional analysis and should be resolved before acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a solid, honest paper that demonstrates a real failure mode—marginal conformal bands under-cover clinically high-risk subgroups in AD forecasting—and offers a simple, sensible repair. The theory is not new (they say so themselves: Prop 1 restates the standard finite-sample identity, Prop 2 formalizes the known conditional coverage gap), but the audit across two cohorts and two very different forecasters is well executed. Subgroups are pre-specified before looking at coverage, folds are subject-disjoint, and the ingredient-wise ablation in Table 2 convincingly isolates the two mechanisms. The comparison with Gibbs–Candes is a nice touch and shows the recipe gets similar coverage with less width.\n\nThe soft spot that matters: the analysis treats patient-visits as exchangeable conformal units. Coverage is count-pooled over the 12/24/48-month horizons, so a subject can contribute up to three scores to both calibration and test. Scores from the same subject are typically correlated—a fast progressor has large residuals at all horizons—so the multiset of visit scores is not exchangeable under arbitrary permutations, and the finite-sample guarantees they cite don't strictly apply. The headline 57/68 under-covered rate and the repair's mean gain are per-visit numbers, not per-patient numbers. The paper never justifies visit-level exchangeability and gives no subject-level analysis. This is addressable: re-run the audit with one randomly selected visit per subject, or with a cluster-robust coverage estimate. My guess is the qualitative conclusions survive—the tail mechanism is distributional and the rarity mechanism is about cell size—but the magnitudes could shift, and the paper's language blurs 'visits' and 'patients' throughout. That is the main thing I'd want changed before trusting the headline numbers.\n\nMinor: the technical supplement (Prop 2 proof, Algorithm 1, Latent-ODE tables) wasn't available, so I couldn't verify those; no code or data are shipped; and the multivalid baseline (Bastani et al.) is mentioned in related work but not compared. None of these change the central finding.\n\nWho this is for: anyone building conformal predictors for medical or longitudinal settings where subgroup fairness matters. It deserves a serious referee; with a subject-level robustness check and some tightening of the unit-of-analysis language, I'd be happy to see it published.","headline":"Solid empirical audit with a real soft spot: the analysis counts visits, not patients, and the unexamined visit-level exchangeability assumption deserves a subject-level robustness check.","tokens_in":15467,"tokens_out":2732,"would_cite":true,"duration_ms":25604,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62G15","62P10"],"pacs":[],"model":"deepseek-v4-flash","headline":"A marginal conformal guarantee can look valid on average while silently under-covering the high-risk subgroups it should protect; this paper finds the pattern in 57 of 68 audited combinations and traces it to rarity and tail-heaviness.","keywords":["conformal prediction","subgroup coverage","Alzheimer's disease","longitudinal prediction","cross-conformal pooling","Mondrian conformal","tail-heaviness","prediction bands"],"falsifier":"Recompute every coverage statistic at the subject level, taking one visit per subject at each horizon, and check whether the 57-of-68 under-coverage count survives; if the deficit mostly vanishes, the reported silent under-coverage is an artifact of pooling multiple visits per patient rather than a property of the conformal method.","tokens_in":14471,"feed_emoji":"🧠","tokens_out":17454,"duration_ms":120777,"temperature":0.7,"pith_summary":"Population-level conformal prediction (a distribution-free way to turn a forecaster's residuals into prediction bands with finite-sample coverage guarantees) guarantees coverage only on average, and this paper shows that in longitudinal Alzheimer's biomarker forecasting that average can hide systematic under-coverage of the patients who matter most. Auditing high-risk subgroups across two Alzheimer's cohorts, two base forecasters, and nine clinical and demographic attributes, the paper finds that marginal prediction bands under-cover high-risk subgroups in 57 of 68 audited combinations while marginal coverage stays at the nominal level. The deficits concentrate on clinical-risk groups such as patients with high genetic risk and more severe disease, with a mean 6.1 percentage-point shortfall, whereas demographic groups on average sit at the target. The paper attributes the failures to two mechanisms, rarity (too few calibration patients in a subgroup, capping coverage at k/(n+1)) and tail-heaviness (a heavy-tailed subgroup needing a wider quantile than the population band provides), and shows that a matched repair of pooling, per-subgroup conditioning, and a marginal floor restores coverage for nearly all high-risk subgroups. If the paper is right, any marginal conformal band used in a setting with known high-risk subgroups needs an explicit subgroup-coverage audit before it can be called clinically fair.","feed_headline":"57 of 68 audited: marginal bands under-cover high-risk AD subgroups","feed_subtitle":"Nominal average coverage hides real deficits for high-risk groups; the paper's three-part conformal repair fixes it.","key_machinery":"The machinery is a pair of coverage identities plus one repair rule. The small-cell coverage law gives the finite-sample coverage of a Mondrian (per-subgroup) conformal cell of size n as k/(n+1), where k = min(ceil((n+1)(1-$\\alpha$)), n); this is the quantitative statement of the rarity mechanism and the reason a cell of n < ceil(1/$\\alpha$)-1 is starved. The tail-driven marginal deficit gives the population-limit coverage of subgroup a under the marginal band as F_a(q_marg), with q_marg = $F^{{-1}}$(1-$\\alpha$); the subgroup under-covers exactly when its own quantile q_a exceeds q_marg, and the shortfall is independent of calibration size. The repair is the three-step recipe applied in order: condition on the sensitive attribute to remove the tail deficit, pool out-of-fold scores across subject-disjoint folds (cross-conformal pooling) to enlarge each stratum cell, and apply the marginal floor q*_a = max(qhat_a, qhat_marg) so that no thin cell falls below pooled marginal coverage. The order matters because conditioning creates the rarity that pooling then removes.","core_discovery":"The central claim is that a conformal band can satisfy the nominal marginal coverage guarantee while systematically failing the high-risk subgroups it is meant to protect. The paper's specific discovery is that this silent under-coverage has two distinct and characterizable causes. Rarity (Proposition 1) says a group-conditional cell calibrated on n exchangeable scores covers a fresh point with probability k/(n+1) with k = min(ceil((n+1)(1-$\\alpha$)), n), so small cells fall below nominal. Tail-heaviness (Proposition 2) says a population-wide quantile q_marg = $F^{{-1}}$(1-$\\alpha$) covers subgroup a exactly with probability F_a(q_marg), which is below 1-$\\alpha$ precisely when the subgroup's own quantile exceeds the marginal one; this deficit is distributional and persists at any sample size. The two mechanisms worsen at opposite tolerance levels, and the paper's mechanism-matched repair — per-subgroup conditioning, leakage-safe cross-conformal pooling, and a one-sided marginal floor — restores high-risk subgroup coverage to at least nominal across both cohorts and both forecasters, except for cells so rare that fewer than about 1/$\\alpha$ calibration patients exist.","pith_inferences":["A subject-level re-audit (one visit per subject per horizon) would separate patient-level risk from visit-multiplicity effects; if the 57-of-68 under-coverage count changes sharply, the headline number is partly an artifact of the paper's visit-level pooling.","The same rarity/tail dichotomy should appear wherever marginal conformal bands are applied to populations with identifiable risk strata, such as rare-disease screening or portfolio tail-risk forecasting, so the audit protocol transfers beyond Alzheimer's disease.","The marginal floor acts as a no-regret safety net relative to pooled split conformal, which suggests the recipe can be run online, updating pooled quantiles and per-subgroup quantiles as new visits arrive while keeping the floor.","The two propositions could be stress-tested in fully synthetic exchangeable data with known heavy tails and small cells, which would tell a reader whether the residual-based Figure 2 confirms the law or is driven by cohort artifacts."],"forward_implications":["Any pipeline that reports only marginal conformal coverage cannot certify coverage for any named high-risk subgroup; the audit shows the marginal rate can be at nominal while the subgroups a clinician actually faces are under-covered.","The rarity and tail mechanisms move in opposite directions with the tolerance level, so no single alpha adjustment fixes both; a repair must combine conditioning with pooling.","The repair is a post-hoc quantile computation that leaves the base forecaster untouched, preserves marginal coverage by construction, and transfers across very different forecasters and cohorts.","For subgroup intersections that are simultaneously extremely rare and heavy-tailed (on the order of a dozen calibration patients), no post-hoc quantile method can certify 1-alpha coverage, and the paper treats this as a data-limited frontier.","Across the audited axes the recipe raises subgroup coverage by a mean 4.3 percentage points (6.7 on clinical-risk axes) relative to the marginal band, and it does so by widening bands only where the tail demands it rather than uniformly."],"supporting_citations":[{"why":"Supplies split conformal prediction and the finite-sample marginal coverage guarantee that the audit tests.","marker":"Vovk, Gammerman, and Shafer 2005"},{"why":"Introduces Mondrian per-subgroup conditioning, the object of the rarity mechanism.","marker":"Vovk et al. 2003"},{"why":"Cross-conformal prediction, the basis for leakage-safe pooling of out-of-fold scores in the repair.","marker":"Vovk 2015"},{"why":"Gives the finite-sample 1-2alpha guarantee for the cross-conformal construction the recipe inherits.","marker":"Barber et al. 2021b"},{"why":"Describes the Alzheimer's cohort used as the primary audit dataset.","marker":"Petersen et al. 2010"},{"why":"Provides the external validation cohort for the audit.","marker":"LaMontagne et al. 2019"},{"why":"One of the two base forecasters whose out-of-fold residuals are calibrated in the audit.","marker":"Rubanova, Chen, and Duvenaud 2019"},{"why":"The other base forecaster (a conditional normalizing flow) used in the audit.","marker":"Dinh, Sohl-Dickstein, and Bengio 2017"}],"fun_headline_variants":["Conformal bands mask AD subgroup under-coverage in 57 of 68 checks","Why marginal coverage fails high-risk AD groups: rarity and tail-heaviness","Silent under-coverage: conformal bands fail high-risk AD subgroups","Rarity and tail-heaviness explain AD subgroup under-coverage"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The audit and repair treat each patient-visit as an exchangeable unit and pool coverage counts over horizons, so the finite-sample guarantees and the measured under-coverage rates could change if visits from the same patient are correlated rather than exchangeable.","fun_headline_variants_meta":{"raw":{"variants":["Conformal bands mask AD subgroup under-coverage in 57 of 68 checks","Why marginal coverage fails high-risk AD groups: rarity and tail-heaviness","Silent under-coverage: conformal bands fail high-risk AD subgroups","Rarity and tail-heaviness explain AD subgroup under-coverage"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000727,"raw_usage":{"total_tokens":3346,"prompt_tokens":1121,"completion_tokens":2225,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":737,"completion_tokens_details":{"reasoning_tokens":2144}},"tokens_in":737,"tokens_out":2225,"duration_ms":14620,"temperature":1.0,"reasoning_tokens":2144,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T14:41:28.929606+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute every coverage statistic at the subject level, taking one visit per subject at each horizon, and check whether the 57-of-68 under-coverage count survives; if the deficit mostly vanishes, the reported silent under-coverage is an artifact of pooling multiple visits per patient rather than a property of the conformal method.","supporting_citations":[{"cited_title":"and Benzinger, Tammie L.S","cited_arxiv_id":null,"evidence_quote":"Provides the external validation cohort for the audit."}],"review_version":1}