Pith. sign in

REVIEW 3 major objections 6 minor 70 references

SafeImpute fills missing clinical lab values and releases only those for which the rate of clinically unacceptable errors can be held to a user-chosen level.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-11 04:56 UTC pith:KMP7IR7J

load-bearing objection Solid packaging of event-graph imputation with conformal FDR selection for clinical labs; the hinge is exchangeability, not a broken result. the 3 major comments →

arxiv 2607.05613 v1 pith:KMP7IR7J submitted 2026-07-06 cs.LG

SafeImpute: Reliable Clinical Data Imputation via Conformal Selection

classification cs.LG
keywords Clinical Data ImputationGraph Neural NetworksConformal SelectionFDR ControlLongitudinal EHRReliable ImputationEvent GraphSelective Release
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Real-world clinical records miss many decision-critical lab values because visits are irregular and tests are ordered opportunistically. Most imputers optimize average accuracy and give little guidance on which filled-in numbers are safe to act on. This paper argues that reliable clinical imputation means both accurate prediction and selective release with statistical control of clinically unacceptable errors. SafeImpute builds an event graph of visits with temporal edges within each patient and trend-aware value edges across patients, then imputes with a two-relation graph network and adaptive fusion. It scores each imputation with a label-free proxy risk that combines instability under structure-preserving edge perturbations and a penalty for sparse neighborhood evidence, converts those scores into conformal p-values, and applies the Benjamini–Hochberg procedure so the false discovery rate of unacceptable errors among released values is controlled at a chosen level. On a private diabetes cohort and two public critical-care datasets, the method is competitive on standard error metrics and supports selective release under the stated FDR target.

Core claim

The paper claims that an event-graph two-relation imputer, paired with conformal selection on a proxy risk score of prediction instability plus evidence scarcity, can produce accurate imputations of a designated lab marker while controlling the false discovery rate of clinically unacceptable errors among the released subset at a user-specified tolerance and target level.

What carries the argument

Conformal selection on a proxy risk score: instability of the imputed value under margin-weighted edge perturbations, plus a degree-based evidence penalty, turned into conformal p-values against a calibration set of truly high-error nodes and filtered by Benjamini–Hochberg to control FDR of risky releases.

Load-bearing premise

Once the model and observed covariates are fixed, calibration and test visits must behave as exchangeable so the conformal p-values stay valid; if shared graph structure or how targets are held out breaks that, the formal error-rate guarantee for the released set does not hold.

What would settle it

On repeated random held-out splits of a cohort with known true lab values, run the full pipeline at a fixed clinical tolerance and target FDR level; if the fraction of released imputations whose absolute error meets or exceeds that tolerance systematically exceeds the target level, the reliability claim fails.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Downstream users can be shown only a subset of imputations whose expected share of clinically bad errors is bounded by a chosen level.
  • The same risk-aware selection can improve simple classical imputers when applied to their outputs, so the selection layer is not tied only to the graph model.
  • When a patient has few visits, cross-patient trend-aware edges supply relational signal that both helps accuracy and informs the risk score.
  • Operating points can be tuned by the clinical error tolerance and the target FDR to trade coverage of released values against residual risk.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Unreleased entries could be treated operationally as “order the lab” rather than “use the impute,” turning the method into a decision gate for care workflows.
  • The same conformal selection layer may transfer to other sparse longitudinal markers if a similarly constructed exchangeable calibration set can be held out.
  • If empirical FDR still tracks the target on real electronic health records even when exchangeability is only approximate, the method could remain useful as a practical filter beyond the formal assumption.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. SafeImpute addresses reliable clinical imputation of a decision-critical lab (HbA1c) under irregular, sparse longitudinal records. It builds an event graph with intra-patient temporal edges and inter-patient trend-aware value edges, learns imputations with a two-relation GNN and adaptive fusion plus an auxiliary masked reconstruction loss, then converts a proxy risk score (perturbation instability plus degree-based evidence penalty) into conformal p-values and applies Benjamini–Hochberg to control the FDR of clinically unacceptable errors (|y−ŷ|≥δ) among released imputations. On Mayo Clinic, MIMIC-III, and MIMIC-IV, the method reports competitive MAE/RMSE and selective-release precision, with empirical FDR tracking target α under stated exchangeability/PRDS conditions (Props. 3.1–3.2, Appendix A).

Significance. The paper formulates a practically important problem—accuracy plus selective release with finite-sample FDR control for clinically unacceptable imputation errors—and couples a domain-motivated multi-relation event graph with conformal selection. Strengths include a clear risk-null formulation, an explicit proxy score design, FDR proofs under stated assumptions, multi-dataset evaluation including private Mayo data, ablations of graph and risk-score components, and public code. If the exchangeability hinge holds in deployment and selective evaluation is fair, this is a useful template for risk-aware clinical imputation beyond average reconstruction error.

major comments (3)
  1. [Sec. 3.3, Assumption 3.1, Prop. 3.2] Assumption 3.1 (Sec. 3.3) is load-bearing for Prop. 3.2: residual dependence from shared temporal/value edges, a jointly trained GNN, and patient-level structure may violate exchangeability even with targets masked. The manuscript cites node-level conformal GNN work and reports empirical FDR≈α on Mayo (Fig. 2), but does not stress-test the guarantee (e.g., patient-blocked cal/test splits, leave-patient-out calibration, or sensitivity when value-edge density changes). Without such checks, the finite-sample FDR claim for S(α) remains only partially supported.
  2. [Sec. 4.1.2, Table 1] Table 1’s selective-release comparison evaluates all baselines on the subset selected by SafeImpute’s proxy risk score, then reports each method’s precision on its own outputs. That protocol measures how other imputers perform on nodes SafeImpute ranks as safe, not whether each method can itself produce an FDR-controlled release set. The claim of outperforming baselines in “FDR-controlled selective-release evaluation” (Abstract; Sec. 4.2) is therefore overstated unless baselines receive their own uncertainty/selection mechanisms or a method-agnostic selection rule is used.
  3. [Sec. 4.1.3, Table 5, Sec. 4.5] Table 5 shows low acceptance fractions and power under the chosen operating points (Mayo acceptance 0.1132 / power 0.1818; MIMIC-III 0.0883 / 0.0982; MIMIC-IV 0.1822 / 0.2163). Sec. 4.1.3 also chooses α from the empirical FDR–acceptance trade-off. Together, these weaken the practical claim of useful selective release: error control may be achieved largely by releasing a small, easy subset, and α is not purely a pre-specified clinical tolerance. The paper should report fixed a priori α/δ settings, larger acceptance regimes, and a clearer utility discussion of when ~10% release is clinically actionable.
minor comments (6)
  1. [Abstract / Sec. 1] Abstract and Sec. 1 write “alse discovery rate”; correct to “false discovery rate.”
  2. [Sec. 3.1, Eq. (8)] Eq. (8) uses both 𝑟𝑖,ℓ and 𝑟(𝑖, ℓ); unify the notation for the most recent prior observation.
  3. [Sec. 3.3, Eq. (33); Appendix A] Eq. (33) is hard to parse as written; align the displayed formula with the standard conformal rank form used in Appendix A Eq. (42).
  4. [Table 1] Table 1 Precision is high for several methods on Mayo (often 0.8333 or 1.0), consistent with a small selected set; state selected-set sizes next to Precision for readability.
  5. [Appendix B, Table 6] Mayo has only 84 patients (Table 6); briefly discuss variance and generalizability of the private-cohort results relative to MIMIC.
  6. [Sec. 5.1] Related work covers imputation and clinical UQ well; a short comparison to SAITS/CSDI/BRITS-style clinical time-series imputers under the same selective protocol would help readers place the event-graph design.

Circularity Check

0 steps flagged

No significant circularity: FDR control is external conformal theory applied to a designed proxy; accuracy is held-out evaluation.

full rationale

SafeImpute’s derivation chain does not reduce its central claims to their own inputs by construction. The imputer (event graph, two-relation GNN, adaptive fusion, auxiliary reconstruction) is trained on observed targets and scored with MAE/RMSE on held-out observed labs—standard supervised evaluation, not a fitted quantity renamed as prediction. Reliability control converts a label-free proxy risk score (perturbation instability plus degree-based evidence penalty) into conformal p-values and applies Benjamini–Hochberg; Propositions 3.1–3.2 and Appendix A reduce validity to exchangeability (Assumption 3.1) and standard conformal-selection/BH theory from independent literature (Jin & Candès; Benjamini–Hochberg; related GNN conformal work), not to a self-citation uniqueness theorem or an ansatz that already encodes the target FDR. Choosing operating α from the empirical FDR–acceptance trade-off and reporting precision on the selected subset are operating-point and evaluation design choices, not self-definitional identities (Eq. X = target by construction). No load-bearing step matches the circularity patterns.

Axiom & Free-Parameter Ledger

6 free parameters · 6 axioms · 3 invented entities

The accuracy claim rests on a hand-designed event graph and GNN objective; the reliability claim rests on exchangeability, a label-free proxy that ranks true error, and classical BH under PRDS. Several numeric knobs (α, δ, edge thresholds, β, perturbation schedule) are chosen by the authors and materially change acceptance and reported precision.

free parameters (6)
  • target FDR level α
    Set to 0.15 (Mayo) and 0.35 (MIMIC) from the empirical FDR–acceptance trade-off; directly defines the selective-release operating point.
  • clinical error tolerance δ
    0.6 pp on Mayo and 1.0 pp on MIMIC define the risky-imputation null; changes which errors count as false discoveries.
  • proxy evidence weight β
    Fixed at 0.1 in ˆκ = S_pred + β S_evid; balances instability vs degree penalty and affects p-value ranking.
  • auxiliary loss weight λ
    Fixed at 0.1 in L = L_target + λ L_aux; regularizes multi-lab reconstruction strength.
  • edge construction thresholds (Δ_t, τ_v, τ_δ)
    Determine temporal and trend-aware value edges; graph topology and thus both accuracy and degrees depend on these cutoffs.
  • perturbation keep-probability schedule (π_min, π_max, γ, K)
    Controls structure-preserving edge drops used to compute prediction instability S_pred; shapes the proxy risk score without labels.
axioms (6)
  • domain assumption Calibration and test nodes are exchangeable given observed covariates and fixed trained parameters (Assumption 3.1), so conformal p-values are super-uniform under the risk null.
    Load-bearing for Prop. 3.1; graph dependence could violate it even with targets masked.
  • standard math Null p-values are independent or satisfy PRDS so BH controls FDR at level α (Prop. 3.2).
    Standard multiple-testing condition cited from Benjamini–Yekutieli / conformal selection literature.
  • domain assumption The label-free proxy ˆκ tends to be larger for truly risky imputations (|y−ŷ|≥δ), making the ranking informative for selection power.
    Not proved; supported empirically by Fig. 3 quantile alignment and the instability-only ablation.
  • domain assumption Clinically unacceptable error is adequately captured by absolute residual exceeding a fixed scalar δ on the target lab.
    Defines the risk hypothesis H_j; real clinical harm may depend on direction, context, and patient state.
  • domain assumption Intra-patient temporal continuity and inter-patient value/trend similarity are useful relational signals under opportunistic missingness.
    Motivates the event-graph design; ablations support it but it is not a theorem.
  • ad hoc to paper GCN message passing with adaptive gating and Huber auxiliary reconstruction is a valid learner for the constructed multi-relation event graph.
    Architectural choice of the paper; alternatives could change accuracy and the proxy score distribution.
invented entities (3)
  • Trend-aware value edges on clinical event nodes no independent evidence
    purpose: Connect visits across patients using RMS-normalized lab-value and trend distances to share information under sparse trajectories.
    Specific edge rule (Eqs. 8–13) is a paper-defined construction, not an independently measured clinical object.
  • Proxy risk score ˆκ = prediction instability + β·evidence penalty no independent evidence
    purpose: Provide a label-free nonconformity signal for conformal testing of risky imputations.
    Defined in Sec. 3.3 from perturbations and degrees; validated only inside this paper’s plots/ablations.
  • Two-relation GNN with per-node adaptive fusion gate for temporal vs value messages no independent evidence
    purpose: Learn event representations that balance intra-patient and inter-patient signals for target-lab imputation.
    Methodological construct of the paper; no external existence claim beyond model performance.

pith-pipeline@v1.1.0-grok45 · 24781 in / 4006 out tokens · 35024 ms · 2026-07-11T04:56:22.552479+00:00 · methodology

0 comments
read the original abstract

Clinical care often relies on key laboratory indicators, yet real-world patient visits are sparse and tests are ordered irregularly, leading to pervasive missingness. While many imputation methods improve average accuracy, they provide limited guidance on which imputed values are reliable enough for high-stakes downstream use. In this work, we study reliable clinical imputation, aiming to produce accurate imputations while selectively releasing the reliable results, with statistical control over clinically unacceptable errors. To achieve this goal, we propose SafeImpute, a reliable imputation framework for irregular and sparse clinical longitudinal records. SafeImpute constructs an event graph that captures both intra-patient temporal trajectories and inter-patient clinical similarity, and learns imputations with a two-relation GNN and adaptive fusion, regularized by an auxiliary masked reconstruction objective. For reliability guarantees, SafeImpute converts a proxy risk score into conformal p-values and applies the Benjamini--Hochberg procedure to control the false discovery rate (FDR) of unacceptable errors among released imputations at a user-specified tolerance. Experiments on our Mayo Clinic data, the public MIMIC-III and MIMIC-IV datasets show that SafeImpute achieves strong imputation accuracy while providing reliable error control, outperforming diverse baselines in both standard imputation evaluation and FDR-controlled selective-release evaluation.

Figures

Figures reproduced from arXiv: 2607.05613 by Curtiss B. Cook, Jingrui He, Junting Wang, Mengting Ai, Xinrui He.

Figure 1
Figure 1. Figure 1: SafeImpute overview for reliable clinical imputation. We convert irregular and sparse longitudinal records into an event graph with temporal edges and trend-aware value edges, then learn with a two-relation GNN combined with an adaptive fusion gating network to impute missing labs. To achieve FDR control over clinically unacceptable errors, we compute a proxy risk score that combines perturbation-induced i… view at source ↗
Figure 2
Figure 2. Figure 2: FDR–power trade-off of conformal selection on Mayo Clinic dataset. Each panel fixes the clinical tolerance [PITH_FULL_IMAGE:figures/full_fig_p008_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Alignment between the proxy risk score 𝜅ˆ and the real imputation error. Selective Release Summary [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗
Figure 7
Figure 7. Figure 7: Additional FDR–power trade-off curves on MIMIC [PITH_FULL_IMAGE:figures/full_fig_p012_7.png] view at source ↗
Figure 5
Figure 5. Figure 5: Left: distribution of the number of visits per patient. [PITH_FULL_IMAGE:figures/full_fig_p012_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Visit-wise missingness for 10 randomly sampled [PITH_FULL_IMAGE:figures/full_fig_p012_6.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

70 extracted references · 6 canonical work pages · 3 internal anchors

  1. [1]

    Anastasios N Angelopoulos, Rina Foygel Barber, and Stephen Bates. 2024. The- oretical foundations of conformal prediction.arXiv preprint arXiv:2411.11824 (2024)

  2. [2]

    Anastasios Nikolas Angelopoulos, Stephen Bates, Adam Fisch, Lihua Lei, and Tal Schuster. 2024. Conformal Risk Control. InThe Twelfth International Conference on Learning Representations. https://openreview.net/forum?id=33XGfHLtZg

  3. [3]

    2014.Introduction to imprecise probabilities

    Thomas Augustin, Frank PA Coolen, Gert De Cooman, and Matthias CM Troffaes. 2014.Introduction to imprecise probabilities. John Wiley & Sons

  4. [4]

    Verona, Curtiss B

    Yikun Ban, Xinrui He, Patricia M. Verona, Curtiss B. Cook, and Jingrui He

  5. [5]

    arXiv:https://doi.org/10.1080/20565623.2025.2567166 doi:10.1080/20565623.2025

    GLM-DM: language model boosted neural networks for HbA1c trend prediction in diabetes mellitus.Future Science OA11, 1 (2025), 2567166. arXiv:https://doi.org/10.1080/20565623.2025.2567166 doi:10.1080/20565623.2025. 2567166 PMID: 41215678

  6. [6]

    Stephen Bates, Emmanuel Candès, Lihua Lei, Yaniv Romano, and Matteo Sesia

  7. [7]

    Testing for outliers with conformal p-values.The Annals of Statistics51, 1 (2023), 149–178

  8. [8]

    Brett K Beaulieu-Jones, Jason H Moore, and Pooled Resource Open-Access ALS Clinical Trials Consortium. 2017. Missing data imputation in the electronic health record using deeply learned autoencoders. InPacific symposium on biocomputing

  9. [9]

    World Scientific, 207–218

  10. [10]

    Yoav Benjamini and Yosef Hochberg. 1995. Controlling the false discovery rate: a practical and powerful approach to multiple testing.Journal of the Royal statistical society: series B (Methodological)57, 1 (1995), 289–300

  11. [11]

    Yoav Benjamini and Daniel Yekutieli. 2001. The control of the false discovery rate in multiple testing under dependency.Annals of statistics(2001), 1165–1188

  12. [12]

    David M Blei, Alp Kucukelbir, and Jon D McAuliffe. 2017. Variational inference: A review for statisticians.Journal of the American statistical Association112, 518 (2017), 859–877

  13. [13]

    Charles Blundell, Julien Cornebise, Koray Kavukcuoglu, and Daan Wierstra. 2015. Weight uncertainty in neural network. InInternational conference on machine learning. PMLR, 1613–1622

  14. [14]

    Andrea Campagner, Elia Mario Biganzoli, Clara Balsano, Cristina Cereda, and Federico Cabitza. 2025. Modeling unknowns: A vision for uncertainty-aware machine learning in healthcare.International Journal of Medical Informatics203 (2025), 106014. doi:10.1016/j.ijmedinf.2025.106014

  15. [15]

    Wei Cao, Dong Wang, Jian Li, Hao Zhou, Lei Li, and Yitan Li. 2018. Brits: Bidirectional recurrent imputation for time series.Advances in neural information processing systems31 (2018)

  16. [16]

    Zhengping Che, Sanjay Purushotham, Kyunghyun Cho, David Sontag, and Yan Liu. 2018. Recurrent neural networks for multivariate time series with missing values.Scientific reports8, 1 (2018), 6085

  17. [17]

    Andrea Cini, Ivan Marisca, and Cesare Alippi. 2021. Filling the g_ap_s: Mul- tivariate time series imputation by graph neural networks.arXiv preprint arXiv:2108.00298(2021)

  18. [18]

    Tianyu Du, Luca Melis, and Ting Wang. 2024. ReMasker: Imputing Tabular Data with Masked Autoencoding. InThe Twelfth International Conference on Learning Representations. https://openreview.net/forum?id=KI9NqjLVDT

  19. [19]

    Wenjie Du, David Côté, and Yan Liu. 2023. Saits: Self-attention-based imputation for time series.Expert Systems with Applications219 (2023), 119619

  20. [20]

    Pamela Giustinelli, Charles F Manski, and Francesca Molinari. 2022. Precise or imprecise probabilities? Evidence from survey response related to late-onset 9 Xinrui He, Mengting Ai, Junting Wang, Curtiss B. Cook and Jingrui He dementia.Journal of the European Economic Association20, 1 (2022), 187–221

  21. [21]

    Lovedeep Gondara and Ke Wang. 2018. Mida: Multiple imputation using de- noising autoencoders. InPacific-Asia conference on knowledge discovery and data mining. Springer, 260–272

  22. [22]

    2013.Monte carlo methods

    John Hammersley. 2013.Monte carlo methods. Springer Science & Business Media

  23. [23]

    Xinrui He, Yikun Ban, Jiaru Zou, Tianxin Wei, Curtiss Cook, and Jingrui He. 2025. LLM-Forest: Ensemble Learning of LLMs with Graph-Augmented Prompts for Data Imputation. InFindings of the Association for Computational Linguistics: ACL 2025, Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar (Eds.). Association for Computational Li...

  24. [24]

    doi:10.18653/v1/2025.findings-acl.361

  25. [25]

    Xinrui He, Tianxin Wei, and Jingrui He. 2023. Robust Basket Recommendation via Noise-tolerated Graph Contrastive Learning. InProceedings of the 32nd ACM International Conference on Information and Knowledge Management(Birming- ham, United Kingdom)(CIKM ’23). Association for Computing Machinery, New York, NY, USA, 709–719. doi:10.1145/3583780.3615039

  26. [26]

    Sepp Hochreiter and Jürgen Schmidhuber. 1997. Long short-term memory.Neural computation9, 8 (1997), 1735–1780

  27. [27]

    Zarin Tahia Hossain and Mostafa Milani. 2025. Beyond Accuracy: An Empirical Study of Uncertainty Estimation in Imputation.arXiv preprint arXiv:2511.21607 (2025)

  28. [28]

    On the Limits of Selective AI Prediction: A Case Study in Clinical Decision Making

    Sarah Jabbour, David Fouhey, Nikola Banovic, Stephanie D. Shepard, Ella Kaze- rooni, Michael W. Sjoding, and Jenna Wiens. 2025. On the Limits of Selective AI Prediction: A Case Study in Clinical Decision Making. arXiv:2508.07617 [cs.HC] https://arxiv.org/abs/2508.07617

  29. [29]

    Daniel Jarrett, Bogdan C Cebere, Tennison Liu, Alicia Curth, and Mihaela van der Schaar. 2022. Hyperimpute: Generalized iterative imputation with automatic model selection. InInternational Conference on Machine Learning. PMLR, 9916– 9937

  30. [30]

    Ying Jin and Emmanuel J Candès. 2023. Selection by prediction with conformal p-values.Journal of Machine Learning Research24, 244 (2023), 1–41

  31. [31]

    Taeho Jo, Eun Hye Lee, Alzheimer’s Disease Neuroimaging Initiative (ADNI), and the Alzheimer’s Disease Sequencing Project (ADSP). 2025. Uncertainty-aware genomic classification of Alzheimer’s disease: a transformer-based ensemble approach with Monte Carlo dropout.Briefings in Bioinformatics26, 6 (2025), bbaf587

  32. [32]

    Alistair EW Johnson, Lucas Bulgarelli, Lu Shen, Alvin Gayles, Ayad Shammout, Steven Horng, Tom J Pollard, Sicheng Hao, Benjamin Moody, Brian Gow, et al

  33. [33]

    MIMIC-IV, a freely accessible electronic health record dataset.Scientific data10, 1 (2023), 1

  34. [34]

    Alistair EW Johnson, Tom J Pollard, Lu Shen, Li-wei H Lehman, Mengling Feng, Mohammad Ghassemi, Benjamin Moody, Peter Szolovits, Leo Anthony Celi, and Roger G Mark. 2016. MIMIC-III, a freely accessible critical care database. Scientific data3, 1 (2016), 1–9

  35. [35]

    Siddhartha Kapuria, Patrick Minot, Ariel Kapusta, Naruhiko Ikoma, and Farshid Alambeigi. 2024. A novel dual layer cascade reliability framework for an informed and intuitive clinician-ai interaction in diagnosis of colorectal cancer polyps. IEEE Journal of Biomedical and Health Informatics28, 4 (2024), 2326–2337

  36. [36]

    Ki-Yeol Kim, Byoung-Jin Kim, and Gwan-Su Yi. 2004. Reuse of imputed data in microarray analysis increases imputation efficiency.BMC bioinformatics5, 1 (2004), 160

  37. [37]

    2019.Statistical analysis with missing data

    Roderick JA Little and Donald B Rubin. 2019.Statistical analysis with missing data. Vol. 793. John Wiley & Sons

  38. [38]

    L López, Shaza Elsharief, Dhiyaa Al Jorf, Firas Darwish, Congbo Ma, and Farah E Shamout. 2025. Uncertainty Quantification for Machine Learning in Healthcare: A Survey.arXiv preprint arXiv:2505.02874(2025)

  39. [39]

    Charles Lu, Andréanne Lemay, Ken Chang, Katharina Höbel, and Jayashree Kalpathy-Cramer. 2022. Fair conformal predictors for applications in medical imaging. InProceedings of the AAAI conference on artificial intelligence, Vol. 36. 12008–12016

  40. [40]

    Pierre-Alexandre Mattei and Jes Frellsen. 2019. MIWAE: Deep generative mod- elling and imputation of incomplete data sets. InInternational conference on machine learning. PMLR, 4413–4423

  41. [41]

    Alireza Mehrtash, William M Wells, Clare M Tempany, Purang Abolmaesumi, and Tina Kapur. 2020. Confidence calibration and predictive uncertainty estimation for deep medical image segmentation.IEEE transactions on medical imaging39, 12 (2020), 3868–3878

  42. [42]

    Alexander S Millar, John Arnn, Sam Himes, and Julio C Facelli. 2024. Uncertainty in breast cancer risk prediction: a conformal prediction study of race stratification. InMEDINFO 2023—The Future Is Accessible. IOS Press, 991–995

  43. [43]

    Kanika Narang, Adit Krishnan, Junting Wang, Chaoqi Yang, Hari Sundaram, and Carolyn Sutter. 2021. Ranking User-Generated Content via Multi-Relational Graph Convolution. InProceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval(Virtual Event, Canada) (SIGIR ’21). Association for Computing Machinery, N...

  44. [44]

    Christina Papangelou, Konstantinos Kyriakidis, Pantelis Natsiavas, Ioanna Chou- varda, and Andigoni Malousi. 2025. Reliable machine learning models in genomic medicine using conformal prediction.Frontiers in BioinformaticsVolume 5 - 2025 (2025). doi:10.3389/fbinf.2025.1507448

  45. [45]

    Riyi Qiu, Yugang Jia, Mirsad Hadzikadic, Michael Dulin, Xi Niu, and Xin Wang

  46. [46]

    Modeling the uncertainty in electronic health records: a Bayesian deep learning approach.arXiv preprint arXiv:1907.06162(2019)

  47. [47]

    Aravind Sankar, Junting Wang, Adit Krishnan, and Hari Sundaram. 2020. Beyond Localized Graph Neural Networks: An Attributed Motif Regularization Frame- work. In2020 IEEE International Conference on Data Mining (ICDM). 472–481. doi:10.1109/ICDM50108.2020.00056

  48. [48]

    Aravind Sankar, Junting Wang, Adit Krishnan, and Hari Sundaram. 2022. Self- supervised role learning for graph neural networks.Knowledge and Information Systems64, 8 (2022), 2091–2121

  49. [49]

    Yishan Shen, Yuyang Ye, Hui Xiong, and Yong Chen. 2025. SAFER: A Cal- ibrated Risk-Aware Multimodal Recommendation Model for Dynamic Treat- ment Regimes. InForty-second International Conference on Machine Learning. https://openreview.net/forum?id=7UqNM85dD6

  50. [50]

    Shariq I Sherwani, Haseeb A Khan, Aishah Ekhzaimy, Afshan Masood, and Meena K Sakharkar. 2016. Significance of HbA1c test in diagnosis and prognosis of diabetic patients.Biomarker insights11 (2016), BMI–S38440

  51. [51]

    Daniel J Stekhoven and Peter Bühlmann. 2012. MissForest—non-parametric missing value imputation for mixed-type data.Bioinformatics28, 1 (2012), 112– 118

  52. [52]

    Yusuke Tashiro, Jiaming Song, Yang Song, and Stefano Ermon. 2021. CSDI: Con- ditional Score-based Diffusion Models for Probabilistic Time Series Imputation. arXiv:2107.03502 [cs.LG] https://arxiv.org/abs/2107.03502

  53. [53]

    Olga Troyanskaya, Michael Cantor, Gavin Sherlock, Pat Brown, Trevor Hastie, Robert Tibshirani, David Botstein, and Russ B Altman. 2001. Missing value estimation methods for DNA microarrays.Bioinformatics17, 6 (2001), 520–525

  54. [54]

    Alireza Vafaei Sadr, Jiang Li, Wenke Hwang, Mohammed Yeasin, Ming Wang, Harold Lehmann, Ramin Zand, and Vida Abedi. 2025. Flexible imputation toolkit for electronic health records.Scientific reports15, 1 (2025), 17176

  55. [55]

    Janette Vazquez and Julio C Facelli. 2022. Conformal prediction in clinical medical sciences.Journal of Healthcare Informatics Research6, 3 (2022), 241–252

  56. [56]

    Pascal Vincent, Hugo Larochelle, Yoshua Bengio, and Pierre-Antoine Manzagol

  57. [57]

    In Proceedings of the 25th international conference on Machine learning

    Extracting and composing robust features with denoising autoencoders. In Proceedings of the 25th international conference on Machine learning. 1096–1103

  58. [58]

    Junting Wang, Chenghuan Guo, Jiao Yang, Yanhui Guo, Yan Gao, and Hari Sun- daram. 2025. Multi-modal Relational Item Representation Learning for Inferring Substitutable and Complementary Items.arXiv preprint arXiv:2507.22268(2025)

  59. [59]

    Dongxia Wu, Liyao Gao, Matteo Chinazzi, Xinyue Xiong, Alessandro Vespignani, Yi-An Ma, and Rose Yu. 2021. Quantifying uncertainty in deep spatiotemporal forecasting. InProceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining. 1841–1851

  60. [60]

    Zhichao Yang, Avijit Mitra, Weisong Liu, Dan Berlowitz, and Hong Yu. 2023. TransformEHR: transformer-based encoder-decoder generative model to en- hance prediction of disease outcomes using electronic health records.Nature communications14, 1 (2023), 7857

  61. [61]

    Jinsung Yoon, James Jordon, and Mihaela van der Schaar. 2018. GAIN: Missing Data Imputation using Generative Adversarial Nets. InProceedings of the 35th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 80), Jennifer Dy and Andreas Krause (Eds.). PMLR, 5689–5698. https://proceedings.mlr.press/v80/yoon18a.html

  62. [62]

    Jiaxuan You, Xiaobai Ma, Yi Ding, Mykel J Kochenderfer, and Jure Leskovec

  63. [63]

    In Advances in Neural Information Processing Systems, H

    Handling Missing Data with Graph Representation Learning. In Advances in Neural Information Processing Systems, H. Larochelle, M. Ran- zato, R. Hadsell, M.F. Balcan, and H. Lin (Eds.), Vol. 33. Curran Associates, Inc., 19075–19087. https://proceedings.neurips.cc/paper_files/paper/2020/file/ dc36f18a9a0a776671d4879cae69b551-Paper.pdf

  64. [64]

    Soroush H Zargarbashi, Simone Antonelli, and Aleksandar Bojchevski. 2023. Conformal prediction sets for graph neural networks. InInternational Conference on Machine Learning. PMLR, 12292–12318

  65. [65]

    Hengrui Zhang, Liancheng Fang, Qitian Wu, and Philip S Yu. 2025. Diffputer: Empowering diffusion models for missing data imputation. InThe Thirteenth International Conference on Learning Representations

  66. [66]

    He Zhao, Ke Sun, Amir Dezfouli, and Edwin V Bonilla. 2023. Transformed distribution matching for missing value imputation. InInternational Conference on Machine Learning. PMLR, 42159–42186

  67. [67]

    Yidong Zhao, Changchun Yang, Artur Schweidtmann, and Qian Tao. 2022. Effi- cient bayesian uncertainty estimation for nnu-net. InInternational Conference on Medical Image Computing and Computer-Assisted Intervention. Springer, 535–544

  68. [68]

    Lecheng Zheng, Baoyu Jing, Zihao Li, Zhichen Zeng, Tianxin Wei, Mengting Ai, Xinrui He, Lihui Liu, Dongqi Fu, Jiaxuan You, et al. 2025. Pyg-ssl: A graph self-supervised learning toolkit. InProceedings of the 34th ACM International Conference on Information and Knowledge Management. 6580–6586

  69. [69]

    Jessica Zhou, Kaeli Rizzo, Trevor Christensen, Ziqi Tang, and Peter K Koo. 2026. Uncertainty-aware genomic deep learning with knowledge distillation.npj 10 SafeImpute: Reliable Clinical Data Imputation via Conformal Selection Artificial Intelligence2, 1 (2026), 3

  70. [70]

    Í𝑚 𝑗=1 𝑇𝑗 𝑅 𝑗 max{1, Í𝑚 𝑗=1 𝑅 𝑗 } # =E

    Jiaru Zou, Dongqi Fu, Sirui Chen, Xinrui He, Zihao Li, Yada Zhu, Jiawei Han, and Jingrui He. 2026. RAG over Tables: Hierarchical Memory Index, Multi-Stage Retrieval, and Benchmarking. InICLR 2026 Workshop on Logical Reasoning of Large Language Models. https://openreview.net/forum?id=LUW0zHCE2W A Proof of FDR Control Goal.Let 𝜅(𝑖)=|𝑦 𝑖 − ˆ𝑦(𝑖)| be the true...