REVIEW 1 major objections 1 minor 112 references
SurvBench: A Standardised Preprocessing Pipeline for Multi-Modal Electronic Health Record Survival Analysis
T0 review · 1 major / 1 minor · reviewed 2026-05-17 · grok-4.3
Pith's one-line read A configurable preprocessing pipeline converts raw electronic health records into consistent tensors for comparing survival models across studies.
desk verdict SurvBench bundles standard EHR preprocessing steps into a YAML-driven pipeline across four databases and multiple modalities, which could help with reproducibility if the config surface is truly complete. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The preprocessing pipeline that standardizes cohort definitions, time discretization, missingness handling with binary masks, and censoring rules for single-risk and competing-risk survival endpoints.
What would settle it
A controlled experiment in which two otherwise identical survival models produce materially different performance rankings when trained on data prepared under alternative but equally defensible preprocessing rules.
Extended reading notes
Core claim
The pipeline converts raw data exports into model-ready tensors for survival analysis, supporting multiple critical care sources and four input modalities of time-series vitals and laboratory values, static demographics, ICD codes, and radiology report embeddings, with every preprocessing decision controlled through YAML configuration files and with imputation, scaling, and feature filtering performed on the training fold only.
Load-bearing premise
That the specific preprocessing choices encoded in the pipeline represent the appropriate standard the field should adopt rather than one reasonable option among many.
Editorial extensions
If this is right
- Model comparisons can isolate the effect of architecture and training choices without preprocessing differences confounding concordance metrics.
- Imputation and scaling statistics are derived exclusively from the training portion of each split to prevent information leakage.
- Missing values are accompanied by explicit binary mask tensors so models can learn from missingness patterns directly.
- The same configuration can be reused to produce harmonized external validation sets across different data sources.
Reading between the lines
- Widespread use could shorten the time researchers spend on data preparation and shift effort toward model innovation.
- The configuration-driven design makes it straightforward to test how small changes in cohort or censoring rules affect final model rankings.
- Community extensions could add new modalities or data sources while preserving the same controlled comparison environment.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents SurvBench, an open-source preprocessing pipeline that converts raw PhysioNet exports from four critical-care databases (MIMIC-IV, eICU, MC-MED, HiRID) into model-ready tensors for survival analysis. It supports four modalities (time-series vitals/labs, static demographics, ICD codes, radiology embeddings), single-risk and competing-risk endpoints, missingness masks, train-only imputation/scaling, and cross-dataset harmonization, with every preprocessing decision controlled through YAML configuration files.
Significance. If the YAML interface fully exposes all decisions as claimed, this artifact would meaningfully advance reproducibility in deep-learning EHR survival work by supplying an explicit, configurable baseline that future studies can adopt for matched comparisons, particularly in multi-modal settings. The open-source release and concrete engineering choices (train-only fits, masks, competing-risk support) are strengths that could reduce the impact of undocumented preprocessing on reported model differences.
major comments (1)
- [Abstract] Abstract: the central claim that 'Every preprocessing decision is controlled through YAML configuration' is load-bearing for the standardisation objective. To substantiate that the pipeline eliminates the reproducibility barrier, the manuscript must demonstrate that cohort inclusion/exclusion logic, time-binning rules, censoring definitions (including for the three-way competing-risks pathway), missingness mask generation, and cross-modality alignment are all fully parameterised in the YAML schema without hard-coded or dataset-specific defaults that remain outside user control.
minor comments (1)
- Consider adding an explicit table or appendix that enumerates all YAML keys, their defaults, and the corresponding preprocessing step they control, to make the configuration surface immediately verifiable by readers.
Simulated Author's Rebuttal
We thank the referee for their constructive feedback and for recognizing the potential of SurvBench to improve reproducibility in multi-modal EHR survival analysis. We agree that explicitly demonstrating the full YAML parameterization is necessary to substantiate the central claim and have prepared revisions accordingly.
read point-by-point responses
-
Referee: [Abstract] Abstract: the central claim that 'Every preprocessing decision is controlled through YAML configuration' is load-bearing for the standardisation objective. To substantiate that the pipeline eliminates the reproducibility barrier, the manuscript must demonstrate that cohort inclusion/exclusion logic, time-binning rules, censoring definitions (including for the three-way competing-risks pathway), missingness mask generation, and cross-modality alignment are all fully parameterised in the YAML schema without hard-coded or dataset-specific defaults that remain outside user control.
Authors: We agree this demonstration is essential. In the revised manuscript we will add a new subsection in Methods (and an appendix table) that explicitly maps each listed decision to its YAML key(s): cohort inclusion/exclusion under 'cohort_selection' (with flags for each database and user-overridable criteria), time-binning under 'time_discretization' (bin_size, alignment, aggregation), censoring definitions for both single-risk and the three-way competing-risks pathway (including home-discharge handling) under 'endpoint_definition', missingness mask generation under 'missingness' (mask creation and propagation rules), and cross-modality alignment under 'modality_alignment' (temporal syncing and feature harmonization). We will include verbatim excerpts from the default config files for MIMIC-IV and eICU, show how every parameter is exposed without hard-coded fallbacks, and reference the exact Python functions that read these keys. The open-source repository already implements this design; the revision will make the mapping transparent in the paper itself. revision: yes
Circularity Check
No circularity: software pipeline defined by external configuration, no derivation chain
full rationale
The manuscript describes an open-source preprocessing pipeline whose behavior is explicitly governed by user-supplied YAML files rather than any internal equations or fitted quantities. No mathematical derivations, predictions, or self-referential definitions appear in the abstract or described contribution. Imputation, scaling, and feature filtering are stated to be fit on the training fold only, but this is a standard data-split practice and does not create a closed loop within the paper's own claims. The central assertion—that every preprocessing decision is controllable via configuration—is an engineering interface claim, not a logical reduction that collapses back onto its own inputs by construction. No self-citations or uniqueness theorems are invoked as load-bearing premises. The work is therefore self-contained as a software artifact.
Assumptions & free parameters
assumptions (1)
- domain assumption Raw exports from MIMIC-IV, eICU, MC-MED and HiRID follow documented PhysioNet schemas for time-series, static demographics, ICD codes and radiology reports.
Cite this review
Pith. "Pith review of SurvBench: A Standardised Preprocessing Pipeline for Multi-Modal Electronic Health Record Survival Analysis." pith.science (2026). https://pith.science/paper/2511.11935
@misc{pith2026251111935,
author = {Pith},
title = {Pith review of: SurvBench: A Standardised Preprocessing Pipeline for Multi-Modal Electronic Health Record Survival Analysis},
year = {2026},
howpublished = {\url{https://pith.science/paper/2511.11935}},
note = {Machine review of arXiv:2511.11935}
}
read the original abstract
Deep-learning survival models for electronic health record (EHR) data are hard to compare across papers because the upstream preprocessing step, which includes cohort definition, time discretisation, missingness handling, and censoring rules, is typically undocumented and inconsistent. A reported difference in concordance between two mortality models can therefore reflect any of these choices rather than a modelling contribution. We present SurvBench, an open-source preprocessing pipeline that converts raw PhysioNet exports into model-ready tensors for survival analysis. SurvBench covers four critical-care databases (MIMIC-IV, eICU, MC-MED, HiRID) and four input modalities: time-series vitals and laboratory values, static demographics, International Classification of Diseases (ICD) codes, and radiology report embeddings. Every preprocessing decision is controlled through YAML configuration. Imputation, scaling, and feature filtering are fit on the training fold only. Missingness is recorded as a binary mask alongside each feature tensor. The pipeline handles single-risk endpoints (in-hospital and in-ICU mortality) and competing-risks endpoints (a three-way emergency-department admission pathway, with home discharge treated as administrative censoring). We also provide support for harmonised cross-dataset external validation between eICU and MIMIC-IV. SurvBench is publicly available at https://github.com/munibmesinovic/SurvBench, providing a robust platform that future deep-learning EHR survival work, especially nascent multi-modal approaches, can be measured against under matched preprocessing.
Reference graph
Works this paper leans on
-
[1]
Journal of the Royal Statistical Society: Series B (Methodological)34(2), 187–220 (1972)
Cox, D.R.: Regression models and life-tables. Journal of the Royal Statistical Society: Series B (Methodological)34(2), 187–220 (1972)
work page 1972
-
[2]
Journal of the American Statistical Association53(282), 457–481 (1958)
Kaplan, E.L., Meier, P.: Nonparametric estimation from incomplete observations. Journal of the American Statistical Association53(282), 457–481 (1958)
work page 1958
-
[3]
Journal of Machine Learning Research20(129), 1–30 (2019)
Kvamme, H., Borgan, Ø., Scheel, I.: Time-to-event prediction with neural networks and Cox regression. Journal of Machine Learning Research20(129), 1–30 (2019)
work page 2019
-
[4]
In: Thirty-Second AAAI Conference on Artificial Intelligence (2018)
Lee, C., Zame, W.R., Yoon, J., Schaar, M.: DeepHit: A deep learning approach to survival analysis with competing risks. In: Thirty-Second AAAI Conference on Artificial Intelligence (2018)
work page 2018
-
[5]
BMC Medical Research Methodology18(1), 24 (2018) https://doi.org/10.1186/ s12874-018-0482-1
Katzman, J.L., Shaham, U., Cloninger, A., Bates, J., Jiang, T., Kluger, Y.: DeepSurv: Personalised treatment recommender system using a Cox proportional hazards deep neural 19 network. BMC Medical Research Methodology18(1), 24 (2018) https://doi.org/10.1186/ s12874-018-0482-1
work page 2018
-
[6]
Mesinovic, M., Watkinson, P., Zhu, T.: Dysurv: dynamic deep learning model for survival analysis with conditional variational inference. Journal of the American Medical Informatics Association, 271 (2024) https://doi.org/10.1093/jamia/ocae271
-
[7]
In: Temporal Graph Learning Workshop@ KDD 2025 (2025)
Mesinovic, M., Watkinson, P., Zhu, T.: Multi-modal interpretable graph for competing risk prediction with electronic health records. In: Temporal Graph Learning Workshop@ KDD 2025 (2025)
work page 2025
-
[8]
Harutyunyan, H., Khachatrian, H., Kale, D.C., Ver Steeg, G., Galstyan, A.: Multitask learning and benchmarking with clinical time series data. Scientific Data6(1), 96 (2019) https://doi. org/10.1038/s41597-019-0103-9
Show all 112 references
-
[9]
NPJ Digital Medicine1(1), 18 (2018) https://doi.org/10.1038/s41746-018-0029-1
Rajkomar, A., Oren, E., Chen, K., Dai, A.M., Hajaj, N., Hardt, M., Liu, P.J., Liu, X., Marcus, J., Sun, M., Sundberg, P., Yee, H., Zhang, K., Zhang, Y., Flores, G., Duggan, G.E., Irvine, J., Le, Q., Litsch, K., Mossin, A., Tansuwan, J., Wang, D., Wexler, J., Wilson, J., Ludwig...
2018 doi
-
[10]
IEEE Journal of Biomedical and Health Informatics22(5), 1589–1604 (2018) https://doi.org/10.1109/JBHI.2017.2767063
Shickel, B., Tighe, P.J., Bihorac, A., Rashidi, P.: Deep EHR: A survey of recent advances in deep learning techniques for electronic health record (EHR) analysis. IEEE Journal of Biomedical and Health Informatics22(5), 1589–1604 (2018) https://doi.org/10.1109/JBHI.2017.2767063
2018 doi
-
[11]
Scientific Reports8(1), 6085 (2018) https://doi.org/ 10.1038/s41598-018-24271-9
Che, Z., Purushotham, S., Cho, K., Sontag, D., Liu, Y.: Recurrent neural networks for mul- tivariate time series with missing values. Scientific Reports8(1), 6085 (2018) https://doi.org/ 10.1038/s41598-018-24271-9
2018 doi
-
[12]
In: International Conference on Learning Representations (2016)
Lipton, Z.C., Kale, D.C., Elkan, C., Wetzel, R.: Learning to diagnose with LSTM recurrent neural networks. In: International Conference on Learning Representations (2016)
2016
-
[13]
eGEMs1(3), 7 (2013) https://doi.org/10.13063/ 2327-9214.1035
Wells, B.J., Chagin, K.M., Nowacki, A.S., Kattan, M.W.: Strategies for handling missing data in electronic health record derived data. eGEMs1(3), 7 (2013) https://doi.org/10.13063/ 2327-9214.1035
2013
-
[14]
In: Proceedings of the AAAI Conference on Artificial Intelligence, pp
Luo, Y., Xin, Y., Joshi, R., Celi, L., Szolovits, P.: Predicting ICU mortality risk by grouping temporal trends from a multivariate panel of physiologic measurements. In: Proceedings of the AAAI Conference on Artificial Intelligence, pp. 42–50 (2016)
2016
-
[15]
Big Data Analytics1(1), 9 (2016) https://doi.org/10.1186/ s41044-016-0014-0
Garc´ ıa, S., Ram´ ırez-Gallego, S., Luengo, J., Ben´ ıtez, J.M., Herrera, F.: Big data prepro- cessing: Methods and prospects. Big Data Analytics1(1), 9 (2016) https://doi.org/10.1186/ s41044-016-0014-0
2016
-
[16]
Proceedings of the IEEE104(2), 444–466 (2016) https://doi.org/10.1109/JPROC.2015.2501978
Johnson, A.E.W., Ghassemi, M.M., Nemati, S., Niehaus, K.E., Clifton, D.A., Clifford, G.D.: Machine learning and decision support in critical care. Proceedings of the IEEE104(2), 444–466 (2016) https://doi.org/10.1109/JPROC.2015.2501978
2016 doi
-
[17]
Scientific Data3, 160035 (2016) https://doi.org/10.1038/sdata.2016.35
Johnson, A.E.W., Pollard, T.J., Shen, L., Lehman, L.-w.H., Feng, M., Ghassemi, M., Moody, B., Szolovits, P., Celi, L.A., Mark, R.G.: MIMIC-III, a freely accessible critical care database. Scientific Data3, 160035 (2016) https://doi.org/10.1038/sdata.2016.35
2016 doi
-
[18]
Statistics in Medicine26(11), 2389–2430 (2007) https://doi.org/10.1002/sim.2712
Putter, H., Fiocco, M., Geskus, R.B.: Tutorial in biostatistics: Competing risks and multi-state models. Statistics in Medicine26(11), 2389–2430 (2007) https://doi.org/10.1002/sim.2712
2007 doi
-
[19]
Statistics in Medicine31(11-12), 1074–1088 (2012) https://doi.org/10
Andersen, P.K., Keiding, N.: Interpretability and importance of functionals in competing risks and multistate models. Statistics in Medicine31(11-12), 1074–1088 (2012) https://doi.org/10. 1002/sim.4385 20
2012
-
[20]
Journal of Machine Learning Research17(1), 2797–2819 (2016)
Wiens, J., Guttag, J., Horvitz, E.: Patient risk stratification with time-varying parameters: A multitask learning approach. Journal of Machine Learning Research17(1), 2797–2819 (2016)
2016
-
[21]
In: Machine Learning for Healthcare Conference, pp
Nestor, B., McDermott, M.B.A., Boag, W., Berner, G., Naumann, T., Hughes, M.C., Ghassemi, M., Szolovits, P.: Feature robustness in non-stationary health records: Caveats to deploy- able model performance in common clinical machine learning tasks. In: Machine Learning for Healt...
2019
-
[22]
In: NeurIPS 2019 Workshop on Machine Learning for Health (ML4H) (2019).https://arxiv.org/abs/1909.02832
Ren, Y., Yang, M., Li, Y., He, L., Liu, W.: DeepWeiSurv: A Weibull-based deep learning model for survival analysis. In: NeurIPS 2019 Workshop on Machine Learning for Health (ML4H) (2019).https://arxiv.org/abs/1909.02832
2019
-
[23]
P., L.P., Gotze, T., Li, H., V
Aastha, Zare, A., He, Z.-J., L. P., L.P., Gotze, T., Li, H., V. S., V.S., Li, X., Sun, J.: Deep- Compete: A deep learning model for competing risks. In: 2021 IEEE International Conference on Bioinformatics and Biomedicine (BIBM), pp. 933–938 (2021). https://doi.org/10.1109/ BI...
2021
-
[24]
In: Pro- ceedings of the 31st ACM International Conference on Information & Knowledge Management (CIKM), pp
Huang, Z., Zhang, A., Hu, Y., Wang, L., Sun, J., Chen, Y.: TransformerJM: A Transformer- based joint model for multivariate longitudinal data and competing risks survival. In: Pro- ceedings of the 31st ACM International Conference on Information & Knowledge Management (CIKM), ...
2022 doi
-
[25]
In: Proceedings of the 2023 SIAM International Conference on Data Mining (SDM), pp
Zhang, H., Wu, Z., Zhao, J.: CAT-Surv: A categorical time-discretization approach for survival analysis. In: Proceedings of the 2023 SIAM International Conference on Data Mining (SDM), pp. 406–414 (2023). https://doi.org/10.1137/1.9781611977653.ch46 . SIAM
2023 doi
-
[26]
Artificial Intelligence Review57(3), 65 (2024) https://doi.org/10.1007/ s10462-023-10681-3
Wiegrebe, S., Kopper, P., Sonabend, R., Bender, A., R¨ ugamer, D.: Deep learning for sur- vival analysis: a review. Artificial Intelligence Review57(3), 65 (2024) https://doi.org/10.1007/ s10462-023-10681-3
2024
-
[27]
JAMA323(4), 305–306 (2020) https://doi.org/10.1001/jama.2019.20866
Beam, A.L., Manrai, A.K., Ghassemi, M.: Challenges to the reproducibility of machine learning models in health care. JAMA323(4), 305–306 (2020) https://doi.org/10.1001/jama.2019.20866
2020 doi
-
[28]
Science334(6060), 1226–1227 (2011) https://doi.org/10.1126/science.1213847
Peng, R.D.: Reproducible research in computational science. Science334(6060), 1226–1227 (2011) https://doi.org/10.1126/science.1213847
2011 doi
-
[29]
NPJ Digital Medicine3(1), 41 (2020) https: //doi.org/10.1038/s41746-020-0253-3
Sendak, M.P., Gao, M., Brajer, N., Balu, S.: Presenting machine learning model information to clinical end users with model facts labels. NPJ Digital Medicine3(1), 41 (2020) https: //doi.org/10.1038/s41746-020-0253-3
2020 doi
-
[30]
Scientific Data10(1), 1 (2023) https://doi.org/10
Johnson, A.E.W., Bulgarelli, L., Shen, L., Gayles, A., Shammout, A., Horng, S., Pollard, T.J., Hao, S., Moody, B., Gow, B., Lehman, L.-w.H., Celi, L.A., Mark, R.G.: MIMIC-IV, a freely accessible electronic health record dataset. Scientific Data10(1), 1 (2023) https://doi.org/1...
2023
-
[31]
In: Machine Learning for Health, pp
Gupta, M., Gallamoza, B., Cutrona, N., Dhakal, P., Poulain, R., Beheshti, R.: An extensive data processing pipeline for MIMIC-IV. In: Machine Learning for Health, pp. 311–325. PMLR, ??? (2022)
2022
-
[32]
McDermott, M.B.A., Nestor, B., Kim, E., Zhang, W., Goldenberg, A., Szolovits, P., Ghassemi, M.: Comprehensive comparative study of multi-label classification methods (2021)
2021
-
[33]
Scientific Data5, 180178 (2018) https://doi.org/10.1038/sdata.2018.178
Pollard, T.J., Johnson, A.E.W., Raffa, J.D., Celi, L.A., Mark, R.G., Badawi, O.: The eICU Col- laborative Research Database, a freely available multi-centre database for critical care research. Scientific Data5, 180178 (2018) https://doi.org/10.1038/sdata.2018.178
2018 doi
-
[34]
Scientific 21 Reports9(1), 15665 (2019) https://doi.org/10.1038/s41598-019-51810-8
Kim, Y., Shachar, S.S., Gayvert, K., Li, P.P., Gannavarapu, A., Lempicki, M., Claassen, J., Castro, M., Steinherz, P., Lis, R., Levy-Lahad, E., Meiner, V., Daly, M.B., Tischkowitz, M., Offit, K., Robson, M., Domchek, S.M., Walsh, M.F., Levine, D.A.: Development and validation ...
2019 doi
-
[35]
The Lancet Respiratory Medicine3(1), 42–52 (2015) https://doi.org/ 10.1016/S2213-2600(14)70239-5
Pirracchio, R., Petersen, M.L., Carone, M., Rigon, M.R., Chevret, S., Laan, M.J.: Mortal- ity prediction in intensive care units with the Super ICU Learner Algorithm (SICULA): A population-based study. The Lancet Respiratory Medicine3(1), 42–52 (2015) https://doi.org/ 10.1016/...
2015 doi
-
[36]
BMC Medical Research Methodology 14, 137 (2014) https://doi.org/10.1186/1471-2288-14-137
Ploeg, T., Austin, P.C., Steyerberg, E.W.: Modern modelling techniques are data hungry: A simulation study for predicting dichotomous endpoints. BMC Medical Research Methodology 14, 137 (2014) https://doi.org/10.1186/1471-2288-14-137
2014 doi
-
[37]
IEEE Journal of Biomedical and Health Informatics24(11), 3268–3275 (2020) https://doi.org/10.1109/JBHI.2020.2984931
Darabi, S., Kachuee, M., Fazeli, S., Sartipi, M.: TAPER: Time-aware patient EHR rep- resentation. IEEE Journal of Biomedical and Health Informatics24(11), 3268–3275 (2020) https://doi.org/10.1109/JBHI.2020.2984931
2020 doi
-
[38]
BMJ361, 1479 (2018) https: //doi.org/10.1136/bmj.k1479
Agniel, D., Kohane, I.S., Weber, G.M.: Biases in electronic health record data due to processes within the healthcare system: Retrospective observational study. BMJ361, 1479 (2018) https: //doi.org/10.1136/bmj.k1479
2018 doi
-
[39]
Nature Medicine26(3), 364–373 (2020) https://doi.org/10.1038/s41591-020-0789-4
Hyland, S.L., Faltys, M., H¨ user, M., Lyu, X., Gumbsch, T., Esteban, C., Bock, C., Horn, M., Moor, M., Rieck, B., Zimmermann, M., Bodenham, D., Borgwardt, K., R¨ atsch, G., Merz, T.M.: Early prediction of circulatory failure in the intensive care unit using machine learning. ...
2020 doi
-
[40]
Journal of Biomedical Informatics83, 112–134 (2018) https://doi.org/10
Purushotham, S., Meng, C., Che, Z., Liu, Y.: Benchmarking deep learning models on large healthcare datasets. Journal of Biomedical Informatics83, 112–134 (2018) https://doi.org/10. 1016/j.jbi.2018.04.007
2018
-
[41]
NPJ Digital Medicine4(1), 86 (2021) https://doi.org/10.1038/s41746-021-00455-y
Rasmy, L., Xiang, Y., Xie, Z., Tao, C., Zhi, D.: Med-BERT: Pretrained contextualized em- beddings on large-scale structured electronic health records for disease prediction. NPJ Digital Medicine4(1), 86 (2021) https://doi.org/10.1038/s41746-021-00455-y
2021 doi
-
[42]
In: Advances in Neural Information Processing Systems, pp
Choi, E., Bahadori, M.T., Sun, J., Kulas, J., Schuetz, A., Stewart, W.: RETAIN: An inter- pretable predictive model for healthcare using reverse time attention mechanism. In: Advances in Neural Information Processing Systems, pp. 3504–3512 (2016)
2016
-
[43]
Huang, K., Altosaar, J., Ranganath, R.: ClinicalBERT: Modeling clinical notes and predicting hospital readmission (2019)
2019
-
[44]
In: Proceedings of the 2nd Clinical Natural Language Processing Workshop, pp
Alsentzer, E., Murphy, J., Boag, W., Weng, W.-H., Jindi, D., Naumann, T., McDermott, M.: Publicly available clinical BERT embeddings. In: Proceedings of the 2nd Clinical Natural Language Processing Workshop, pp. 72–78 (2019)
2019
-
[45]
Circulation101(23), 215–220 (2000)
Goldberger, A.L., Amaral, L.A.N., Glass, L., Hausdorff, J.M., Ivanov, P.C., Mark, R.G., Mietus, J.E., Moody, G.B., Peng, C.-K., Stanley, H.E.: PhysioBank, PhysioToolkit, and PhysioNet: Components of a new research resource for complex physiologic signals. Circulation101(23), 2...
2000
-
[46]
In: 2020 IEEE-EMBS International Con- ference on Biomedical and Health Informatics (BHI), pp
Sheikhalishahi, S., K¨ ok, I., Luo, Y., Peelen, L., Slooter, A.J.C., G, K.C.: Benchmarking critical care datasets: A comparison of eICU and MIMIC-III. In: 2020 IEEE-EMBS International Con- ference on Biomedical and Health Informatics (BHI), pp. 1–4 (2020). https://doi.org/10.1...
2020
-
[47]
In: International Conference on Machine Learning, pp
Futoma, J., Hariharan, S., Heller, K.: Learning to detect sepsis with a multitask Gaussian process RNN classifier. In: International Conference on Machine Learning, pp. 1174–1182 (2017)
2017
-
[48]
Scientific Data12(1), 1094 (2025) https://doi.org/ 10.1038/s41597-025-05419-5 22
Kansal, A., Chen, E., Jin, B.T., Rajpurkar, P., Kim, D.A.: MC-MED, multimodal clinical monitoring in the emergency department. Scientific Data12(1), 1094 (2025) https://doi.org/ 10.1038/s41597-025-05419-5 22
2025 doi
-
[49]
PhysioNet (2025)
Kansal, A., Chen, E., Jin, T., Rajpurkar, P., Kim, D.: MC-MED, multimodal clinical monitor- ing in the emergency department (version 1.0.1). PhysioNet (2025). https://doi.org/10.13026/ wvyw-g663 . https://physionet.org/content/mc-med/1.0.1/
2025
-
[50]
Journal of the American Medical Informatics Association20(3), 494–502 (2013) https://doi
Malin, B., El Emam, K.: Detecting identity disclosure in high-dimensional (biomedical) data. Journal of the American Medical Informatics Association20(3), 494–502 (2013) https://doi. org/10.1136/amiajnl-2012-001031
2013 doi
-
[51]
Critical Care18(4), 421 (2014) https://doi.org/10.1186/cc13993
Verburg, I.W.M., Atlen, J., Keizer, N.F., Escoredo-Luks, N.L., H¨ okby, N., Jonge, E., Meurs, M.M.: Impact of arbitrary ICU admission and discharge criteria on outcome analysis. Critical Care18(4), 421 (2014) https://doi.org/10.1186/cc13993
2014 doi
-
[52]
Journal of Clinical Anesthesia 13(8), 563–568 (2001) https://doi.org/10.1016/s0952-8180(01)00344-9
Rosenberg, A.L., Ho, V., Lema, J.V., O’Brien, J.F.: Severity of illness and the timing of admis- sion and discharge for patients who die in the intensive care unit. Journal of Clinical Anesthesia 13(8), 563–568 (2001) https://doi.org/10.1016/s0952-8180(01)00344-9
2001 doi
-
[53]
Anaesthesia54(6), 558–564 (1999) https://doi.org/10.1046/j.1365-2044.1999.00843.x
Goldhill, D.R., Sumner, A.: Mortality and length of stay in intensive care: a comparison of three units. Anaesthesia54(6), 558–564 (1999) https://doi.org/10.1046/j.1365-2044.1999.00843.x
1999 doi
-
[54]
JMIR Medical Informatics4(3), 28 (2016) https://doi.org/10.2196/medinform.5909
Desautels, T., Calvert, J., Hoffman, J., Jay, M., Kerem, Y., Shieh, L., Shimabukuro, D., Chet- tipally, U., Feldman, M.D., Barton, C., Wales, D.J., Das, R.: Prediction of sepsis in the intensive care unit with minimal electronic health record data: A machine learning approach....
2016 doi
-
[55]
Critical Care Medicine 34(11), 2735–2741 (2006) https://doi.org/10.1097/01.CCM.0000240974.74548.D8
Kramer, A.A., Zimmerman, J.E.: A comparison of three methods for predicting discharge status: The importance of both patient case mix and provider-level effects. Critical Care Medicine 34(11), 2735–2741 (2006) https://doi.org/10.1097/01.CCM.0000240974.74548.D8
2006 doi
-
[56]
Intensive Care Medicine38(10), 1654–1661 (2012) https://doi.org/10.1007/ s00134-012-2629-6
Fuchs, L., Chronaki, C.E., Park, S., Novack, V., Baumfeld, Y., Scott, D., McLennan, S., Tal- mor, D., Celi, L.: ICU admission characteristics and mortality rates among elderly and very elderly patients. Intensive Care Medicine38(10), 1654–1661 (2012) https://doi.org/10.1007/ s...
2012
-
[57]
Epidemiology20(4), 555–561 (2009) https://doi.org/10.1097/EDE.0b013e3181a39056
Wolbers, M., Koller, M.T., Witteman, J.C.M., Steyerberg, E.W.: Prognostic models with com- peting risks: Methods and application to coronary risk prediction. Epidemiology20(4), 555–561 (2009) https://doi.org/10.1097/EDE.0b013e3181a39056
2009 doi
-
[58]
Medical Care48(6 Suppl), 96–105 (2010) https://doi.org/10.1097/MLR
Varadhan, R., Weiss, C.O., Segal, J.B., Wu, A.W., Scharfstein, D., Boyd, C.: Evaluating health outcomes in the presence of competing risks: A review of statistical methods and clinical applications. Medical Care48(6 Suppl), 96–105 (2010) https://doi.org/10.1097/MLR. 0b013e3181d99107
2010 doi
-
[59]
Orthopedics36(1), 15–21 (2013) https://doi.org/10.3928/01477447-20121217-08
Cushing, T.A., Bruce, S.E., Kaimraj, M., Gaskill, T.R., Cripe, M.J., Gaski, G.E., Johnson, W.J., Salyers, E., Zody, R.D., Siders, C.A.,et al.: Comparison of pediatric and adult open-fracture management. Orthopedics36(1), 15–21 (2013) https://doi.org/10.3928/01477447-20121217-08
2013 doi
-
[60]
Current Pediatric Reviews11(3), 195–200 (2015) https: //doi.org/10.2174/1573396311666150722104113
Martin, B., Kling, P., Gonzalez, R., Sgromolo, T., Pil-Kim, C.: Pediatric-specific metabolic and hematologic responses to trauma. Current Pediatric Reviews11(3), 195–200 (2015) https: //doi.org/10.2174/1573396311666150722104113
2015 doi
-
[61]
Annals of Emergency Medicine58(1), 33–40 (2011) https://doi.org/10.1016/j.annemergmed.2010.08.040
Welch, S.J., Asplin, B.R., Stone-Griffith, S., Davidson, S.J., Augustine, J., Schuur, J.: Emer- gency department operational metrics, measures and definitions: Results of the Second Performance Measures and Benchmarking Summit. Annals of Emergency Medicine58(1), 33–40 (2011) h...
2011 doi
-
[62]
Academic Emergency Medicine18(8), 848–855 (2011) https://doi.org/10
Rowe, B.H., McRae, A.D., Yaghoubi, M., Forgie, P.J., Shao, J., Johnson, C., Holroyd, B.R., Yoon, P., O’Brien, D.M., Oh, P.: Characteristics of patients who leave emergency departments without being seen. Academic Emergency Medicine18(8), 848–855 (2011) https://doi.org/10. 1111...
2011
-
[63]
John Wiley & Sons, ??? (2006)
Pintilie, M.: Analysing Competing Risks Data. John Wiley & Sons, ??? (2006). https://doi. org/10.1002/0470870716
2006 doi
-
[64]
Journal of the American Statistical Association94(446), 496–509 (1999) https://doi.org/ 10.1080/01621459.1999.10474144
Fine, J.P., Gray, R.J.: A proportional hazards model for the subdistribution of a competing risk. Journal of the American Statistical Association94(446), 496–509 (1999) https://doi.org/ 10.1080/01621459.1999.10474144
1999 doi
-
[65]
In: International Conference on Artificial Intelligence and Statistics (2024)
Mesinovic, M., Cui, C., Zenati, M.A., Cannesson, M., Burdjalov, V.K.: MM-GraphSurv: Multi- modal graph-based survival analysis for critical care. In: International Conference on Artificial Intelligence and Statistics (2024)
2024
-
[66]
Critical Care Medicine24(4), 650–654 (1996) https://doi
Vincent, J.-L., De Mendon¸ ca, A., Cantraine, F., Moreno, R., Takala, J., Suter, P.M., V-A, D.B., Thijs, L.G., Sprung, C.L.: The SOFA (sepsis-related organ failure assessment) score to describe organ dysfunction/failure. Critical Care Medicine24(4), 650–654 (1996) https://doi....
1996 doi
-
[67]
Critical Care Medicine34(5), 1297–1310 (2006) https://doi.org/10.1097/01.CCM
Zimmerman, J.E., Kramer, A.A., McNair, D.S., Malila, F.M.: Acute Physiology and Chronic Health Evaluation (APACHE) IV: Hospital mortality assessment for today’s critically ill patients. Critical Care Medicine34(5), 1297–1310 (2006) https://doi.org/10.1097/01.CCM. 0000215112.84523.F0
2006 doi
-
[68]
Neural Computation9(8), 1735–1780 (1997) https://doi.org/10.1162/neco.1997.9.8.1735
Hochreiter, S., Schmidhuber, J.: Long short-term memory. Neural Computation9(8), 1735–1780 (1997) https://doi.org/10.1162/neco.1997.9.8.1735
1997 doi
-
[69]
Bai, S., Kolter, J.Z., Koltun, V.: An empirical evaluation of generic convolutional and recurrent networks for sequence modeling (2018)
2018
-
[70]
In: 2017 International Conference on Pervasive Computing (ICPC), pp
Potdar, K., Pardawala, T.S., Pai, C.D.: A comparative study of categorical data encoding techniques for supervised machine learning. In: 2017 International Conference on Pervasive Computing (ICPC), pp. 1–7 (2017). https://doi.org/10.1109/PERVASIVE.2017.83pervasive. 2017.83 . IEEE
2017 doi
-
[71]
Science366(6464), 447–453 (2019) https://doi.org/ 10.1126/science.aax2342
Obermeyer, Z., Powers, B., Vogeli, C., Mullainathan, S.: Dissecting racial bias in an algorithm used to manage the health of populations. Science366(6464), 447–453 (2019) https://doi.org/ 10.1126/science.aax2342
2019 doi
-
[72]
Annals of Internal Medicine169(12), 866–872 (2018) https: //doi.org/10.7326/M18-1990
Rajkomar, A., Hardt, M., Howell, M.D., Corrado, G., Chin, M.H.: Ensuring fairness in machine learning to advance health equity. Annals of Internal Medicine169(12), 866–872 (2018) https: //doi.org/10.7326/M18-1990
2018 doi
-
[73]
ACM SIGKDD Explorations Newsletter6(1), 20–29 (2004) https://doi.org/10.1145/1007730.1007735
Batista, G.E.A.P.A., Prati, R.C., Monard, M.C.: A study of the behavior of several methods for balancing machine learning training data. ACM SIGKDD Explorations Newsletter6(1), 20–29 (2004) https://doi.org/10.1145/1007730.1007735
2004 doi
-
[74]
Intelligent Data Analysis6(5), 429–449 (2002) https://doi.org/10.3233/IDA-2002-6504
Japkowicz, N., Stephen, S.: The class imbalance problem: A systematic study. Intelligent Data Analysis6(5), 429–449 (2002) https://doi.org/10.3233/IDA-2002-6504
2002 doi
-
[75]
Journal of Machine Learning Research3, 1157–1182 (2003)
Guyon, I., Elisseeff, A.: An introduction to variable and feature selection. Journal of Machine Learning Research3, 1157–1182 (2003)
2003
-
[76]
Bioinformatics23(19), 2507–2517 (2007) https://doi.org/10.1093/bioinformatics/btm344
Saeys, Y., Inza, I., Larra˜ naga, P.: A review of feature selection techniques in bioinformatics. Bioinformatics23(19), 2507–2517 (2007) https://doi.org/10.1093/bioinformatics/btm344
2007 doi
-
[77]
Proceedings of the National Academy of Sciences99(10), 6562–6566 (2002) https://doi.org/10.1073/pnas.102102699
Ambroise, C., McLachlan, G.J.: Selection bias in gene extraction on the basis of microarray gene- expression data. Proceedings of the National Academy of Sciences99(10), 6562–6566 (2002) https://doi.org/10.1073/pnas.102102699
2002 doi
-
[78]
Journal of Machine Learning Research3, 1371–1382 (2003)
Reunanen, J.: Overfitting in making comparisons between variable selection methods. Journal of Machine Learning Research3, 1371–1382 (2003)
2003
-
[79]
Science 354(6317), 1240–1241 (2016) https://doi.org/10.1126/science.aah6168
Stodden, V., McNutt, M., Bailey, D.H., Deelman, E., Gil, Y., Hanson, B., Heroux, M.A., 24 Ioannidis, J.P.A., Taufer, M.: Enhancing reproducibility for computational methods. Science 354(6317), 1240–1241 (2016) https://doi.org/10.1126/science.aah6168
2016 doi
-
[80]
Nature Reviews Genetics13(6), 395–405 (2012) https://doi.org/ 10.1038/nrg3208
Jensen, P.B., Jensen, L.J., Brunak, S.: Mining electronic health records: Towards better research applications and clinical care. Nature Reviews Genetics13(6), 395–405 (2012) https://doi.org/ 10.1038/nrg3208
2012 doi
-
[81]
Nature585(7825), 357–362 (2020) https://doi.org/10.1038/s41586-020-2649-2
Harris, C.R., Millman, K.J., Walt, S.J., Gommers, R., Virtanen, P., Cournapeau, D., Wieser, E., Taylor, J., Berg, S., Smith, N.J., Kern, R., Picus, M., Hoyer, S., Kerkwijk, M.H., Brett, M., Haldane, A., R´ ıo, J.F., Wiebe, M., Peterson, P., G´ erard-Marchant, P., Sheppard, K.,...
2020 doi
-
[82]
In: Proceedings of the 9th Python in Science Conference, pp
McKinney, W.: Data structures for statistical computing in Python. In: Proceedings of the 9th Python in Science Conference, pp. 56–61 (2010). https://doi.org/10.25080/ Majora-92bf1922-00a
2010
-
[83]
Journal of Machine Learning Research12, 2825–2830 (2011)
Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., Vanderplas, J., Passos, A., Cournapeau, D., Brucher, M., Perrot, M., Duchesnay, ´E.: Scikit-learn: Machine learning in Python. Journal of Mac...
2011
-
[84]
In: Advances in Neural Information Processing Systems, pp
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Kopf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., Chintala, S.: PyTorch: An imperativ...
2019
-
[85]
ACM SIGOPS Operating Systems Review49(1), 71–79 (2015) https://doi.org/10.1145/2723872.2723882
Boettiger, C.: An introduction to Docker for reproducible research. ACM SIGOPS Operating Systems Review49(1), 71–79 (2015) https://doi.org/10.1145/2723872.2723882
2015 doi
-
[86]
PLoS Computational Biology9(10), 1003285 (2013) https://doi.org/10.1371/ journal.pcbi.1003285
Sandve, G.K., Nekrutenko, A., Taylor, J., Hovig, E.: Ten simple rules for reproducible compu- tational research. PLoS Computational Biology9(10), 1003285 (2013) https://doi.org/10.1371/ journal.pcbi.1003285
2013
-
[87]
PLoS Computational Biology13(6), 1005510 (2017) https: //doi.org/10.1371/journal.pcbi.1005510
Wilson, G., Bryan, J., Cranston, K., Kitzes, J., Nederbragt, L., Teal, T.K.: Good enough practices in scientific computing. PLoS Computational Biology13(6), 1005510 (2017) https: //doi.org/10.1371/journal.pcbi.1005510
2017 doi
-
[88]
PeerJ7, 6257 (2019) https://doi.org/10.7717/peerj.6257
Gensheimer, M.F., Narasimhan, B.: A scalable discrete-time survival model for neural networks. PeerJ7, 6257 (2019) https://doi.org/10.7717/peerj.6257
2019 doi
-
[89]
In: Neural Networks: Tricks of the Trade, pp
LeCun, Y., Bottou, L., Orr, G.B., M¨ uller, K.-R.: Efficient backprop. In: Neural Networks: Tricks of the Trade, pp. 9–48. Springer, ??? (2012). https://doi.org/10.1007/978-3-642-35289-8 3
2012 doi
-
[90]
Journal of Statistical Software59(10), 1–23 (2014) https://doi.org/ 10.18637/jss.v059.i10
Wickham, H.: Tidy data. Journal of Statistical Software59(10), 1–23 (2014) https://doi.org/ 10.18637/jss.v059.i10
2014 doi
-
[91]
Statistics and Computing
Wilkinson, L.: The Grammar of Graphics. Statistics and Computing. Springer, ??? (2005). https://doi.org/10.1007/0-387-28695-0
2005 doi
-
[92]
Journal of the American Medical Informatics Association23(e1), 139–146 (2016) https://doi.org/10.1093/jamia/ocv145
Kahn, M.G., Callahan, T.J., Barnard, J., Bauck, A.E., Brown, J., Davidson, B.N., Estiri, H., Goecks, J., H collaborators, P.: A pragmatic data quality assessment framework for research- quality data. Journal of the American Medical Informatics Association23(e1), 139–146 (2016)...
2016 doi
-
[93]
Journal of the American Medical Informatics Association20(1), 144–151 (2013) https://doi.org/10.1136/amiajnl-2011-000681 25
Weiskopf, N.G., Weng, C.: Methods and dimensions of electronic health record data quality assessment: Enabling reuse for clinical research. Journal of the American Medical Informatics Association20(1), 144–151 (2013) https://doi.org/10.1136/amiajnl-2011-000681 25
2013 doi
-
[94]
Chest100(6), 1619–1636 (1991) https://doi.org/10.1378/chest.100.6.1619
Knaus, W.A., Wagner, D.P., Draper, E.A., Zimmerman, J.E., Bergner, M., Bastos, P.G., Sirio, C.A., Murphy, D.J., Lotring, T., Damiano, A., Harrell Jr, F.E.: The APACHE III prognostic system: Risk prediction of hospital mortality for critically ill hospitalized adults. Chest100(...
1991 doi
-
[95]
Critical Care Medicine44(2), 368–374 (2016) https://doi.org/10.1097/CCM.0000000000001571
Churpek, M.M., Yuen, T.C., Winslow, C., Meltzer, D.O., Kattan, M.W., Edelson, D.P.: Multicenter comparison of machine learning methods and conventional regression for pre- dicting clinical deterioration on the wards. Critical Care Medicine44(2), 368–374 (2016) https://doi.org/...
2016 doi
-
[96]
Computer25(10), 40–51 (1992) https://doi.org/10
Meyer, B.: Applying ”design by contract”. Computer25(10), 40–51 (1992) https://doi.org/10. 1109/2.161109
1992
-
[97]
Sculley, D., Holt, G., Golovin, D., Davydov, E., Phillips, T., Ebner, D., Chaudhary, V., Young, M., Crespo, J.-F., Dennison, D.: Hidden technical debt in machine learning systems, 2503–2511 (2015)
2015
-
[98]
Paleyes, A., Urma, R.-G., Lawrence, N.D.: Challenges in deploying machine learning: A survey of case studies (2020)
2020
-
[99]
Medical Care50, 21–29 (2012) https://doi.org/10.1097/MLR.0b013e318257dd67
Kahn, M.G., Raebel, M.A., Glanz, J.M., Riedlinger, K., Steiner, J.F.: A pragmatic framework for single-site and multisite data quality assessment in electronic health record-based clinical research. Medical Care50, 21–29 (2012) https://doi.org/10.1097/MLR.0b013e318257dd67
2012 doi
-
[100]
Springer, ??? (2003)
Klein, J.P., Moeschberger, M.L.: Survival Analysis: Techniques for Censored and Truncated Data, 2nd edn. Springer, ??? (2003). https://doi.org/10.1007/b97377
2003 doi
-
[101]
MIT Press, ??? (2009)
Qui˜ nonero-Candela, J., Sugiyama, M., Schwaighofer, A., Lawrence, N.D.: Dataset Shift in Machine Learning. MIT Press, ??? (2009)
2009
-
[102]
Pattern Recognition45(1), 521–530 (2012) https://doi
Moreno-Torres, J.G., Raeder, T., Alaiz-Rodr´ ıguez, R., Chawla, N.V., Herrera, F.: A unifying view on dataset shift in classification. Pattern Recognition45(1), 521–530 (2012) https://doi. org/10.1016/j.patcog.2011.06.019
2012 doi
-
[103]
BMJ338, 2393 (2009) https://doi.org/10.1136/bmj.b2393
Sterne, J.A.C., White, I.R., Carlin, J.B., Spratt, M., Royston, P., Kenward, M.G., Wood, A.M., Carpenter, J.R.: Multiple imputation for missing data in epidemiological and clinical research: Potential and pitfalls. BMJ338, 2393 (2009) https://doi.org/10.1136/bmj.b2393
2009 doi
-
[104]
CRC Press, ??? (2018)
Buuren, S.: Flexible Imputation of Missing Data, 2nd edn. CRC Press, ??? (2018). https://doi. org/10.1201/9780429492259
2018 doi
-
[105]
New England Journal of Medicine385(3), 283–286 (2021) https://doi.org/10.1056/NEJMc2104626
Finlayson, S.G., Subbaswamy, A., Singh, K., Bowers, J., Kupke, A., Zittrain, J., Kohane, I.S., Saria, S.: The clinician and dataset shift in artificial intelligence. New England Journal of Medicine385(3), 283–286 (2021) https://doi.org/10.1056/NEJMc2104626
2021 doi
-
[106]
In: International Conference on Machine Learning (ICML), pp
Subbaswamy, A., Saria, S.: Preventing dataset shift from hurting model performance. In: International Conference on Machine Learning (ICML), pp. 6031–6040 (2019). PMLR
2019
-
[107]
Technical Report RFC 1952, RFC Editor (1996)
Deutsch, P.: GZIP file format specification version 4.3. Technical Report RFC 1952, RFC Editor (1996). https://doi.org/10.17487/RFC1952
1952 doi
-
[108]
Pearson, ??? (2014)
Tanenbaum, A.S., Bos, H.: Modern Operating Systems, 4th edn. Pearson, ??? (2014)
2014
-
[109]
ACM Transactions on Knowledge Discovery from Data6(4), 15 (2012) https://doi.org/10.1145/2382577.2382579
Kaufman, S., Rosset, S., Perlich, C., Stitelman, O.: Leakage in data mining: Formulation, de- tection, and avoidance. ACM Transactions on Knowledge Discovery from Data6(4), 15 (2012) https://doi.org/10.1145/2382577.2382579
2012 doi
-
[110]
Patterns4(9), 100804 (2023) https://doi.org/10.1016/j.patter.2023.100804 26
Kapoor, S., Narayanan, A.: Leakage and the reproducibility crisis in machine-learning-based science. Patterns4(9), 100804 (2023) https://doi.org/10.1016/j.patter.2023.100804 26
2023 doi
-
[111]
Journal of Biomedical Informatics 46(5), 837–848 (2013) https://doi.org/10.1016/j.jbi.2013.06.011
Rothman, M.J., Rothman, S.I., Beals IV, J.: Development and validation of a continuous meas- ure of patient condition using the Electronic Medical Record. Journal of Biomedical Informatics 46(5), 837–848 (2013) https://doi.org/10.1016/j.jbi.2013.06.011
2013 doi
-
[112]
Circulation138(20), 618–651 (2018) https://doi.org/10.1161/CIR.0000000000000617 Acknowledgements.The authors would like to thank Max Buhlan for assistance with Figure 2
Thygesen, K., Alpert, J.S., Jaffe, A.S., Chaitman, B.R., Bax, J.J., Morrow, D.A., White, H.D., Myocardial Infarction, E.G.: Fourth universal definition of myocardial infarction (2018). Circulation138(20), 618–651 (2018) https://doi.org/10.1161/CIR.0000000000000617 Acknowledgem...
2018 doi
Reviewed May 17, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.