REVIEW 2 major objections 7 minor 39 references
Surrogate Transformer Opens the Black Box of EHR Foundation Models
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · glm-5.2
2026-07-08 14:34 UTC pith:EMWBO7MA
load-bearing objection Surrogate outperforms the FEMR on ground-truth metrics (Table 5), indicating it learns a different decision boundary; without a prediction-agreement fidelity metric, SHAP explanations cannot be validly attributed to the FEMR. the 2 major comments →
X-FEMR: A Token-level Explainable Approach for Electronic Health Records Foundation Models using Transformer-based Models
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper's central discovery is that a Transformer-based surrogate model can approximate a FEMR's predictions closely enough that SHAP-based token attribution on the surrogate yields clinically meaningful explanations. The most frequently identified important tokens—heart rate, systolic blood pressure, body temperature, respiratory rate, and oxygen saturation—are established clinical predictors for the two prediction tasks studied, and the clinical validated events ratio of ~0.3 quantifies this alignment.
What carries the argument
Transformer-based surrogate model trained on FEMR input-output pairs, SHAP token attribution applied to the surrogate, and the clinical validated events ratio (R) defined as the proportion of SHAP-identified tokens that match clinically validated features.
Load-bearing premise
The paper assumes that token importance derived via SHAP from the surrogate model reflects the decision logic of the original FEMR. Since the surrogate is an approximation with measurably different performance characteristics, it could learn a different decision boundary—one that relies on spurious correlations that happen to achieve similar accuracy—making the extracted explanations invalid for the original model even if they align with clinical knowledge.
What would settle it
If the surrogate model's decision boundary diverges from the FEMR's in a way that produces similar predictions but via different features, the SHAP-based token importances would not explain the FEMR at all, and the clinical alignment would be a property of the surrogate rather than the original model.
If this is right
- If the surrogate-fidelity assumption holds, clinicians could use token-level explanations to audit FEMR predictions for individual patients, identifying which specific vital signs or lab values drove a given risk score.
- The clinical validated events ratio provides a reusable quantitative metric for evaluating whether any future explainability method for EHR models produces clinically grounded explanations, not just statistically faithful ones.
- The hard-label vs. soft-label supervision comparison suggests that different training regimes for the surrogate expose different aspects of the FEMR's decision logic, which could be used diagnostically to detect when a model relies on spurious correlations versus clinically meaningful features.
- The approach could be extended to other FEMRs as they become available, enabling comparative explainability studies across different foundation model architectures for EHR data.
Where Pith is reading between the lines
- The clinical validated events ratio of 0.3 means that roughly 70% of SHAP-identified important tokens are not clinically validated features. Whether this represents clinically irrelevant noise, spurious correlations, or clinically meaningful features not yet formally validated is not addressed but is critical for trusting the explanations.
- If the surrogate model achieves higher AUROC than the original FEMR (as observed in Table 5 for hard labels), it may have learned a different decision boundary rather than faithfully imitating the FEMR, which would mean the explanations reflect the surrogate's logic rather than the FEMR's.
- The event-level tokenization combines code, numeric value, text, and time delta into a single token, which means SHAP attributions cannot distinguish whether the numeric value or the event type is driving importance—a granularity limitation that future work could address with finer-grained tokenization.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes X-FEMR, a token-level explainability framework for Foundation Models for Electronic Health Records (FEMRs). The approach trains a Transformer-based surrogate model on input-output pairs from CLMBR-T-Base (the target FEMR) across two clinical prediction tasks (length-of-stay and ICU transfer), then applies SHAP to the surrogate to extract token-level importance scores. A novel 'clinical validated events ratio' (R) is introduced to quantify how well the identified tokens correspond to clinically validated features. The authors report that the surrogate achieves comparable or superior predictive performance to the FEMR (Tables 4–5) and that the top SHAP-identified tokens align with known clinical predictors (Tables 6–7).
Significance. Token-level explainability for structured EHR foundation models is a genuinely underexplored area, and the paper targets a real gap. The use of a Transformer-based surrogate that preserves temporal dynamics is a reasonable methodological choice over simpler surrogates (e.g., linear models). The introduction of a quantitative clinical alignment metric, while currently flawed in design, is a step toward rigorous evaluation of explanations. The pipeline is applied to a publicly available FEMR (CLMBR-T-Base) on a public benchmark (EHRSHOT), which aids reproducibility. However, the central claim—that SHAP attributions on the surrogate validly explain the FEMR's decision logic—is not adequately supported by the current evidence, primarily because surrogate-to-FEMR fidelity is never directly measured.
major comments (2)
- §5.2, Tables 4–5: The surrogate model outperforms CLMBR-T-Base on ground-truth labels (e.g., LOS hard-label AUROC 0.7637 vs. 0.7046; AUPRC 0.4434 vs. 0.4173). The paper frames this as evidence that the surrogate 'closely approximates' the FEMR, but a surrogate that exceeds the teacher's performance against ground truth is not necessarily faithful to the teacher's decision boundary—it may be learning a different (and better) function. The paper never reports prediction-agreement metrics between the surrogate and CLMBR-T-Base (e.g., agreement rate, KL divergence on output distributions). Without such a fidelity metric, the claim that SHAP attributions on the surrogate reflect the FEMR's reasoning is unsupported. This is load-bearing for the paper's central contribution. The authors should report direct surrogate–FEMR agreement and discuss whether the observed performance gap undermines the
- §3.4, Eqs. (5)–(6): The clinical validated events ratio R is defined as the count of clinically validated token occurrences divided by total token occurrences among SHAP-attributed tokens. However, Eq. (5) counts raw occurrence frequency of token f_j in the dataset, not the frequency with which f_j receives non-trivial SHAP importance. The text does not specify a SHAP importance threshold for inclusion in F. As written, R measures the prevalence of clinically validated codes in the data, not their attribution-weighted importance. Additionally, no baseline (e.g., the ratio of clinically validated codes among all tokens in the dataset) is reported, making R = 0.31 difficult to interpret—if the base rate of clinically validated tokens is also ~0.30, the metric carries no signal. The metric should be redefined to weight by SHAP magnitude and compared against a null baseline.
minor comments (7)
- Table 5, ICU transfer soft-label row: Precision, Recall, and F1 are all 0.0000, indicating the model makes no positive predictions. This row is effectively unusable and should be flagged or removed with explanation.
- §3.4: The set F is described as 'the set of all tokens from SHAP analysis,' but it is unclear whether this means all tokens that receive any non-zero SHAP value, or tokens above some threshold. This ambiguity affects the interpretation of R.
- §4.2: The surrogate model architecture (3 layers, 128 hidden, 4 heads) is much smaller than CLMBR-T-Base (141M parameters). The potential capacity mismatch and its effect on fidelity should be briefly discussed.
- §5.3: The paper states that 'numerical values and time delta contributed most significantly' but does not present SHAP magnitude statistics to support this claim. Summary statistics or a figure showing SHAP value distributions per feature type would strengthen this.
- Table 2: The Hospital Frailty Risk Score is listed as '109 specific ICD-10 diagnostic codes' but the codes themselves are not enumerated. A supplementary table would improve reproducibility.
- §1: The phrase 'the first token-level explainability approach for FEMRs' is a strong novelty claim. The authors should verify that no concurrent or prior work on token-level attribution for structured EHR models (e.g., Med-BERT or ETHOS) exists.
- Eq. (4): The notation mixes subscripts (e_t, E_pos(t)) and the gated injection g(·) is not fully defined. Specifying the gate's functional form would improve readability.
Simulated Author's Rebuttal
We thank the referee for a careful and constructive review. The two major comments identify genuine gaps in the manuscript: (1) the absence of direct surrogate-to-FEMR fidelity metrics, and (2) a flaw in the clinical validated events ratio R that conflates raw token prevalence with SHAP-weighted importance. We agree with both points and will revise the manuscript accordingly.
read point-by-point responses
-
Referee: §5.2, Tables 4–5: The surrogate model outperforms CLMBR-T-Base on ground-truth labels... surrogate-to-FEMR fidelity is never directly measured... The authors should report direct surrogate–FEMR agreement and discuss whether the observed performance gap undermines the central claim.
Authors: The referee is correct. Ground-truth performance metrics (AUROC, AUPRC, etc.) demonstrate that the surrogate is a competent predictor, but they do not establish that the surrogate faithfully replicates the FEMR's decision boundary. A surrogate that exceeds the teacher's performance against ground truth may have learned a different—and potentially better—function, in which case SHAP attributions on the surrogate would not validly explain the FEMR's reasoning. This is a load-bearing gap in our evidence chain, and we accept it as such. In the revised manuscript, we will report direct surrogate-to-FEMR agreement metrics, including prediction agreement rate and KL divergence on output probability distributions, for both hard-label and soft-label settings. We will also add an explicit discussion of the fidelity gap implied by the performance difference and temper the claim that the surrogate 'closely approximates' FEMR predictions to reflect what the fidelity metrics actually show. We note that the soft-label setting, which trains on the FEMR's probability outputs rather than binary decisions, is partially designed to mitigate this issue by matching the full output distribution; the fidelity metrics will allow us to quantify how well this works in practice. revision: yes
-
Referee: §3.4, Eqs. (5)–(6): The clinical validated events ratio R is defined as the count of clinically validated token occurrences divided by total token occurrences among SHAP-attributed tokens. However, Eq. (5) counts raw occurrence frequency of token f_j in the dataset, not the frequency with which f_j receives non-trivial SHAP importance. The text does not specify a SHAP importance threshold for inclusion in F. As written, R measures the prevalence of clinically validated codes in the data, not their attribution-weighted importance. Additionally, no baseline is reported, making R = 0.31 difficult to interpret. The metric should be redefined to weight by SHAP magnitude and compared against a null baseline.
Authors: The referee's analysis is accurate on both counts. As written, Eq. (5) counts raw occurrence frequency of tokens in the dataset without conditioning on whether those tokens actually received non-trivial SHAP importance, and no SHAP importance threshold is specified for inclusion in the set F. Consequently, R as currently defined conflates the base prevalence of clinically validated codes in the data with their attribution-weighted importance. Furthermore, without a null baseline (e.g., the ratio of clinically validated codes among all tokens in the dataset), R = 0.31 is uninterpretable—if the base rate is also approximately 0.30, the metric carries no signal about whether SHAP attributions preferentially select clinically validated features. We will revise the metric in the next version. Specifically, we will: (1) redefine the counting function c(f_j; D) to count only occurrences where f_j receives SHAP importance above a specified threshold, (2) define and report this threshold explicitly, (3) weight by SHAP magnitude as the referee suggests, and (4) compute and report a null baseline ratio of clinically validated codes among all tokens in the dataset. We will recompute R under the revised definition and re-evaluate whether the results still support the claim that SHAP-identified tokens align with clinical knowledge. revision: yes
Circularity Check
No circularity found: the clinical features are externally defined, the surrogate is trained on FEMR outputs (standard practice), and the alignment metric is not forced by construction.
full rationale
The paper's derivation chain proceeds as follows: (1) CLMBR-T-Base (FEMR) produces predictions on clinical tasks; (2) a Transformer surrogate is trained on the FEMR's input-output pairs; (3) SHAP is applied to the surrogate to identify important tokens; (4) a clinical validated events ratio R (Eq. 6) measures the fraction of SHAP-identified tokens that fall within an externally defined set F_c of clinically validated features (Tables 2, 3, drawn from the clinical literature [7, 5, 27, 32]). None of these steps reduce to their inputs by construction. The set F_c is defined independently of the model or its outputs — it comes from established clinical scoring systems (NEWS2, ICD-10 frailty codes, etc.). The metric R is not forced: it yields 0.31 and 0.27 for the two tasks, which would not be the case if the result were tautological. The surrogate is trained on FEMR outputs, which is the standard surrogate-model paradigm (not circular — it is an approximation whose fidelity is empirically assessed, albeit imperfectly). The concern that the surrogate may learn a different decision boundary than the FEMR (since it outperforms the FEMR on ground-truth metrics in Table 5) is a validity/correctness risk, not a circularity issue. The self-citation [37] (co-author Yin) concerns healthcare process variability and is not load-bearing for any claim in this paper. No step in the derivation chain is equivalent to its inputs by definition or by self-citation.
Axiom & Free-Parameter Ledger
free parameters (4)
- Surrogate model architecture =
3 layers, 128 hidden dim, 4 heads
- Learning rate =
1e-4
- Max sequence length =
2048
- Clinical validated feature sets (F_c) =
Tables 2, 3
axioms (3)
- domain assumption A surrogate model that achieves similar predictive performance to a FEMR has sufficiently similar internal decision logic for SHAP explanations to transfer.
- domain assumption The clinically validated features listed in Tables 2 and 3 are the correct ground truth for model explanations.
- standard math SHAP values provide meaningful token-level attribution for Transformer models.
read the original abstract
Foundation Models for Electronic Health Records (FEMRs) are pretrained on large-scale structured patient data, enabling them to convert longitudinal patient trajectories into generalizable representations for diverse clinical prediction tasks. Despite their effectiveness, FEMRs remain black-box models, raising concerns about bias, interpretability, and clinical trust. To address this, we propose the first token-level explainability approach for FEMRs. We train a Transformer-based surrogate model on input-output pairs from the FEMR across two prediction tasks, approximating its behavior while preserving temporal dynamics. We identify the most influential tokens, providing insights into how FEMRs leverage different aspects of patient history for predictions. To evaluate clinical relevance, we introduce a novel clinical alignment metric that quantifies the correspondence between the surrogate model's key tokens and clinically validated features. Our results demonstrate that the surrogate closely approximates FEMR predictions and that token-level explanations align well with clinical knowledge, offering a practical framework for interpretable and trustworthy clinical AI.
Figures
Reference graph
Works this paper leans on
-
[1]
Qaiser Abbas, Woonyoung Jeong, and Seung Won Lee. Explainable ai in clinical decision support systems: A meta-analysis of methods, applications, and usability challenges.Healthcare, 13(17), 2025
work page 2025
-
[2]
Sajid Ali, Tamer Abuhmed, Shaker El-Sappagh, Khan Muhammad, Jose M. Alonso-Moral, Roberto Confalonieri, Riccardo Guidotti, Javier Del Ser, Natalia D ´ıaz-Rodr´ıguez, and Francisco Herrera. Ex- plainable artificial intelligence (xai): What we know and what is left to attain trustworthy artificial intel- ligence.Information Fusion, 99:101805, 2023
work page 2023
-
[3]
Prediction of 30-day hospital readmission with clinical notes and ehr information
Tiago Almeida, Plinio Moreno, and Catarina Barata. Prediction of 30-day hospital readmission with clinical notes and ehr information. In Nuno Gonc ¸alves, H´elder P. Oliveira, and Joan Andreu S´anchez, editors,Pattern Recognition and Image Analysis, pages 220–232, Cham, 2026. Springer Nature Switzer- land
work page 2026
-
[4]
Nadine Bienefeld, Jens Michael Boss, Rahel L ¨uthy, Dominique Brodbeck, Jan Azzati, Mirco Blaser, Jan Willms, and Emanuela Keller. Solving the explainable AI conundrum by bridging clinicians’ needs and developers’ goals.npj Digital Medicine, 6(1):94, May 2023
work page 2023
-
[5]
A national early warning score for acutely ill patients.BMJ, 345, 2012
BMJ. A national early warning score for acutely ill patients.BMJ, 345, 2012
work page 2012
-
[6]
Guilherme J. Cavalcante, Jos ´e Gabriel A. Moreira, Gabriel A. B. do Nascimento, Vincent Dong, Alex Nguyen, Tha´ıs G. do Rˆego, Yuri Malheiros, Telmo M. Silva Filho, Carla R. Zeballos Torrez, James C. Gee, Anne Marie McCarthy, Andrew D. A. Maidment, and Bruno Barufaldi. Toward explainable ai approaches for breast imaging: adapting foundation models to div...
work page 2025
-
[7]
Mieke Deschepper, Chlo ¨e De Smedt, and Kirsten Colpaert. A literature-based approach to predict con- tinuous hospital length of stay in adult acute care patients using admission variables: A single university center experience.International Journal of Medical Informatics, 193:105678, 2025
work page 2025
-
[8]
Swapna Gokhale, David Taylor, Jaskirath Gill, Yanan Hu, Nikolajs Zeps, Vincent Lequertier, Luis Prado, Helena Teede, and Joanne Enticott. Hospital length of stay prediction tools for all hospital ad- missions and general medicine populations: systematic review and meta-analysis.Frontiers in Medicine, 10:1192969, August 2023
work page 2023
-
[9]
Lin Lawrence Guo, Jason Fries, Ethan Steinberg, Scott Lanyon Fleming, Keith Morse, Catherine Af- tandilian, Jose Posada, Nigam Shah, and Lillian Sung. A multi-center study on the adaptability of a shared foundation model for electronic health records.NPJ Digital Medicine, 7(1):171, 2024
work page 2024
-
[10]
Foundation models in bioinformatics.Natl
Fei Guo, Renchu Guan, Yaohang Li, Qi Liu, Xiaowo Wang, Can Yang, and Jianxin Wang. Foundation models in bioinformatics.Natl. Sci. Rev., 12(4):nwaf028, April 2025
work page 2025
-
[11]
Niklas, Li Yu-Chuan, Stang Paul E., Madigan David, and Ryan Patrick B
Hripcsak George, Duke Jon D., Shah Nigam H., Reich Christian G., Huser V ojtech, Schuemie Mar- tijn J., Suchard Marc A., Park Rae Woong, Wong Ian Chi Kei, Rijnbeek Peter R., Van Der Lei Johan, Pratt Nicole, Norén G. Niklas, Li Yu-Chuan, Stang Paul E., Madigan David, and Ryan Patrick B. Observational Health Data Sciences and Informatics (OHDSI): Opp...
work page 2015
-
[12]
Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. Scaling Laws for Neural Language Models, January
-
[13]
arXiv:2001.08361 [cs]
work page internal anchor Pith review Pith/arXiv arXiv 2001
-
[14]
Hyeokjong Lee, Jaewon Kim, Sangmin Kwak, Azka Rehman, Sang Min Park, and Jooyoung Chang. Optimizing retinal images based carotid atherosclerosis prediction with explainable foundation models. npj Digital Medicine, 8(1):582, September 2025
work page 2025
-
[15]
Jianfang Liu, Elaine Larson, Amanda Hessels, Bevin Cohen, Philip Zachariah, David Caplan, and Jingjing Shang. Comparison of Measures to Predict Mortality and Length of Stay in Hospitalized Pa- tients.Nursing Research, 68(3):200–209, May 2019
work page 2019
-
[16]
A unified approach to interpreting model predictions
Scott M Lundberg and Su-In Lee. A unified approach to interpreting model predictions. In I. Guyon, U. V on Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, editors,Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc., 2017
work page 2017
-
[17]
Clement J McDonald, Stanley M Huff, Jeffrey G Suico, Gilbert Hill, Dennis Leavelle, Raymond Aller, Arden Forrey, Kathy Mercer, Georges DeMoor, John Hook, Warren Williams, James Case, Pat Maloney, and for the Laboratory LOINC Developers. LOINC, a Universal Standard for Identifying Laboratory Observations: A 5-Year Update.Clinical Chemistry, 49(4):624–633, ...
work page 2003
-
[18]
Stuart J Nelson, Kelly Zeng, John Kilbourne, Tammy Powell, and Robin Moore. Normalized names for clinical drugs: Rxnorm at 6 years.Journal of the American Medical Informatics Association, 18(4):441– 448, 04 2011
work page 2011
-
[19]
Laila Rasmy, Yang Xiang, Ziqian Xie, Cui Tao, and Degui Zhi. Med-BERT: pretrained contextual- ized embeddings on large-scale structured electronic health records for disease prediction.npj Digital Medicine, 4(1):86, May 2021. 13
work page 2021
-
[20]
Zero shot health trajectory prediction using transformer.NPJ Digit
Pawel Renc, Yugang Jia, Anthony E Samir, Jaroslaw Was, Quanzheng Li, David W Bates, and Arka- diusz Sitek. Zero shot health trajectory prediction using transformer.NPJ Digit. Med., 7(1):256, Septem- ber 2024
work page 2024
-
[21]
Foundation model of electronic medical records for adaptive risk estimation
Pawel Renc, Michal K Grzeszczyk, Nassim Oufattole, Deirdre Goode, Yugang Jia, Szymon Biegan- ski, Matthew B A McDermott, Jaroslaw Was, Anthony E Samir, Jonathan W Cunningham, David W Bates, and Arkadiusz Sitek. Foundation model of electronic medical records for adaptive risk estimation. Gigascience, 14(giaf107), January 2025
work page 2025
-
[22]
”Why Should I Trust You?”: Explaining the Predictions of Any Classifier
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. ”Why Should I Trust You?”: Explaining the Predictions of Any Classifier. InProceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pages 1135–1144, San Francisco California USA, August 2016. ACM
work page 2016
-
[23]
Yucheng Ruan, Daniel J Tan, See-Kiong Ng, Ling Huang, and Mengling Feng. Towards accurate and reliable icu outcome prediction: a multimodal learning framework based on belief function theory using structured ehrs and free-text notes.Journal of Healthcare Informatics Research, pages 1–42, 2025
work page 2025
-
[24]
Shahab S Band, Atefeh Yarahmadi, Chung-Chian Hsu, Meghdad Biyari, Mehdi Sookhak, Rasoul Ameri, Iman Dehzangi, Anthony Theodore Chronopoulos, and Huey-Wen Liang. Application of explain- able artificial intelligence in medical health: A systematic review of interpretability methods.Informatics in Medicine Unlocked, 40:101286, 2023
work page 2023
-
[25]
Explainable AI for healthcare 5.0: Opportunities and challenges.IEEE Access, 10:84486–84517, 2022
Deepti Saraswat, Pronaya Bhattacharya, Ashwin Verma, Vivek Kumar Prasad, Sudeep Tanwar, Gul- shan Sharma, Pitshou N Bokoro, and Ravi Sharma. Explainable AI for healthcare 5.0: Opportunities and challenges.IEEE Access, 10:84486–84517, 2022
work page 2022
-
[26]
Iqbal H Sarker. Machine learning: Algorithms, real-world applications and research directions.SN computer science, 2(3):160, 2021
work page 2021
-
[27]
Deep Inside Convolutional Networks: Visualising Image Classification Models and Saliency Maps
Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman. Deep Inside Convolutional Networks: Visualising Image Classification Models and Saliency Maps, April 2014. arXiv:1312.6034 [cs]
work page internal anchor Pith review Pith/arXiv arXiv 2014
-
[28]
The National Early Warning Score 2 (NEWS2)
Gary B Smith, Oliver C Redfern, Marco Af Pimentel, Stephen Gerry, Gary S Collins, James Malycha, David Prytherch, Paul E Schmidt, and Peter J Watkinson. The National Early Warning Score 2 (NEWS2). Clinical Medicine, 19(3):260, May 2019
work page 2019
-
[29]
Andrew Street, Laia Maynou, Joanna M Blodgett, and Simon Conroy. Association between Hospi- tal Frailty Risk Score and length of hospital stay, hospital mortality, and hospital costs for all adults in England: a nationally representative, retrospective, observational cohort study.The Lancet Healthy Longevity, 6(8):100740, August 2025
work page 2025
-
[30]
Axiomatic attribution for deep networks
Mukund Sundararajan, Ankur Taly, and Qiqi Yan. Axiomatic attribution for deep networks. In Proceedings of the 34th International Conference on Machine Learning - V olume 70, ICML’17, page 3319–3328. JMLR.org, 2017
work page 2017
-
[31]
Mohan Timilsina, Samuele Buosi, Muhammad Asif Razzaq, Rafiqul Haque, Conor Judge, and Edward Curry. Harmonizing foundation models in healthcare: A comprehensive survey of their roles, relation- ships, and impact in artificial intelligence’s advancing terrain.Computers in Biology and Medicine, 189:109925, 2025. 14
work page 2025
-
[32]
McCradden, and Anna Goldenberg
Sana Tonekaboni, Shalmali Joshi, Melissa D. McCradden, and Anna Goldenberg. What clinicians want: Contextualizing explainable machine learning for clinical end use. In Finale Doshi-Velez, Jim Fackler, Ken Jung, David Kale, Rajesh Ranganath, Byron Wallace, and Jenna Wiens, editors,Proceed- ings of the 4th Machine Learning for Healthcare Conference, volume ...
work page 2019
-
[33]
Amol A. Verma. Toward the Rigorous Evaluation of Early Warning Scores.JAMA Network Open, 7(10):e2438966, October 2024
work page 2024
-
[34]
Song Wang, Yishu Wei, Haotian Ma, Max Lovitt, Kelly Deng, Yuan Meng, Zihan Xu, Jingze Zhang, Yunyu Xiao, Ying Ding, Xuhai Xu, Joydeep Ghosh, and Yifan Peng. A multi-stage large language model framework for extracting suicide-related social determinants of health.Commun. Med. (Lond.), 5(1):404, September 2025
work page 2025
-
[35]
Ben Wellner, Joan Grand, Elizabeth Canzone, Matt Coarr, Patrick W Brady, Jeffrey Simmons, Eric Kirkendall, Nathan Dean, Monica Kleinman, Peter Sylvester, et al. Predicting unplanned transfers to the intensive care unit: a machine learning approach leveraging diverse clinical elements.JMIR medical informatics, 5(4):e8680, 2017
work page 2017
-
[36]
Ehrshot: An ehr benchmark for few-shot evaluation of foundation models
Michael Wornow, Rahul Thapa, Ethan Steinberg, Jason Fries, and Nigam Shah. Ehrshot: An ehr benchmark for few-shot evaluation of foundation models. 2023
work page 2023
-
[37]
Michael Wornow, Yizhe Xu, Rahul Thapa, Birju Patel, Ethan Steinberg, Scott Fleming, Michael A Pfeffer, Jason Fries, and Nigam H Shah. The shaky foundations of large language models and foundation models for electronic health records.npj digital medicine, 6(1):135, 2023
work page 2023
-
[38]
Pengfei Yin, Abel Armas Cervantes, and Daniel Capurro. Measuring and visualizing healthcare pro- cess variability.Journal of Biomedical Informatics, 170:104918, 2025
work page 2025
-
[39]
Michael P. Young, Valerie J. Gooder, Karen McBride, Brent James, and Elliott S. Fisher. Inpatient transfers to the intensive care unit: Delays are associated with increased mortality and morbidity.Journal of General Internal Medicine, 18(2):77–83, February 2003. 15
work page 2003
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.