Pith. sign in

REVIEW 3 major objections 5 minor 105 references

Pretraining that forces models to recover waveform shape, not just signal values, yields more transferable ECG and pulse-oximetry representations than reconstruction or generic self-supervised learning.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-14 16:30 UTC pith:QV5YXRCU

load-bearing objection Solid multimodal SSL recipe for ECG+SpO2 with consistent MIMIC gains; the morphology-masking story is the real claim and is under-specified, but the paper is still worth a referee. the 3 major comments →

arxiv 2607.09749 v1 pith:QV5YXRCU submitted 2026-07-03 eess.SP cs.AIcs.LG

MorphologyFM: A Foundation Model for Morphology-Aware Representation Learning from ECG and Pulse Oximetry Waveforms

classification eess.SP cs.AIcs.LG
keywords foundation modelsECGpulse oximetrywaveform morphologyself-supervised learningmultimodal representation learningphysiological monitoringmasked autoencoders
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper argues that the clinically useful content in ECG and pulse-oximetry signals lives mainly in their morphology—the shapes of waves, intervals, slopes, and beat-to-beat patterns—not in sample-by-sample amplitudes. Existing self-supervised methods mostly reconstruct or forecast raw signals, so they can copy local continuity without learning those shapes. MorphologyFM is a multimodal foundation model pretrained on paired ECG and SpO2 from MIMIC that masks whole morphological components, aligns the two modalities in latent space, and uses contrastive learning so similar morphologies cluster together. Across arrhythmia classification, hypoxemia prediction, mortality, and length-of-stay estimation, it reports consistent gains over MAE, contrastive learning, Barlow Twins, and JEPA, and joint ECG–SpO2 pretraining beats either modality alone. The sympathetic reading is that morphology is a strong, annotation-free inductive bias for continuous physiological monitoring.

Core claim

Large-scale self-supervised pretraining that explicitly preserves clinically meaningful waveform morphology on paired ECG and SpO2 produces more transferable physiological representations than objectives based primarily on reconstruction or generic embedding alignment. MorphologyFM combines morphology-guided masking of complete waveform components, cross-modal latent alignment of synchronized ECG–SpO2 windows, and contrastive organization of morphologically similar segments; the resulting encoder improves multiple downstream clinical prediction tasks and benefits from multimodal joint pretraining and scale of unlabeled waveforms.

What carries the argument

Morphology-aware masking: instead of random sample masks, approximate cardiac cycles (R-peaks) and aligned pulse peaks define masks over whole ECG components (P, QRS, ST, T) and SpO2 regions (systolic upstroke, diastolic decay, pulse peak, respiratory modulation), forcing the encoder to infer higher-order physiological structure. This is trained jointly with masked reconstruction, cross-modal alignment, and contrastive morphology objectives.

Load-bearing premise

The method assumes that automatic peak detection and masking of complete morphological components on noisy ICU waveforms reliably isolates the clinically meaningful structure that drives the reported gains, rather than other unmeasured training differences.

What would settle it

Re-run the same backbone, data, and protocol with morphology-aware masking replaced by random or temporal-block masking at matched mask rates and augmentations; if average downstream AUROC/F1 no longer favors MorphologyFM over MAE and JEPA (Table 1 vs Table 3 setup), the morphology-inductive-bias claim fails.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Waveform foundation models should treat morphology preservation as a first-class pretraining objective, not only reconstruction or forecasting fidelity.
  • Paired ECG–SpO2 pretraining should yield more transferable cardiovascular representations than single-modality pretraining of comparable size.
  • Downstream arrhythmia, hypoxemia, mortality, and length-of-stay models can be improved by fine-tuning a morphology-pretrained encoder with less labeled data than task-specific supervised training from scratch.
  • Scaling unlabeled paired waveforms (hundreds of thousands to millions of segments) should continue to improve transfer until diminishing returns set in.
  • Patient-similarity retrieval in the latent space should become more clinically coherent when embeddings are organized by morphology rather than reconstruction alone.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If component-level masking is the real driver, the same idea should transfer to other periodic clinical waveforms (arterial pressure, capnography, EEG envelopes) once reliable landmark detectors exist.
  • Brittle R-peak or pulse-peak detection in artifact-heavy ICU data would systematically weaken the claimed inductive bias; robustness of the masker is as important as the transformer backbone.
  • A natural next test is whether morphology-pretrained embeddings improve early-warning tasks that clinicians already read from shape (ST change, pulse-pressure variation) under device and hospital shift.
  • Joint latent alignment of synchronized modalities may act as free multi-view supervision that partly substitutes for labels in critical-care monitoring.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. MorphologyFM is a multimodal foundation model for ECG and SpO2 waveforms pretrained on MIMIC with a morphology-aware self-supervised objective. The method combines morphology-guided masking of complete ECG components (P/QRS/ST/T) and SpO2 regions (systolic/diastolic/respiratory), cross-modal latent alignment of synchronized pairs, and a contrastive morphology objective, jointly optimized with a reconstruction loss over masked regions only. After pretraining, the encoder is fine-tuned on arrhythmia classification, hypoxemia prediction, mortality prediction, and length-of-stay estimation. Against MAE, SimCLR-style contrastive learning, Barlow Twins, and JEPA under a shared transformer backbone and fine-tuning protocol, MorphologyFM reports the strongest average transfer (89.6 vs. 83.8–86.5 in Table 1), with further gains from joint ECG+SpO2 pretraining (Table 2), morphology-aware vs. random/temporal-block masking (Table 3), full multi-term objective (Table 4), nearest-neighbor retrieval (Table 5), and scaling of unlabeled pretraining data (Table 6).

Significance. If the gains are truly attributable to morphology as an inductive bias rather than unreported protocol differences, the paper would strengthen the case for structure-preserving SSL on continuous physiological signals and provide a practical multimodal backbone for ICU monitoring. The experimental design is useful: same backbone and fine-tuning protocol across MAE, contrastive, Barlow Twins, and JEPA; multimodal vs. single-modality comparison; masking and objective ablations; retrieval and scaling analyses. These elements make the work a credible contribution to physiological foundation models even if some causal claims need tighter support. The absence of released code or external validation is a limitation but does not erase the empirical pattern on MIMIC.

major comments (3)
  1. [§3.4 / Table 3] §3.4 and Table 3: The central claim that morphology-aware masking of complete P/QRS/ST/T and SpO2 systolic/diastolic/respiratory regions is the inductive bias behind the gains is under-specified. The paper does not report R-peak/component detection success rates or failure modes on noisy MIMIC ICU waveforms, nor the realized fraction of samples masked under morphology-aware vs. random vs. temporal-block policies. Without those controls, Table 3 is consistent with a different effective mask rate or difficulty rather than a morphology-specific effect, which is load-bearing for the title and strongest claim.
  2. [§3.8 / §4.1] §3.8 and §4.1: Free parameters that define the method—λ_rec, λ_con, λ_align, InfoNCE temperature τ, morphology-component masking probabilities, and the augmentation mixture—are not given numerical values or selection criteria. The claim that identical backbone, optimization schedule, and fine-tuning protocol isolate the pretraining objective therefore cannot be verified from the manuscript. At minimum, report the hyperparameter settings used for MorphologyFM and confirm that baselines received comparable tuning budgets.
  3. [Table 1 / §4.1] Table 1 and §4.1: Length of stay is described as estimation yet scored with AUROC, which requires an unstated binarization (or multi-class reduction). Also, results are averaged over five independent runs with no standard deviations, confidence intervals, or significance tests. Both issues weaken the quantitative support for the ranking MorphologyFM > JEPA > Barlow > contrastive > MAE.
minor comments (5)
  1. [References] Related work cites several arXiv items with 2025–2026 dates and placeholder-style entries (e.g., [25?]); clean the bibliography for consistency and verifiability.
  2. [§3.2] §3.2: State the common sampling frequency, window length in samples/seconds, and exact signal-quality discard criteria so the pretraining corpus is reproducible.
  3. [§3.3] §3.3: Specify transformer depth, width, number of heads, patch size, and whether the class token is learned or a mean pool; Table 2 parameter counts alone are insufficient.
  4. [Abstract / Table 6] Abstract and §1 claim ‘millions of waveform segments’ while Table 6 tops out at 10M; align the prose with the scaling table and state the exact pretraining corpus size used for the main results.
  5. [Table 2] Table 2 reports ‘Average AUROC’ while Table 1 mixes F1 and AUROC; clarify how the average is computed when metrics differ across tasks.

Circularity Check

0 steps flagged

No significant circularity: MorphologyFM is empirical SSL evaluated on external supervised labels, not a derivation that redefines its target.

full rationale

The paper’s load-bearing claims are experimental: morphology-guided masking, cross-modal alignment, and contrastive objectives are pretrained on unlabeled MIMIC ECG/SpO2, then transferred to separate supervised tasks (arrhythmia F1, hypoxemia/mortality/LOS AUROC) against MAE, contrastive, Barlow Twins, and JEPA with a shared backbone (Tables 1–6). Downstream metrics are not algebraically forced by the pretraining loss (L = λ_rec L_rec + λ_con L_con + λ_align L_align); reconstruction is restricted to masked regions, contrastive pairs come from morphology-preserving augmentations, and alignment uses synchronized modality pairs—none of these redefine the clinical labels used at evaluation. There is no self-definitional loop (morphology is operationalized via R-peak/component masking, not via the reported AUROC/F1), no fitted-input-called-prediction (pretraining does not fit to the downstream targets), no uniqueness theorem imported from the authors, and no ansatz smuggled in as a forced mathematical result. The single same-author citation (Chronoformer [87]) is background temporal modeling and is not load-bearing for the MorphologyFM objective or tables. Weaknesses (under-specified mask rates, detection reliability, hyperparameter schedules) are correctness/confounding risks, not circularity. The experimental chain is self-contained against external task labels.

Axiom & Free-Parameter Ledger

5 free parameters · 5 axioms · 2 invented entities

The central claim rests on empirical SSL design choices and domain premises about clinical morphology, not on a closed-form theory. Load-bearing free parameters are the loss weights, temperature, and masking/augmentation knobs that define the objective. Domain axioms include that automatic peak/component detection yields clinically meaningful masks and that MIMIC paired waveforms are a sufficient pretraining distribution. MorphologyFM is an engineered system, not a new physical entity; independent evidence is only the paper’s own tables.

free parameters (5)
  • λ_rec, λ_con, λ_align (joint loss weights)
    §3.8 defines L = λ_rec Lrec + λ_con Lcon + λ_align Lalign; values are not reported but control the claimed objective balance.
  • InfoNCE temperature τ
    §3.7 uses τ in the contrastive loss; temperature strongly affects embedding geometry and is not specified.
  • Morphology-component masking probabilities
    §3.4 says complete morphological components are selected with predefined probabilities; those probabilities are free design choices that define the method.
  • Augmentation mixture (noise, stretch, scale, crop, baseline wander)
    §3.2 stochastic augmentations shape the contrastive positives; strengths/schedules are unspecified free knobs.
  • Transformer/tokenizer capacity and optimization schedule
    Only ~89–95M parameters are given (Table 2); depth, width, patch size, LR, epochs, and batch size are free and unreported.
axioms (5)
  • domain assumption Clinically meaningful information in ECG/SpO2 is primarily morphological (shape, intervals, slopes, beat variability) rather than sample-wise amplitude continuity.
    Stated throughout Introduction and Discussion; motivates replacing pure reconstruction with morphology-aware objectives.
  • domain assumption Automatic R-peak detection and SpO2 peak alignment can identify complete morphological components well enough for masking on ICU data.
    §3.4 depends on this without reporting detector accuracy or failure modes.
  • domain assumption MIMIC critical-care paired ECG–SpO2 segments are a representative unlabeled corpus for transferable physiological representations.
    All pretraining and evaluation are on MIMIC (§3.1, §4.1); generalization is assumed, not tested externally.
  • ad hoc to paper Same transformer backbone + same fine-tuning protocol isolates pretraining-objective effects across MAE, contrastive, Barlow Twins, JEPA, and MorphologyFM.
    §4.1 asserts fair comparison; without hyperparameter search details this is an experimental axiom.
  • standard math Standard SSL constructions (masked reconstruction, InfoNCE, latent multimodal alignment) are valid representation-learning objectives.
    Used as background machinery in §3.5–3.8 without re-derivation.
invented entities (2)
  • MorphologyFM (morphology-aware multimodal waveform foundation model) no independent evidence
    purpose: Name the combined encoder, morphology-guided masking, cross-modal alignment, and contrastive objective as a general-purpose physiological backbone.
    The entity is the paper’s system; evidence is internal downstream tables only, with no external independent validation or released artifact.
  • Morphology-aware masking operator over ECG/SpO2 components no independent evidence
    purpose: Replace random/temporal masks with clinically structured masks (P/QRS/ST/T; systolic/diastolic/respiratory regions).
    Defined in §3.4 as the key inductive bias; no external benchmark isolates this operator beyond the paper’s ablation Table 3.

pith-pipeline@v1.1.0-grok45 · 19920 in / 3869 out tokens · 46157 ms · 2026-07-14T16:30:13.692460+00:00 · methodology

0 comments
read the original abstract

Foundation models have recently emerged as a powerful paradigm for learning transferable representations from large scale biomedical data, yet existing approaches for physiological waveforms primarily optimize reconstruction or forecasting objectives that do not explicitly preserve clinically meaningful waveform morphology. Electrocardiograms (ECGs) and pulse oximetry (SpO2) waveforms encode rich cardiovascular and hemodynamic information through their morphological structure. In this work, we introduce MorphologyFM, a multimodal foundation model pretrained on paired ECG and SpO2 waveforms from the MIMIC critical care database using a morphology aware self supervised learning objective. MorphologyFM combines morphology guided masking, cross modal representation learning, and contrastive latent alignment to learn representations that capture clinically relevant physiological structure without requiring manual annotations. We evaluate MorphologyFM across multiple downstream prediction tasks, including arrhythmia classification, hypoxemia prediction, mortality prediction, and length of stay estimation, demonstrating consistent improvements over representative self supervised learning methods, including Masked Autoencoders (MAE), contrastive learning, Barlow Twins, and Joint Embedding Predictive Architectures (JEPA). Furthermore, we show that jointly modeling ECG and SpO2 waveforms produces more transferable representations than single modality pretraining. Our results establish waveform morphology as a powerful inductive bias for self supervised physiological representation learning and introduce MorphologyFM as a general purpose foundation model for continuous physiological monitoring.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

105 extracted references · 28 linked inside Pith

  1. [1]

    Bert: Pre-training of deep bidirectional transformers for language understanding

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. InProceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers), pages 4171–4186, 2019

  2. [2]

    Masked autoencoders are scalable vision learners

    Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick. Masked autoencoders are scalable vision learners. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 16000–16009, 2022

  3. [3]

    Deep residual learning for image recognition, 2015

    Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition, 2015. URLhttps://arxiv.org/abs/1512.03385

  4. [4]

    Multi-scale 3d deep convolutional neural network for hyperspectral image classification

    Mingyi He, Bo Li, and Huahui Chen. Multi-scale 3d deep convolutional neural network for hyperspectral image classification. In2017 IEEE International Conference on Image Processing (ICIP), pages 3904–3908. IEEE, 2017

  5. [5]

    Bag of tricks for image classification with convolutional neural networks

    Tong He, Zhi Zhang, Hang Zhang, Zhongyue Zhang, Junyuan Xie, and Mu Li. Bag of tricks for image classification with convolutional neural networks. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 558–567, 2019

  6. [6]

    A foundational vision transformer improves diagnostic performance for electrocardiograms.NPJ Digital Medicine, 6(1):108, 2023

    Akhil Vaid, Joy Jiang, Ashwin Sawant, Stamatios Lerakis, Edgar Argulian, Yuri Ahuja, Joshua Lampert, Alexander Charney, Hayit Greenspan, Jagat Narula, et al. A foundational vision transformer improves diagnostic performance for electrocardiograms.NPJ Digital Medicine, 6(1):108, 2023

  7. [7]

    Foundation models in healthcare: Opportunities, risks & strategies forward

    Anja Thieme, Aditya Nori, Marzyeh Ghassemi, Rishi Bommasani, Tariq Osman Andersen, and Ewa Luger. Foundation models in healthcare: Opportunities, risks & strategies forward. InExtended abstracts of the 2023 CHI conference on human factors in computing systems, pages 1–4, 2023

  8. [8]

    Foundation models for time series forecasting.International IT Journal of Research, ISSN: 3007-6706, 2(4):144–156, 2024

    Suresh Chandra Thakur. Foundation models for time series forecasting.International IT Journal of Research, ISSN: 3007-6706, 2(4):144–156, 2024

  9. [9]

    A foun- dation model for intensive care: Unlocking generalization across tasks and domains at scale

    Manuel Burger, Daphné Chopard, Malte Londschien, Fedor Sergeev, Hugo Yèche, Rita Kuznetsova, Martin Faltys, Eike Gerdes, Polina Leshetkina, Peter Bühlmann, et al. A foun- dation model for intensive care: Unlocking generalization across tasks and domains at scale. medRxiv, pages 2025–07, 2025

  10. [10]

    Foundation models defining a new era in vision: a survey and outlook.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025

    Muhammad Awais, Muzammal Naseer, Salman Khan, Rao Muhammad Anwer, Hisham Cholakkal, Mubarak Shah, Ming-Hsuan Yang, and Fahad Shahbaz Khan. Foundation models defining a new era in vision: a survey and outlook.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025

  11. [11]

    Foundation model for advancing healthcare: challenges, opportunities and future directions.IEEE Reviews in Biomedical Engineering, 2024

    Yuting He, Fuxiang Huang, Xinrui Jiang, Yuxiang Nie, Minghao Wang, Jiguang Wang, and Hao Chen. Foundation model for advancing healthcare: challenges, opportunities and future directions.IEEE Reviews in Biomedical Engineering, 2024

  12. [12]

    Foundation models for electronic health records: repre- sentation dynamics and transferability.arXiv preprint arXiv:2504.10422, 2025

    Michael C Burkhart, Bashar Ramadan, Zewei Liao, Kaveri Chhikara, Juan C Rojas, William F Parker, and Brett K Beaulieu-Jones. Foundation models for electronic health records: repre- sentation dynamics and transferability.arXiv preprint arXiv:2504.10422, 2025. 11

  13. [13]

    Foundation models for time series analysis: A tutorial and survey

    Yuxuan Liang, Haomin Wen, Yuqi Nie, Yushan Jiang, Ming Jin, Dongjin Song, Shirui Pan, and Qingsong Wen. Foundation models for time series analysis: A tutorial and survey. In Proceedings of the 30th ACM SIGKDD conference on knowledge discovery and data mining, pages 6555–6565, 2024

  14. [14]

    Foundation models in bioinformatics.National science review, 12(4):nwaf028, 2025

    Fei Guo, Renchu Guan, Yaohang Li, Qi Liu, Xiaowo Wang, Can Yang, and Jianxin Wang. Foundation models in bioinformatics.National science review, 12(4):nwaf028, 2025

  15. [15]

    Physiology-aware masked cross-modal reconstruction for biosignal representation learning.arXiv preprint arXiv:2605.00973, May 2026

    Hao Zhou, Simon A Lee, Cyrus Tanade, Keum San Chun, Juhyeon Lee, Migyeong Gwak, Megha Thukral, Justin Sung, Eugene Hwang, Mehrab Bin Morshed, et al. Physiology-aware masked cross-modal reconstruction for biosignal representation learning.arXiv preprint arXiv:2605.00973, May 2026

  16. [16]

    Large-scale training of foundation models for wearable biosignals

    Salar Abbaspourazad, Oussama Elachqar, Andrew C Miller, Saba Emrani, Udhyakumar Nallasamy, and Ian Shapiro. Large-scale training of foundation models for wearable biosignals. arXiv preprint arXiv:2312.05409, 2023

  17. [17]

    Sonata: A hybrid world model for inertial kinematics under clinical data scarcity

    Blaise Delaney, Salil Patel, Yuji Xing, Dominic Dootson, Karin Sevegnani, and Chrystalina Antoniades. Sonata: A hybrid world model for inertial kinematics under clinical data scarcity. arXiv preprint arXiv:2604.18058, 2026

  18. [18]

    Edge ppg quality gating with validity-labeled outputs and dual-channel communication on a raspberry pi zero 2 w wearable device

    Haoyang Zhou, Yi Ding, Jiajun Li, Yifan Yi, and Chun Xiao. Edge ppg quality gating with validity-labeled outputs and dual-channel communication on a raspberry pi zero 2 w wearable device. InJournal of Physics: Conference Series, volume 3235, page 012025. IOP Publishing, 2026

  19. [19]

    Himae: Hierarchical masked autoencoders discover resolution-specific structure in wearable time series.arXiv preprint arXiv:2510.25785, 2025

    Simon A Lee, Cyrus Tanade, Hao Zhou, Juhyeon Lee, Megha Thukral, Minji Han, Rachel Choi, Md Sazzad Hissain Khan, Baiying Lu, Migyeong Gwak, et al. Himae: Hierarchical masked autoencoders discover resolution-specific structure in wearable time series.arXiv preprint arXiv:2510.25785, 2025

  20. [20]

    Multimodal self-supervised learning for wearable sleep staging using photoplethysmography and accelerometer signals

    Juhyeon Lee, Simon A Lee, Cyrus Tanade, Viswam Nathan, Megha Thukral, Hao Zhou, Keum San Chun, and Sharanya Arcot Desai. Multimodal self-supervised learning for wearable sleep staging using photoplethysmography and accelerometer signals. InICASSP 2026-2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 7852–7856. ...

  21. [21]

    Glucofm: A dual- stream foundation model for continuous glucose monitoring.arXiv preprint arXiv:2605.30865, 2026

    Zechen Li, Keerthana Natarajan, Weizhi Zhang, Menglian Zhou, Simon A Lee, Yuwei Zhang, Maxwell A Xu, Zeinab Esmaeilpour, Flora D Salim, Mark Malhotra, et al. Glucofm: A dual- stream foundation model for continuous glucose monitoring.arXiv preprint arXiv:2605.30865, 2026

  22. [22]

    Trust and transparency in healthcare ai: A systematic review of explainable nlp for clinical decision support (2023–2025)

    MS Tharini and Jane Rubel Angelina Jeyaraj. Trust and transparency in healthcare ai: A systematic review of explainable nlp for clinical decision support (2023–2025). InInternational Conference on Edge Computing and Applications, pages 451–470. Springer, 2025

  23. [23]

    Smart assistive technologies for neurodisorders: A review on ai, iot, and wearable systems for enhanced patient care.Neurological Sciences, 47(2):211, 2026

    Sandeep Chouhan, Deepika Ghai, Ramandeep Sandhu, and Suman Lata Tripathi. Smart assistive technologies for neurodisorders: A review on ai, iot, and wearable systems for enhanced patient care.Neurological Sciences, 47(2):211, 2026

  24. [24]

    A preoperative data sentenceization method for postoperative major adverse cardiovascular event prediction

    Yuelin Luo, Yaqiang Wang, Xiran Peng, Ruihao Zhou, Guo Chen, Xuechao Hao, and Tao Zhu. A preoperative data sentenceization method for postoperative major adverse cardiovascular event prediction. In2025 IEEE International Conference on Big Data (BigData), pages 8060–8069. IEEE, 2025

  25. [25]

    A large language model for electronic health records.NPJ digital medicine, 5(1):194, 2022

    Xi Yang, Aokun Chen, Nima PourNejatian, Hoo Chang Shin, Kaleb E Smith, Christopher Parisien, Colin Compas, Cheryl Martin, Anthony B Costa, Mona G Flores, et al. A large language model for electronic health records.NPJ digital medicine, 5(1):194, 2022

  26. [26]

    Foundation models for physiological signals: Opportunities and challenges

    Simon A Lee and Kai Akamatsu. Foundation models for physiological signals: Opportunities and challenges. August 2025. 12

  27. [27]

    The shaky foundations of large language models and foundation models for electronic health records.npj digital medicine, 6(1):135, 2023

    Michael Wornow, Yizhe Xu, Rahul Thapa, Birju Patel, Ethan Steinberg, Scott Fleming, Michael A Pfeffer, Jason Fries, and Nigam H Shah. The shaky foundations of large language models and foundation models for electronic health records.npj digital medicine, 6(1):135, 2023

  28. [28]

    Jepa-dna: Grounding ge- nomic foundation models through joint-embedding predictive architectures.arXiv preprint arXiv:2602.17162, 2026

    Ariel Larey, Elay Dahan, Amit Bleiweiss, Raizy Kellerman, Guy Leib, Omri Nayshool, Dan Ofer, Tal Zinger, Dan Dominissini, Gideon Rechavi, et al. Jepa-dna: Grounding ge- nomic foundation models through joint-embedding predictive architectures.arXiv preprint arXiv:2602.17162, 2026

  29. [29]

    Medical radiology report generation: A systematic review of current deep learning methods, trends, and future directions.Artificial intelligence in medicine, page 103220, 2025

    Amaan Izhar, Norisma Idris, and Nurul Japar. Medical radiology report generation: A systematic review of current deep learning methods, trends, and future directions.Artificial intelligence in medicine, page 103220, 2025

  30. [30]

    Raptor: Scalable train-free embeddings for 3d medical volumes leveraging pretrained 2d foundation models.arXiv preprint arXiv:2507.08254, 2025

    Ulzee An, Moonseong Jeong, Simon A Lee, Aditya Gorla, Yuzhe Yang, and Sriram Sankarara- man. Raptor: Scalable train-free embeddings for 3d medical volumes leveraging pretrained 2d foundation models.arXiv preprint arXiv:2507.08254, 2025

  31. [31]

    Seq vs seq: An open suite of paired encoders and decoders.arXiv preprint arXiv:2507.11412, 2025

    Orion Weller, Kathryn Ricci, Marc Marone, Antoine Chaffin, Dawn Lawrie, and Benjamin Van Durme. Seq vs seq: An open suite of paired encoders and decoders.arXiv preprint arXiv:2507.11412, 2025

  32. [32]

    A survey of large language models

    Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al. A survey of large language models. arXiv preprint arXiv:2303.18223, 1(2), 2023

  33. [33]

    Neuroseg meets dinov3: Transferring 2d self- supervised visual priors to 3d neuron segmentation via dinov3 initialization

    Yik San Cheng, Runkai Zhao, and Weidong Cai. Neuroseg meets dinov3: Transferring 2d self- supervised visual priors to 3d neuron segmentation via dinov3 initialization. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 30053–30064, 2026

  34. [34]

    Clinical decision support using pseudo-notes from multiple streams of ehr data.npj Digital Medicine, 8(1):394, July 2025

    Simon A Lee, Sujay Jain, Alex Chen, Kyoka Ono, Arabdha Biswas, Ákos Rudas, Jennifer Fang, and Jeffrey N Chiang. Clinical decision support using pseudo-notes from multiple streams of ehr data.npj Digital Medicine, 8(1):394, July 2025

  35. [35]

    Towards on-device foundation models for raw wearable signals

    Simon A Lee, Cyrus Tanade, Hao Zhou, Juhyeon Lee, Megha Thukral, Baiying Lu, and Sharanya Arcot Desai. Towards on-device foundation models for raw wearable signals. In NeurIPS 2025 Workshop on Learning from Time Series for Health, 2025

  36. [36]

    Ccnet: Criss-cross attention for semantic segmentation

    Zilong Huang, Xinggang Wang, Lichao Huang, Chang Huang, Yunchao Wei, and Wenyu Liu. Ccnet: Criss-cross attention for semantic segmentation. InProceedings of the IEEE/CVF international conference on computer vision, pages 603–612, 2019

  37. [37]

    Crossvit: Cross-attention multi- scale vision transformer for image classification

    Chun-Fu Richard Chen, Quanfu Fan, and Rameswar Panda. Crossvit: Cross-attention multi- scale vision transformer for image classification. InProceedings of the IEEE/CVF international conference on computer vision, pages 357–366, 2021

  38. [38]

    Cross attention network for few-shot classification.Advances in neural information processing systems, 32, 2019

    Ruibing Hou, Hong Chang, Bingpeng Ma, Shiguang Shan, and Xilin Chen. Cross attention network for few-shot classification.Advances in neural information processing systems, 32, 2019

  39. [39]

    A survey of large language models for tabular data imputation: Tuning paradigms and challenges

    Subham Jha, Varun Goyal, and Shweta Meena. A survey of large language models for tabular data imputation: Tuning paradigms and challenges. InInternational Conference on Data Analytics & Management, pages 224–240. Springer, 2025

  40. [40]

    Interpretable large language models for early prediction of antimicrobial multidrug resistance.Health Information Science and Systems, 14(1):11, 2025

    Lucía Carmona-Martos, Paula Martín-Palomeque, Óscar Escudero-Arnanz, and Cristina Soguero-Ruiz. Interpretable large language models for early prediction of antimicrobial multidrug resistance.Health Information Science and Systems, 14(1):11, 2025

  41. [41]

    M Rafly Rahman, Setio Basuki, Muhammad Ilham Perdana, and La Febry Andira Rose Cynthia. Integrating tabular data and textual representations for clinical risk prediction using machine learning and large language models.Kinetik: Game Technology, Information System, Computer Network, Computing, Electronics, and Control, 2026. 13

  42. [42]

    A foundation model for wearable movement data in mental health research.IEEE Journal of Biomedical and Health Informatics, 2026

    Franklin Y Ruan, Aiwei Zhang, Jenny Y Oh, SouYoung Jin, and Nicholas C Jacobson. A foundation model for wearable movement data in mental health research.IEEE Journal of Biomedical and Health Informatics, 2026

  43. [43]

    Physicase: Development and dual-layer validation of synthetic cases for health professional education: A pilot study leveraging generative ai.medRxiv, pages 2026–06, 2026

    Oyindolapo O Komolafe, Angela C Roberts, Jacob Shelley, and Andrews K Tawiah. Physicase: Development and dual-layer validation of synthetic cases for health professional education: A pilot study leveraging generative ai.medRxiv, pages 2026–06, 2026

  44. [44]

    The multimodal paradox: how added and missing modalities shape bias and performance in multimodal ai

    Kishore Sampath, Ayaazuddin Mohammad, Resmi Ramachandranpillai, et al. The multimodal paradox: how added and missing modalities shape bias and performance in multimodal ai. arXiv preprint arXiv:2505.03020, 2025

  45. [45]

    A case study exploring the current landscape of synthetic medical record generation with commercial llms.arXiv preprint arXiv:2504.14657, 2025

    Yihan Lin, Zhirong Bella Yu, and Simon Lee. A case study exploring the current landscape of synthetic medical record generation with commercial llms.arXiv preprint arXiv:2504.14657, 2025

  46. [46]

    Training-free zero-shot anomaly detection in 3d brain mri with 2d foundation models.arXiv preprint arXiv:2602.15315, 2026

    Tai Le-Gia and Jaehyun Ahn. Training-free zero-shot anomaly detection in 3d brain mri with 2d foundation models.arXiv preprint arXiv:2602.15315, 2026

  47. [47]

    Beyond english benchmarks: clinical llm evaluation in brazilian portuguese.arXiv preprint arXiv:2606.07853, 2026

    Giordano de Pinho Souza, Glaucia Melo, Josefino Cabral Melo Lima, and Daniel Schneider. Beyond english benchmarks: clinical llm evaluation in brazilian portuguese.arXiv preprint arXiv:2606.07853, 2026

  48. [48]

    Measuring the gap: correlating synthetic-to-real drift with phi de-identification performance.Genomics & Informatics, 24(1):10, 2026

    Joseph Cornelius and Fabio Rinaldi. Measuring the gap: correlating synthetic-to-real drift with phi de-identification performance.Genomics & Informatics, 24(1):10, 2026

  49. [49]

    On the problem of consistent anomalies in zero-shot anomaly detection.arXiv preprint arXiv:2512.02520, 2025

    Tai Le-Gia. On the problem of consistent anomalies in zero-shot anomaly detection.arXiv preprint arXiv:2512.02520, 2025

  50. [50]

    Reuniting forcibly separated families through shared memories with machine learning.Available at SSRN, 2025

    Huifeng Su, Lesley Meng, and Edieal J Pinker. Reuniting forcibly separated families through shared memories with machine learning.Available at SSRN, 2025

  51. [51]

    Semantic insurance pricing with large language models.arXiv preprint arXiv:2606.29371, 2026

    Christopher Blier-Wong and Derek Kusmenko. Semantic insurance pricing with large language models.arXiv preprint arXiv:2606.29371, 2026

  52. [52]

    Ethical framework for responsible foundational models in medical imaging.Frontiers in Medicine, 12:1544501, 2025

    Debesh Jha, Gorkem Durak, Abhijit Das, Jasmer Sanjotra, Onkar Susladkar, Suramyaa Sarkar, Ashish Rauniyar, Nikhil Kumar Tomar, Linkai Peng, Sirui Li, et al. Ethical framework for responsible foundational models in medical imaging.Frontiers in Medicine, 12:1544501, 2025

  53. [53]

    One loss to rule them all: Marked time-to-event for structured ehr foundation models.arXiv preprint arXiv:2602.00541, 2026

    Zilin Jing, Vincent Jeanselme, Yuta Kobayashi, Simon A Lee, Chao Pang, Aparajita Kashyap, Yanwei Li, Xinzhuo Jiang, and Shalmali Joshi. One loss to rule them all: Marked time-to-event for structured ehr foundation models.arXiv preprint arXiv:2602.00541, 2026

  54. [54]

    Ehr foundation models improve robustness in the presence of temporal distribution shift.Scientific Reports, 13(1):3767, 2023

    Lin Lawrence Guo, Ethan Steinberg, Scott Lanyon Fleming, Jose Posada, Joshua Lemmon, Stephen R Pfohl, Nigam Shah, Jason Fries, and Lillian Sung. Ehr foundation models improve robustness in the presence of temporal distribution shift.Scientific Reports, 13(1):3767, 2023

  55. [55]

    Ehr-r1: A reasoning-enhanced foundational language model for electronic health record analysis, 2025

    Yusheng Liao, Chaoyi Wu, Junwei Liu, Shuyang Jiang, Pengcheng Qiu, Haowen Wang, Yun Yue, Shuai Zhen, Jian Wang, Qianrui Fan, Jinjie Gu, Ya Zhang, Yanfeng Wang, Yu Wang, and Weidi Xie. Ehr-r1: A reasoning-enhanced foundational language model for electronic health record analysis, 2025. URLhttps://arxiv.org/abs/2510.25628

  56. [56]

    Ehragent: Code empowers large language models for few-shot complex tabular reasoning on electronic health records

    Wenqi Shi, Ran Xu, Yuchen Zhuang, Yue Yu, Jieyu Zhang, Hang Wu, Yuanda Zhu, Joyce C Ho, Carl Yang, and May Dongmei Wang. Ehragent: Code empowers large language models for few-shot complex tabular reasoning on electronic health records. InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 22315–22339, 2024

  57. [57]

    Learning the natural history of human disease with generative transformers.Nature, 647(8088):248–256, 2025

    Artem Shmatko, Alexander Wolfgang Jung, Kumar Gaurav, Søren Brunak, Laust Hvas Mortensen, Ewan Birney, Tom Fitzgerald, and Moritz Gerstung. Learning the natural history of human disease with generative transformers.Nature, 647(8088):248–256, 2025. 14

  58. [58]

    Context clues: Evaluating long context models for clinical prediction tasks on ehrs.arXiv preprint arXiv:2412.16178, 2024

    Michael Wornow, Suhana Bedi, Miguel Angel Fuentes Hernandez, Ethan Steinberg, Jason Alan Fries, Christopher Ré, Sanmi Koyejo, and Nigam H Shah. Context clues: Evaluating long context models for clinical prediction tasks on ehrs.arXiv preprint arXiv:2412.16178, 2024

  59. [59]

    Core-behrt: A carefully optimized and rigorously evaluated behrt

    Mikkel Odgaard, Kiril Vadimovic Klein, Sanne Møller Thysen, Espen Jimenez-Solem, Martin Sillesen, and Mads Nielsen. Core-behrt: A carefully optimized and rigorously evaluated behrt. arXiv preprint arXiv:2404.15201, 2024

  60. [60]

    Med-bert: pretrained contextu- alized embeddings on large-scale structured electronic health records for disease prediction

    Laila Rasmy, Yang Xiang, Ziqian Xie, Cui Tao, and Degui Zhi. Med-bert: pretrained contextu- alized embeddings on large-scale structured electronic health records for disease prediction. NPJ digital medicine, 4(1):86, 2021

  61. [61]

    Language models are an effective representation learning technique for electronic health record data.Journal of biomedical informatics, 113:103637, 2021

    Ethan Steinberg, Ken Jung, Jason A Fries, Conor K Corbin, Stephen R Pfohl, and Nigam H Shah. Language models are an effective representation learning technique for electronic health record data.Journal of biomedical informatics, 113:103637, 2021

  62. [62]

    Ehrshot: An ehr benchmark for few-shot evaluation of foundation models.Advances in Neural Information Processing Systems, 36, 2024

    Michael Wornow, Rahul Thapa, Ethan Steinberg, Jason Fries, and Nigam Shah. Ehrshot: An ehr benchmark for few-shot evaluation of foundation models.Advances in Neural Information Processing Systems, 36, 2024

  63. [63]

    Context clues: Evaluating long context models for clinical prediction tasks on ehr data

    Michael Wornow, Suhana Bedi, Miguel Angel Fuentes Hernandez, Ethan Steinberg, Jason Alan Fries, Christopher Re, Sanmi Koyejo, and Nigam Shah. Context clues: Evaluating long context models for clinical prediction tasks on ehr data. InThe Thirteenth International Conference on Learning Representations, 2025

  64. [64]

    Evidence extraction for automated medical coding: preliminary evaluation

    Xiaorui Jiang, Kulsoom Khan, Sumithra Thinakara Vasantha, and Sajjad Haider. Evidence extraction for automated medical coding: preliminary evaluation. InProceedings of the 2024 8th International Conference on Natural Language Processing and Information Retrieval, pages 18–23, 2024

  65. [65]

    Utilizing large language models to predict icd-10 diagnosis codes from patient medical records

    Rudransh Pathak, Gabriel Vald, Yusuf Sermet, and Ibrahim Demir. Utilizing large language models to predict icd-10 diagnosis codes from patient medical records. In2024 IEEE MIT Undergraduate Research Technology Conference (URTC), pages 1–5. IEEE, 2024

  66. [66]

    Can large language models abstract medical coded language?arXiv preprint arXiv:2403.10822, 2024

    Simon A Lee and Timothy Lindsey. Can large language models abstract medical coded language?arXiv preprint arXiv:2403.10822, 2024

  67. [67]

    Clinical modernbert: An efficient and long context encoder for biomedical text.arXiv preprint arXiv:2504.03964, 2025

    Simon A Lee, Anthony Wu, and Jeffrey N Chiang. Clinical modernbert: An efficient and long context encoder for biomedical text.arXiv preprint arXiv:2504.03964, 2025

  68. [68]

    Clinical text embeddings: A systematic review of methods, applications, and future directions.International Journal of Medical Informatics, page 106505, 2026

    Hyunwoo Jo, Seon Kim, Hyunwoo Son, and Jongchan Kim. Clinical text embeddings: A systematic review of methods, applications, and future directions.International Journal of Medical Informatics, page 106505, 2026

  69. [69]

    Aicoder: Exploring automated icd coding on chinese emrs with a multi-agent framework

    Zhenpeng Liang, Hongjiao Guan, Weiyu Zhang, Ying Lian, Bing Xu, and Wenpeng Lu. Aicoder: Exploring automated icd coding on chinese emrs with a multi-agent framework. In2025 IEEE International Conference on Bioinformatics and Biomedicine (BIBM), pages 3823–3827. IEEE, 2025

  70. [70]

    Gate: Graph and text exchange for zero-shot ecg classification with llm prompts.IEEE Journal of Biomedical and Health Informatics, 2026

    Ying An, Shiyu Tang, Xianlai Chen, and Lin Guo. Gate: Graph and text exchange for zero-shot ecg classification with llm prompts.IEEE Journal of Biomedical and Health Informatics, 2026

  71. [71]

    Cp-env: Evaluating large language models on clinical pathways in a controllable hospital environment.arXiv preprint arXiv:2512.10206, 2025

    Yakun Zhu, Zhongzhen Huang, Qianhan Feng, Linjie Mu, Yannian Gu, Shaoting Zhang, Qi Dou, and Xiaofan Zhang. Cp-env: Evaluating large language models on clinical pathways in a controllable hospital environment.arXiv preprint arXiv:2512.10206, 2025

  72. [72]

    An analytical review of optimization techniques in information retrieval for enhanced decision support.Decision Analytics Journal, page 100657, 2025

    Kemal Lazovi´c, Filipe Madeira, Eftim Zdravevski, Luis Augusto Silva, Paulo Jorge Coelho, and Ivan Miguel Pires. An analytical review of optimization techniques in information retrieval for enhanced decision support.Decision Analytics Journal, page 100657, 2025

  73. [73]

    Reinventing clinical dialogue: Agentic paradigms for llm enabled healthcare communication.arXiv preprint arXiv:2512.01453, 2025

    Xiaoquan Zhi, Hongke Zhao, Likang Wu, Chuang Zhao, and Hengshu Zhu. Reinventing clinical dialogue: Agentic paradigms for llm enabled healthcare communication.arXiv preprint arXiv:2512.01453, 2025. 15

  74. [74]

    Oncopt: long-context transformer models for in hospital tumor phenotype extraction from pathology reports.npj Digital Medicine, 2026

    Thanh Duong, Dung Le, V onetta Williams, Sandra Stewart, Yayi Zhao, Muntasir Zitu, Issam El Naqa, Dana Rollison, and Thanh Thieu. Oncopt: long-context transformer models for in hospital tumor phenotype extraction from pathology reports.npj Digital Medicine, 2026

  75. [75]

    Emerging properties in self-supervised vision transformers

    Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerging properties in self-supervised vision transformers. InProceedings of the International Conference on Computer Vision (ICCV), 2021

  76. [76]

    Dinov2: Learning robust visual features without supervision.arXiv preprint arXiv:2304.07193, 2023

    Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy V o, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, et al. Dinov2: Learning robust visual features without supervision.arXiv preprint arXiv:2304.07193, 2023

  77. [77]

    Learning transferable visual models from natural language supervision

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. InInternational conference on machine learning, pages 8748–8763. PMLR, 2021

  78. [78]

    High-resolution image synthesis with latent diffusion models, 2021

    Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models, 2021

  79. [79]

    Photorealistic text-to-image diffusion models with deep language understanding.Advances in neural information processing systems, 35:36479–36494, 2022

    Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L Denton, Kamyar Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, et al. Photorealistic text-to-image diffusion models with deep language understanding.Advances in neural information processing systems, 35:36479–36494, 2022

  80. [80]

    Contrastive learning of preferences with a contextual infonce loss, 2024

    Timo Bertram, Johannes Fürnkranz, and Martin Müller. Contrastive learning of preferences with a contextual infonce loss, 2024. URLhttps://arxiv.org/abs/2407.05898

Showing first 80 references.