REVIEW 3 major objections 5 minor 105 references
Pretraining that forces models to recover waveform shape, not just signal values, yields more transferable ECG and pulse-oximetry representations than reconstruction or generic self-supervised learning.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-14 16:30 UTC pith:QV5YXRCU
load-bearing objection Solid multimodal SSL recipe for ECG+SpO2 with consistent MIMIC gains; the morphology-masking story is the real claim and is under-specified, but the paper is still worth a referee. the 3 major comments →
MorphologyFM: A Foundation Model for Morphology-Aware Representation Learning from ECG and Pulse Oximetry Waveforms
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Large-scale self-supervised pretraining that explicitly preserves clinically meaningful waveform morphology on paired ECG and SpO2 produces more transferable physiological representations than objectives based primarily on reconstruction or generic embedding alignment. MorphologyFM combines morphology-guided masking of complete waveform components, cross-modal latent alignment of synchronized ECG–SpO2 windows, and contrastive organization of morphologically similar segments; the resulting encoder improves multiple downstream clinical prediction tasks and benefits from multimodal joint pretraining and scale of unlabeled waveforms.
What carries the argument
Morphology-aware masking: instead of random sample masks, approximate cardiac cycles (R-peaks) and aligned pulse peaks define masks over whole ECG components (P, QRS, ST, T) and SpO2 regions (systolic upstroke, diastolic decay, pulse peak, respiratory modulation), forcing the encoder to infer higher-order physiological structure. This is trained jointly with masked reconstruction, cross-modal alignment, and contrastive morphology objectives.
Load-bearing premise
The method assumes that automatic peak detection and masking of complete morphological components on noisy ICU waveforms reliably isolates the clinically meaningful structure that drives the reported gains, rather than other unmeasured training differences.
What would settle it
Re-run the same backbone, data, and protocol with morphology-aware masking replaced by random or temporal-block masking at matched mask rates and augmentations; if average downstream AUROC/F1 no longer favors MorphologyFM over MAE and JEPA (Table 1 vs Table 3 setup), the morphology-inductive-bias claim fails.
If this is right
- Waveform foundation models should treat morphology preservation as a first-class pretraining objective, not only reconstruction or forecasting fidelity.
- Paired ECG–SpO2 pretraining should yield more transferable cardiovascular representations than single-modality pretraining of comparable size.
- Downstream arrhythmia, hypoxemia, mortality, and length-of-stay models can be improved by fine-tuning a morphology-pretrained encoder with less labeled data than task-specific supervised training from scratch.
- Scaling unlabeled paired waveforms (hundreds of thousands to millions of segments) should continue to improve transfer until diminishing returns set in.
- Patient-similarity retrieval in the latent space should become more clinically coherent when embeddings are organized by morphology rather than reconstruction alone.
Where Pith is reading between the lines
- If component-level masking is the real driver, the same idea should transfer to other periodic clinical waveforms (arterial pressure, capnography, EEG envelopes) once reliable landmark detectors exist.
- Brittle R-peak or pulse-peak detection in artifact-heavy ICU data would systematically weaken the claimed inductive bias; robustness of the masker is as important as the transformer backbone.
- A natural next test is whether morphology-pretrained embeddings improve early-warning tasks that clinicians already read from shape (ST change, pulse-pressure variation) under device and hospital shift.
- Joint latent alignment of synchronized modalities may act as free multi-view supervision that partly substitutes for labels in critical-care monitoring.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. MorphologyFM is a multimodal foundation model for ECG and SpO2 waveforms pretrained on MIMIC with a morphology-aware self-supervised objective. The method combines morphology-guided masking of complete ECG components (P/QRS/ST/T) and SpO2 regions (systolic/diastolic/respiratory), cross-modal latent alignment of synchronized pairs, and a contrastive morphology objective, jointly optimized with a reconstruction loss over masked regions only. After pretraining, the encoder is fine-tuned on arrhythmia classification, hypoxemia prediction, mortality prediction, and length-of-stay estimation. Against MAE, SimCLR-style contrastive learning, Barlow Twins, and JEPA under a shared transformer backbone and fine-tuning protocol, MorphologyFM reports the strongest average transfer (89.6 vs. 83.8–86.5 in Table 1), with further gains from joint ECG+SpO2 pretraining (Table 2), morphology-aware vs. random/temporal-block masking (Table 3), full multi-term objective (Table 4), nearest-neighbor retrieval (Table 5), and scaling of unlabeled pretraining data (Table 6).
Significance. If the gains are truly attributable to morphology as an inductive bias rather than unreported protocol differences, the paper would strengthen the case for structure-preserving SSL on continuous physiological signals and provide a practical multimodal backbone for ICU monitoring. The experimental design is useful: same backbone and fine-tuning protocol across MAE, contrastive, Barlow Twins, and JEPA; multimodal vs. single-modality comparison; masking and objective ablations; retrieval and scaling analyses. These elements make the work a credible contribution to physiological foundation models even if some causal claims need tighter support. The absence of released code or external validation is a limitation but does not erase the empirical pattern on MIMIC.
major comments (3)
- [§3.4 / Table 3] §3.4 and Table 3: The central claim that morphology-aware masking of complete P/QRS/ST/T and SpO2 systolic/diastolic/respiratory regions is the inductive bias behind the gains is under-specified. The paper does not report R-peak/component detection success rates or failure modes on noisy MIMIC ICU waveforms, nor the realized fraction of samples masked under morphology-aware vs. random vs. temporal-block policies. Without those controls, Table 3 is consistent with a different effective mask rate or difficulty rather than a morphology-specific effect, which is load-bearing for the title and strongest claim.
- [§3.8 / §4.1] §3.8 and §4.1: Free parameters that define the method—λ_rec, λ_con, λ_align, InfoNCE temperature τ, morphology-component masking probabilities, and the augmentation mixture—are not given numerical values or selection criteria. The claim that identical backbone, optimization schedule, and fine-tuning protocol isolate the pretraining objective therefore cannot be verified from the manuscript. At minimum, report the hyperparameter settings used for MorphologyFM and confirm that baselines received comparable tuning budgets.
- [Table 1 / §4.1] Table 1 and §4.1: Length of stay is described as estimation yet scored with AUROC, which requires an unstated binarization (or multi-class reduction). Also, results are averaged over five independent runs with no standard deviations, confidence intervals, or significance tests. Both issues weaken the quantitative support for the ranking MorphologyFM > JEPA > Barlow > contrastive > MAE.
minor comments (5)
- [References] Related work cites several arXiv items with 2025–2026 dates and placeholder-style entries (e.g., [25?]); clean the bibliography for consistency and verifiability.
- [§3.2] §3.2: State the common sampling frequency, window length in samples/seconds, and exact signal-quality discard criteria so the pretraining corpus is reproducible.
- [§3.3] §3.3: Specify transformer depth, width, number of heads, patch size, and whether the class token is learned or a mean pool; Table 2 parameter counts alone are insufficient.
- [Abstract / Table 6] Abstract and §1 claim ‘millions of waveform segments’ while Table 6 tops out at 10M; align the prose with the scaling table and state the exact pretraining corpus size used for the main results.
- [Table 2] Table 2 reports ‘Average AUROC’ while Table 1 mixes F1 and AUROC; clarify how the average is computed when metrics differ across tasks.
Circularity Check
No significant circularity: MorphologyFM is empirical SSL evaluated on external supervised labels, not a derivation that redefines its target.
full rationale
The paper’s load-bearing claims are experimental: morphology-guided masking, cross-modal alignment, and contrastive objectives are pretrained on unlabeled MIMIC ECG/SpO2, then transferred to separate supervised tasks (arrhythmia F1, hypoxemia/mortality/LOS AUROC) against MAE, contrastive, Barlow Twins, and JEPA with a shared backbone (Tables 1–6). Downstream metrics are not algebraically forced by the pretraining loss (L = λ_rec L_rec + λ_con L_con + λ_align L_align); reconstruction is restricted to masked regions, contrastive pairs come from morphology-preserving augmentations, and alignment uses synchronized modality pairs—none of these redefine the clinical labels used at evaluation. There is no self-definitional loop (morphology is operationalized via R-peak/component masking, not via the reported AUROC/F1), no fitted-input-called-prediction (pretraining does not fit to the downstream targets), no uniqueness theorem imported from the authors, and no ansatz smuggled in as a forced mathematical result. The single same-author citation (Chronoformer [87]) is background temporal modeling and is not load-bearing for the MorphologyFM objective or tables. Weaknesses (under-specified mask rates, detection reliability, hyperparameter schedules) are correctness/confounding risks, not circularity. The experimental chain is self-contained against external task labels.
Axiom & Free-Parameter Ledger
free parameters (5)
- λ_rec, λ_con, λ_align (joint loss weights)
- InfoNCE temperature τ
- Morphology-component masking probabilities
- Augmentation mixture (noise, stretch, scale, crop, baseline wander)
- Transformer/tokenizer capacity and optimization schedule
axioms (5)
- domain assumption Clinically meaningful information in ECG/SpO2 is primarily morphological (shape, intervals, slopes, beat variability) rather than sample-wise amplitude continuity.
- domain assumption Automatic R-peak detection and SpO2 peak alignment can identify complete morphological components well enough for masking on ICU data.
- domain assumption MIMIC critical-care paired ECG–SpO2 segments are a representative unlabeled corpus for transferable physiological representations.
- ad hoc to paper Same transformer backbone + same fine-tuning protocol isolates pretraining-objective effects across MAE, contrastive, Barlow Twins, JEPA, and MorphologyFM.
- standard math Standard SSL constructions (masked reconstruction, InfoNCE, latent multimodal alignment) are valid representation-learning objectives.
invented entities (2)
-
MorphologyFM (morphology-aware multimodal waveform foundation model)
no independent evidence
-
Morphology-aware masking operator over ECG/SpO2 components
no independent evidence
read the original abstract
Foundation models have recently emerged as a powerful paradigm for learning transferable representations from large scale biomedical data, yet existing approaches for physiological waveforms primarily optimize reconstruction or forecasting objectives that do not explicitly preserve clinically meaningful waveform morphology. Electrocardiograms (ECGs) and pulse oximetry (SpO2) waveforms encode rich cardiovascular and hemodynamic information through their morphological structure. In this work, we introduce MorphologyFM, a multimodal foundation model pretrained on paired ECG and SpO2 waveforms from the MIMIC critical care database using a morphology aware self supervised learning objective. MorphologyFM combines morphology guided masking, cross modal representation learning, and contrastive latent alignment to learn representations that capture clinically relevant physiological structure without requiring manual annotations. We evaluate MorphologyFM across multiple downstream prediction tasks, including arrhythmia classification, hypoxemia prediction, mortality prediction, and length of stay estimation, demonstrating consistent improvements over representative self supervised learning methods, including Masked Autoencoders (MAE), contrastive learning, Barlow Twins, and Joint Embedding Predictive Architectures (JEPA). Furthermore, we show that jointly modeling ECG and SpO2 waveforms produces more transferable representations than single modality pretraining. Our results establish waveform morphology as a powerful inductive bias for self supervised physiological representation learning and introduce MorphologyFM as a general purpose foundation model for continuous physiological monitoring.
Reference graph
Works this paper leans on
-
[1]
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. InProceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers), pages 4171–4186, 2019
2019
-
[2]
Masked autoencoders are scalable vision learners
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick. Masked autoencoders are scalable vision learners. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 16000–16009, 2022
2022
-
[3]
Deep residual learning for image recognition, 2015
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition, 2015. URLhttps://arxiv.org/abs/1512.03385
Pith/arXiv arXiv 2015
-
[4]
Multi-scale 3d deep convolutional neural network for hyperspectral image classification
Mingyi He, Bo Li, and Huahui Chen. Multi-scale 3d deep convolutional neural network for hyperspectral image classification. In2017 IEEE International Conference on Image Processing (ICIP), pages 3904–3908. IEEE, 2017
2017
-
[5]
Bag of tricks for image classification with convolutional neural networks
Tong He, Zhi Zhang, Hang Zhang, Zhongyue Zhang, Junyuan Xie, and Mu Li. Bag of tricks for image classification with convolutional neural networks. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 558–567, 2019
2019
-
[6]
A foundational vision transformer improves diagnostic performance for electrocardiograms.NPJ Digital Medicine, 6(1):108, 2023
Akhil Vaid, Joy Jiang, Ashwin Sawant, Stamatios Lerakis, Edgar Argulian, Yuri Ahuja, Joshua Lampert, Alexander Charney, Hayit Greenspan, Jagat Narula, et al. A foundational vision transformer improves diagnostic performance for electrocardiograms.NPJ Digital Medicine, 6(1):108, 2023
2023
-
[7]
Foundation models in healthcare: Opportunities, risks & strategies forward
Anja Thieme, Aditya Nori, Marzyeh Ghassemi, Rishi Bommasani, Tariq Osman Andersen, and Ewa Luger. Foundation models in healthcare: Opportunities, risks & strategies forward. InExtended abstracts of the 2023 CHI conference on human factors in computing systems, pages 1–4, 2023
2023
-
[8]
Foundation models for time series forecasting.International IT Journal of Research, ISSN: 3007-6706, 2(4):144–156, 2024
Suresh Chandra Thakur. Foundation models for time series forecasting.International IT Journal of Research, ISSN: 3007-6706, 2(4):144–156, 2024
2024
-
[9]
A foun- dation model for intensive care: Unlocking generalization across tasks and domains at scale
Manuel Burger, Daphné Chopard, Malte Londschien, Fedor Sergeev, Hugo Yèche, Rita Kuznetsova, Martin Faltys, Eike Gerdes, Polina Leshetkina, Peter Bühlmann, et al. A foun- dation model for intensive care: Unlocking generalization across tasks and domains at scale. medRxiv, pages 2025–07, 2025
2025
-
[10]
Foundation models defining a new era in vision: a survey and outlook.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025
Muhammad Awais, Muzammal Naseer, Salman Khan, Rao Muhammad Anwer, Hisham Cholakkal, Mubarak Shah, Ming-Hsuan Yang, and Fahad Shahbaz Khan. Foundation models defining a new era in vision: a survey and outlook.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025
2025
-
[11]
Foundation model for advancing healthcare: challenges, opportunities and future directions.IEEE Reviews in Biomedical Engineering, 2024
Yuting He, Fuxiang Huang, Xinrui Jiang, Yuxiang Nie, Minghao Wang, Jiguang Wang, and Hao Chen. Foundation model for advancing healthcare: challenges, opportunities and future directions.IEEE Reviews in Biomedical Engineering, 2024
2024
-
[12]
Michael C Burkhart, Bashar Ramadan, Zewei Liao, Kaveri Chhikara, Juan C Rojas, William F Parker, and Brett K Beaulieu-Jones. Foundation models for electronic health records: repre- sentation dynamics and transferability.arXiv preprint arXiv:2504.10422, 2025. 11
Pith/arXiv arXiv 2025
-
[13]
Foundation models for time series analysis: A tutorial and survey
Yuxuan Liang, Haomin Wen, Yuqi Nie, Yushan Jiang, Ming Jin, Dongjin Song, Shirui Pan, and Qingsong Wen. Foundation models for time series analysis: A tutorial and survey. In Proceedings of the 30th ACM SIGKDD conference on knowledge discovery and data mining, pages 6555–6565, 2024
2024
-
[14]
Foundation models in bioinformatics.National science review, 12(4):nwaf028, 2025
Fei Guo, Renchu Guan, Yaohang Li, Qi Liu, Xiaowo Wang, Can Yang, and Jianxin Wang. Foundation models in bioinformatics.National science review, 12(4):nwaf028, 2025
2025
-
[15]
Hao Zhou, Simon A Lee, Cyrus Tanade, Keum San Chun, Juhyeon Lee, Migyeong Gwak, Megha Thukral, Justin Sung, Eugene Hwang, Mehrab Bin Morshed, et al. Physiology-aware masked cross-modal reconstruction for biosignal representation learning.arXiv preprint arXiv:2605.00973, May 2026
Pith/arXiv arXiv 2026
-
[16]
Large-scale training of foundation models for wearable biosignals
Salar Abbaspourazad, Oussama Elachqar, Andrew C Miller, Saba Emrani, Udhyakumar Nallasamy, and Ian Shapiro. Large-scale training of foundation models for wearable biosignals. arXiv preprint arXiv:2312.05409, 2023
Pith/arXiv arXiv 2023
-
[17]
Sonata: A hybrid world model for inertial kinematics under clinical data scarcity
Blaise Delaney, Salil Patel, Yuji Xing, Dominic Dootson, Karin Sevegnani, and Chrystalina Antoniades. Sonata: A hybrid world model for inertial kinematics under clinical data scarcity. arXiv preprint arXiv:2604.18058, 2026
Pith/arXiv arXiv 2026
-
[18]
Edge ppg quality gating with validity-labeled outputs and dual-channel communication on a raspberry pi zero 2 w wearable device
Haoyang Zhou, Yi Ding, Jiajun Li, Yifan Yi, and Chun Xiao. Edge ppg quality gating with validity-labeled outputs and dual-channel communication on a raspberry pi zero 2 w wearable device. InJournal of Physics: Conference Series, volume 3235, page 012025. IOP Publishing, 2026
2026
-
[19]
Simon A Lee, Cyrus Tanade, Hao Zhou, Juhyeon Lee, Megha Thukral, Minji Han, Rachel Choi, Md Sazzad Hissain Khan, Baiying Lu, Migyeong Gwak, et al. Himae: Hierarchical masked autoencoders discover resolution-specific structure in wearable time series.arXiv preprint arXiv:2510.25785, 2025
arXiv 2025
-
[20]
Multimodal self-supervised learning for wearable sleep staging using photoplethysmography and accelerometer signals
Juhyeon Lee, Simon A Lee, Cyrus Tanade, Viswam Nathan, Megha Thukral, Hao Zhou, Keum San Chun, and Sharanya Arcot Desai. Multimodal self-supervised learning for wearable sleep staging using photoplethysmography and accelerometer signals. InICASSP 2026-2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 7852–7856. ...
2026
-
[21]
Zechen Li, Keerthana Natarajan, Weizhi Zhang, Menglian Zhou, Simon A Lee, Yuwei Zhang, Maxwell A Xu, Zeinab Esmaeilpour, Flora D Salim, Mark Malhotra, et al. Glucofm: A dual- stream foundation model for continuous glucose monitoring.arXiv preprint arXiv:2605.30865, 2026
Pith/arXiv arXiv 2026
-
[22]
Trust and transparency in healthcare ai: A systematic review of explainable nlp for clinical decision support (2023–2025)
MS Tharini and Jane Rubel Angelina Jeyaraj. Trust and transparency in healthcare ai: A systematic review of explainable nlp for clinical decision support (2023–2025). InInternational Conference on Edge Computing and Applications, pages 451–470. Springer, 2025
2023
-
[23]
Smart assistive technologies for neurodisorders: A review on ai, iot, and wearable systems for enhanced patient care.Neurological Sciences, 47(2):211, 2026
Sandeep Chouhan, Deepika Ghai, Ramandeep Sandhu, and Suman Lata Tripathi. Smart assistive technologies for neurodisorders: A review on ai, iot, and wearable systems for enhanced patient care.Neurological Sciences, 47(2):211, 2026
2026
-
[24]
A preoperative data sentenceization method for postoperative major adverse cardiovascular event prediction
Yuelin Luo, Yaqiang Wang, Xiran Peng, Ruihao Zhou, Guo Chen, Xuechao Hao, and Tao Zhu. A preoperative data sentenceization method for postoperative major adverse cardiovascular event prediction. In2025 IEEE International Conference on Big Data (BigData), pages 8060–8069. IEEE, 2025
2025
-
[25]
A large language model for electronic health records.NPJ digital medicine, 5(1):194, 2022
Xi Yang, Aokun Chen, Nima PourNejatian, Hoo Chang Shin, Kaleb E Smith, Christopher Parisien, Colin Compas, Cheryl Martin, Anthony B Costa, Mona G Flores, et al. A large language model for electronic health records.NPJ digital medicine, 5(1):194, 2022
2022
-
[26]
Foundation models for physiological signals: Opportunities and challenges
Simon A Lee and Kai Akamatsu. Foundation models for physiological signals: Opportunities and challenges. August 2025. 12
2025
-
[27]
The shaky foundations of large language models and foundation models for electronic health records.npj digital medicine, 6(1):135, 2023
Michael Wornow, Yizhe Xu, Rahul Thapa, Birju Patel, Ethan Steinberg, Scott Fleming, Michael A Pfeffer, Jason Fries, and Nigam H Shah. The shaky foundations of large language models and foundation models for electronic health records.npj digital medicine, 6(1):135, 2023
2023
-
[28]
Ariel Larey, Elay Dahan, Amit Bleiweiss, Raizy Kellerman, Guy Leib, Omri Nayshool, Dan Ofer, Tal Zinger, Dan Dominissini, Gideon Rechavi, et al. Jepa-dna: Grounding ge- nomic foundation models through joint-embedding predictive architectures.arXiv preprint arXiv:2602.17162, 2026
Pith/arXiv arXiv 2026
-
[29]
Medical radiology report generation: A systematic review of current deep learning methods, trends, and future directions.Artificial intelligence in medicine, page 103220, 2025
Amaan Izhar, Norisma Idris, and Nurul Japar. Medical radiology report generation: A systematic review of current deep learning methods, trends, and future directions.Artificial intelligence in medicine, page 103220, 2025
2025
-
[30]
Ulzee An, Moonseong Jeong, Simon A Lee, Aditya Gorla, Yuzhe Yang, and Sriram Sankarara- man. Raptor: Scalable train-free embeddings for 3d medical volumes leveraging pretrained 2d foundation models.arXiv preprint arXiv:2507.08254, 2025
Pith/arXiv arXiv 2025
-
[31]
Seq vs seq: An open suite of paired encoders and decoders.arXiv preprint arXiv:2507.11412, 2025
Orion Weller, Kathryn Ricci, Marc Marone, Antoine Chaffin, Dawn Lawrie, and Benjamin Van Durme. Seq vs seq: An open suite of paired encoders and decoders.arXiv preprint arXiv:2507.11412, 2025
arXiv 2025
-
[32]
A survey of large language models
Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al. A survey of large language models. arXiv preprint arXiv:2303.18223, 1(2), 2023
Pith/arXiv arXiv 2023
-
[33]
Neuroseg meets dinov3: Transferring 2d self- supervised visual priors to 3d neuron segmentation via dinov3 initialization
Yik San Cheng, Runkai Zhao, and Weidong Cai. Neuroseg meets dinov3: Transferring 2d self- supervised visual priors to 3d neuron segmentation via dinov3 initialization. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 30053–30064, 2026
2026
-
[34]
Clinical decision support using pseudo-notes from multiple streams of ehr data.npj Digital Medicine, 8(1):394, July 2025
Simon A Lee, Sujay Jain, Alex Chen, Kyoka Ono, Arabdha Biswas, Ákos Rudas, Jennifer Fang, and Jeffrey N Chiang. Clinical decision support using pseudo-notes from multiple streams of ehr data.npj Digital Medicine, 8(1):394, July 2025
2025
-
[35]
Towards on-device foundation models for raw wearable signals
Simon A Lee, Cyrus Tanade, Hao Zhou, Juhyeon Lee, Megha Thukral, Baiying Lu, and Sharanya Arcot Desai. Towards on-device foundation models for raw wearable signals. In NeurIPS 2025 Workshop on Learning from Time Series for Health, 2025
2025
-
[36]
Ccnet: Criss-cross attention for semantic segmentation
Zilong Huang, Xinggang Wang, Lichao Huang, Chang Huang, Yunchao Wei, and Wenyu Liu. Ccnet: Criss-cross attention for semantic segmentation. InProceedings of the IEEE/CVF international conference on computer vision, pages 603–612, 2019
2019
-
[37]
Crossvit: Cross-attention multi- scale vision transformer for image classification
Chun-Fu Richard Chen, Quanfu Fan, and Rameswar Panda. Crossvit: Cross-attention multi- scale vision transformer for image classification. InProceedings of the IEEE/CVF international conference on computer vision, pages 357–366, 2021
2021
-
[38]
Cross attention network for few-shot classification.Advances in neural information processing systems, 32, 2019
Ruibing Hou, Hong Chang, Bingpeng Ma, Shiguang Shan, and Xilin Chen. Cross attention network for few-shot classification.Advances in neural information processing systems, 32, 2019
2019
-
[39]
A survey of large language models for tabular data imputation: Tuning paradigms and challenges
Subham Jha, Varun Goyal, and Shweta Meena. A survey of large language models for tabular data imputation: Tuning paradigms and challenges. InInternational Conference on Data Analytics & Management, pages 224–240. Springer, 2025
2025
-
[40]
Interpretable large language models for early prediction of antimicrobial multidrug resistance.Health Information Science and Systems, 14(1):11, 2025
Lucía Carmona-Martos, Paula Martín-Palomeque, Óscar Escudero-Arnanz, and Cristina Soguero-Ruiz. Interpretable large language models for early prediction of antimicrobial multidrug resistance.Health Information Science and Systems, 14(1):11, 2025
2025
-
[41]
M Rafly Rahman, Setio Basuki, Muhammad Ilham Perdana, and La Febry Andira Rose Cynthia. Integrating tabular data and textual representations for clinical risk prediction using machine learning and large language models.Kinetik: Game Technology, Information System, Computer Network, Computing, Electronics, and Control, 2026. 13
2026
-
[42]
A foundation model for wearable movement data in mental health research.IEEE Journal of Biomedical and Health Informatics, 2026
Franklin Y Ruan, Aiwei Zhang, Jenny Y Oh, SouYoung Jin, and Nicholas C Jacobson. A foundation model for wearable movement data in mental health research.IEEE Journal of Biomedical and Health Informatics, 2026
2026
-
[43]
Physicase: Development and dual-layer validation of synthetic cases for health professional education: A pilot study leveraging generative ai.medRxiv, pages 2026–06, 2026
Oyindolapo O Komolafe, Angela C Roberts, Jacob Shelley, and Andrews K Tawiah. Physicase: Development and dual-layer validation of synthetic cases for health professional education: A pilot study leveraging generative ai.medRxiv, pages 2026–06, 2026
2026
-
[44]
The multimodal paradox: how added and missing modalities shape bias and performance in multimodal ai
Kishore Sampath, Ayaazuddin Mohammad, Resmi Ramachandranpillai, et al. The multimodal paradox: how added and missing modalities shape bias and performance in multimodal ai. arXiv preprint arXiv:2505.03020, 2025
Pith/arXiv arXiv 2025
-
[45]
Yihan Lin, Zhirong Bella Yu, and Simon Lee. A case study exploring the current landscape of synthetic medical record generation with commercial llms.arXiv preprint arXiv:2504.14657, 2025
Pith/arXiv arXiv 2025
-
[46]
Tai Le-Gia and Jaehyun Ahn. Training-free zero-shot anomaly detection in 3d brain mri with 2d foundation models.arXiv preprint arXiv:2602.15315, 2026
Pith/arXiv arXiv 2026
-
[47]
Giordano de Pinho Souza, Glaucia Melo, Josefino Cabral Melo Lima, and Daniel Schneider. Beyond english benchmarks: clinical llm evaluation in brazilian portuguese.arXiv preprint arXiv:2606.07853, 2026
Pith/arXiv arXiv 2026
-
[48]
Measuring the gap: correlating synthetic-to-real drift with phi de-identification performance.Genomics & Informatics, 24(1):10, 2026
Joseph Cornelius and Fabio Rinaldi. Measuring the gap: correlating synthetic-to-real drift with phi de-identification performance.Genomics & Informatics, 24(1):10, 2026
2026
-
[49]
Tai Le-Gia. On the problem of consistent anomalies in zero-shot anomaly detection.arXiv preprint arXiv:2512.02520, 2025
arXiv 2025
-
[50]
Reuniting forcibly separated families through shared memories with machine learning.Available at SSRN, 2025
Huifeng Su, Lesley Meng, and Edieal J Pinker. Reuniting forcibly separated families through shared memories with machine learning.Available at SSRN, 2025
2025
-
[51]
Semantic insurance pricing with large language models.arXiv preprint arXiv:2606.29371, 2026
Christopher Blier-Wong and Derek Kusmenko. Semantic insurance pricing with large language models.arXiv preprint arXiv:2606.29371, 2026
Pith/arXiv arXiv 2026
-
[52]
Ethical framework for responsible foundational models in medical imaging.Frontiers in Medicine, 12:1544501, 2025
Debesh Jha, Gorkem Durak, Abhijit Das, Jasmer Sanjotra, Onkar Susladkar, Suramyaa Sarkar, Ashish Rauniyar, Nikhil Kumar Tomar, Linkai Peng, Sirui Li, et al. Ethical framework for responsible foundational models in medical imaging.Frontiers in Medicine, 12:1544501, 2025
2025
-
[53]
Zilin Jing, Vincent Jeanselme, Yuta Kobayashi, Simon A Lee, Chao Pang, Aparajita Kashyap, Yanwei Li, Xinzhuo Jiang, and Shalmali Joshi. One loss to rule them all: Marked time-to-event for structured ehr foundation models.arXiv preprint arXiv:2602.00541, 2026
Pith/arXiv arXiv 2026
-
[54]
Ehr foundation models improve robustness in the presence of temporal distribution shift.Scientific Reports, 13(1):3767, 2023
Lin Lawrence Guo, Ethan Steinberg, Scott Lanyon Fleming, Jose Posada, Joshua Lemmon, Stephen R Pfohl, Nigam Shah, Jason Fries, and Lillian Sung. Ehr foundation models improve robustness in the presence of temporal distribution shift.Scientific Reports, 13(1):3767, 2023
2023
-
[55]
Ehr-r1: A reasoning-enhanced foundational language model for electronic health record analysis, 2025
Yusheng Liao, Chaoyi Wu, Junwei Liu, Shuyang Jiang, Pengcheng Qiu, Haowen Wang, Yun Yue, Shuai Zhen, Jian Wang, Qianrui Fan, Jinjie Gu, Ya Zhang, Yanfeng Wang, Yu Wang, and Weidi Xie. Ehr-r1: A reasoning-enhanced foundational language model for electronic health record analysis, 2025. URLhttps://arxiv.org/abs/2510.25628
arXiv 2025
-
[56]
Ehragent: Code empowers large language models for few-shot complex tabular reasoning on electronic health records
Wenqi Shi, Ran Xu, Yuchen Zhuang, Yue Yu, Jieyu Zhang, Hang Wu, Yuanda Zhu, Joyce C Ho, Carl Yang, and May Dongmei Wang. Ehragent: Code empowers large language models for few-shot complex tabular reasoning on electronic health records. InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 22315–22339, 2024
2024
-
[57]
Learning the natural history of human disease with generative transformers.Nature, 647(8088):248–256, 2025
Artem Shmatko, Alexander Wolfgang Jung, Kumar Gaurav, Søren Brunak, Laust Hvas Mortensen, Ewan Birney, Tom Fitzgerald, and Moritz Gerstung. Learning the natural history of human disease with generative transformers.Nature, 647(8088):248–256, 2025. 14
2025
-
[58]
Michael Wornow, Suhana Bedi, Miguel Angel Fuentes Hernandez, Ethan Steinberg, Jason Alan Fries, Christopher Ré, Sanmi Koyejo, and Nigam H Shah. Context clues: Evaluating long context models for clinical prediction tasks on ehrs.arXiv preprint arXiv:2412.16178, 2024
Pith/arXiv arXiv 2024
-
[59]
Core-behrt: A carefully optimized and rigorously evaluated behrt
Mikkel Odgaard, Kiril Vadimovic Klein, Sanne Møller Thysen, Espen Jimenez-Solem, Martin Sillesen, and Mads Nielsen. Core-behrt: A carefully optimized and rigorously evaluated behrt. arXiv preprint arXiv:2404.15201, 2024
Pith/arXiv arXiv 2024
-
[60]
Med-bert: pretrained contextu- alized embeddings on large-scale structured electronic health records for disease prediction
Laila Rasmy, Yang Xiang, Ziqian Xie, Cui Tao, and Degui Zhi. Med-bert: pretrained contextu- alized embeddings on large-scale structured electronic health records for disease prediction. NPJ digital medicine, 4(1):86, 2021
2021
-
[61]
Language models are an effective representation learning technique for electronic health record data.Journal of biomedical informatics, 113:103637, 2021
Ethan Steinberg, Ken Jung, Jason A Fries, Conor K Corbin, Stephen R Pfohl, and Nigam H Shah. Language models are an effective representation learning technique for electronic health record data.Journal of biomedical informatics, 113:103637, 2021
2021
-
[62]
Ehrshot: An ehr benchmark for few-shot evaluation of foundation models.Advances in Neural Information Processing Systems, 36, 2024
Michael Wornow, Rahul Thapa, Ethan Steinberg, Jason Fries, and Nigam Shah. Ehrshot: An ehr benchmark for few-shot evaluation of foundation models.Advances in Neural Information Processing Systems, 36, 2024
2024
-
[63]
Context clues: Evaluating long context models for clinical prediction tasks on ehr data
Michael Wornow, Suhana Bedi, Miguel Angel Fuentes Hernandez, Ethan Steinberg, Jason Alan Fries, Christopher Re, Sanmi Koyejo, and Nigam Shah. Context clues: Evaluating long context models for clinical prediction tasks on ehr data. InThe Thirteenth International Conference on Learning Representations, 2025
2025
-
[64]
Evidence extraction for automated medical coding: preliminary evaluation
Xiaorui Jiang, Kulsoom Khan, Sumithra Thinakara Vasantha, and Sajjad Haider. Evidence extraction for automated medical coding: preliminary evaluation. InProceedings of the 2024 8th International Conference on Natural Language Processing and Information Retrieval, pages 18–23, 2024
2024
-
[65]
Utilizing large language models to predict icd-10 diagnosis codes from patient medical records
Rudransh Pathak, Gabriel Vald, Yusuf Sermet, and Ibrahim Demir. Utilizing large language models to predict icd-10 diagnosis codes from patient medical records. In2024 IEEE MIT Undergraduate Research Technology Conference (URTC), pages 1–5. IEEE, 2024
2024
-
[66]
Can large language models abstract medical coded language?arXiv preprint arXiv:2403.10822, 2024
Simon A Lee and Timothy Lindsey. Can large language models abstract medical coded language?arXiv preprint arXiv:2403.10822, 2024
Pith/arXiv arXiv 2024
-
[67]
Simon A Lee, Anthony Wu, and Jeffrey N Chiang. Clinical modernbert: An efficient and long context encoder for biomedical text.arXiv preprint arXiv:2504.03964, 2025
Pith/arXiv arXiv 2025
-
[68]
Clinical text embeddings: A systematic review of methods, applications, and future directions.International Journal of Medical Informatics, page 106505, 2026
Hyunwoo Jo, Seon Kim, Hyunwoo Son, and Jongchan Kim. Clinical text embeddings: A systematic review of methods, applications, and future directions.International Journal of Medical Informatics, page 106505, 2026
2026
-
[69]
Aicoder: Exploring automated icd coding on chinese emrs with a multi-agent framework
Zhenpeng Liang, Hongjiao Guan, Weiyu Zhang, Ying Lian, Bing Xu, and Wenpeng Lu. Aicoder: Exploring automated icd coding on chinese emrs with a multi-agent framework. In2025 IEEE International Conference on Bioinformatics and Biomedicine (BIBM), pages 3823–3827. IEEE, 2025
2025
-
[70]
Gate: Graph and text exchange for zero-shot ecg classification with llm prompts.IEEE Journal of Biomedical and Health Informatics, 2026
Ying An, Shiyu Tang, Xianlai Chen, and Lin Guo. Gate: Graph and text exchange for zero-shot ecg classification with llm prompts.IEEE Journal of Biomedical and Health Informatics, 2026
2026
-
[71]
Yakun Zhu, Zhongzhen Huang, Qianhan Feng, Linjie Mu, Yannian Gu, Shaoting Zhang, Qi Dou, and Xiaofan Zhang. Cp-env: Evaluating large language models on clinical pathways in a controllable hospital environment.arXiv preprint arXiv:2512.10206, 2025
arXiv 2025
-
[72]
An analytical review of optimization techniques in information retrieval for enhanced decision support.Decision Analytics Journal, page 100657, 2025
Kemal Lazovi´c, Filipe Madeira, Eftim Zdravevski, Luis Augusto Silva, Paulo Jorge Coelho, and Ivan Miguel Pires. An analytical review of optimization techniques in information retrieval for enhanced decision support.Decision Analytics Journal, page 100657, 2025
2025
-
[73]
Xiaoquan Zhi, Hongke Zhao, Likang Wu, Chuang Zhao, and Hengshu Zhu. Reinventing clinical dialogue: Agentic paradigms for llm enabled healthcare communication.arXiv preprint arXiv:2512.01453, 2025. 15
arXiv 2025
-
[74]
Oncopt: long-context transformer models for in hospital tumor phenotype extraction from pathology reports.npj Digital Medicine, 2026
Thanh Duong, Dung Le, V onetta Williams, Sandra Stewart, Yayi Zhao, Muntasir Zitu, Issam El Naqa, Dana Rollison, and Thanh Thieu. Oncopt: long-context transformer models for in hospital tumor phenotype extraction from pathology reports.npj Digital Medicine, 2026
2026
-
[75]
Emerging properties in self-supervised vision transformers
Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerging properties in self-supervised vision transformers. InProceedings of the International Conference on Computer Vision (ICCV), 2021
2021
-
[76]
Dinov2: Learning robust visual features without supervision.arXiv preprint arXiv:2304.07193, 2023
Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy V o, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, et al. Dinov2: Learning robust visual features without supervision.arXiv preprint arXiv:2304.07193, 2023
Pith/arXiv arXiv 2023
-
[77]
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. InInternational conference on machine learning, pages 8748–8763. PMLR, 2021
2021
-
[78]
High-resolution image synthesis with latent diffusion models, 2021
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models, 2021
2021
-
[79]
Photorealistic text-to-image diffusion models with deep language understanding.Advances in neural information processing systems, 35:36479–36494, 2022
Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L Denton, Kamyar Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, et al. Photorealistic text-to-image diffusion models with deep language understanding.Advances in neural information processing systems, 35:36479–36494, 2022
2022
-
[80]
Contrastive learning of preferences with a contextual infonce loss, 2024
Timo Bertram, Johannes Fürnkranz, and Martin Müller. Contrastive learning of preferences with a contextual infonce loss, 2024. URLhttps://arxiv.org/abs/2407.05898
Pith/arXiv arXiv 2024
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.