REVIEW 1 major objections 1 minor 4 cited by
Physiology-Aware Masked Cross-Modal Reconstruction for Biosignal Representation Learning
T0 review · 1 major / 1 minor · reviewed 2026-05-09 · grok-4.3
Pith's one-line read Pretraining with masked cross-modal reconstruction between temporally ordered biosignals like ECG and PPG produces representations that outperform unimodal and multimodal baselines on 15 of 19 downstream tasks.
desk verdict xMAE adds a temporal-order constraint to cross-modal masked reconstruction for ECG-PPG pairs and reports gains on most downstream tasks, but the evidence tying those gains specifically to the physiology-aware delay is thin. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The masked cross-modal reconstruction objective that reconstructs one temporally delayed biosignal (such as PPG) from masked patches of an earlier signal (such as ECG) to embed directional physiological timing.
What would settle it
A control model pretrained with standard masked reconstruction or contrastive objectives but without any cross-modal ordering constraint achieves equal or higher accuracy on the same 19 downstream tasks.
Extended reading notes
Core claim
xMAE is a biosignal pretraining framework that leverages masked cross-modal reconstruction across temporally ordered biosignals as a training-time constraint to encourage physiologically meaningful timing structure in the learned representations. Pretraining with xMAE yields representations that outperform both unimodal and multimodal baselines on 15 of 19 downstream tasks, including cardiovascular outcome prediction, abnormal laboratory test detection, sleep staging, and demographic inference, while generalizing across devices, body locations, and acquisition settings. Further analysis indicates that the ECG-PPG timing structure is reflected in the learned PPG representations.
Load-bearing premise
That the directional timing relationship between signals can be effectively enforced as a reconstruction constraint during pretraining and will produce representations that measurably improve downstream task performance.
Editorial extensions
If this is right
- Representations transfer to 15 of 19 tasks spanning outcome prediction, anomaly detection, sleep staging, and demographics.
- Performance gains hold when models are tested on new devices, sensor sites, and acquisition protocols.
- Learned PPG embeddings encode measurable ECG-to-PPG timing offsets.
- The approach applies to any multimodal biosignals that observe successive stages of one underlying process.
Reading between the lines
- The same ordering-aware reconstruction could be applied to other causally linked signal pairs such as respiratory effort before oxygen saturation changes.
- Wearable systems might benefit from pretraining on paired ECG-PPG streams to improve real-time fusion without explicit alignment modules.
- Similar constraints may help in other domains where one modality precedes another, such as audio preceding video in speech events.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces xMAE, a self-supervised pretraining framework for biosignals that performs masked cross-modal reconstruction between temporally ordered signals (e.g., ECG preceding PPG due to vascular delay) to learn physiologically structured representations. It reports that this approach outperforms unimodal and multimodal baselines on 15 of 19 downstream tasks spanning cardiovascular outcome prediction, abnormal lab test detection, sleep staging, and demographic inference, with generalization across devices, body locations, and settings. Additional analysis indicates that the learned PPG representations reflect the ECG-PPG timing structure.
Significance. If the empirical results are robust and the directional timing mechanism is shown to be causal for the gains, this work would be significant for advancing multimodal self-supervised learning in biosignals by incorporating physiological priors rather than treating signals as interchangeable views. The release of code supports reproducibility. It could influence pretraining strategies for other temporally structured multimodal data in healthcare.
major comments (1)
- Experiments section: No ablation study isolates the effect of the directional temporal ordering (e.g., by randomizing PPG relative to ECG or using symmetric bidirectional reconstruction without delay modeling) while holding masking, architecture, and other factors fixed. This is load-bearing for the central claim, as the reported gains on 15 of 19 tasks (including cardiovascular, sleep, and lab tasks) could arise from generic cross-modal pretraining rather than the physiology-aware timing structure.
minor comments (1)
- Abstract: The claim of outperformance on 15 of 19 tasks is stated without reference to specific baseline definitions, number of runs, or statistical tests, which would strengthen the summary for readers.
Simulated Author's Rebuttal
We thank the referee for the constructive feedback on our manuscript. We address the major comment below and will incorporate revisions to strengthen the evidence for our central claim.
read point-by-point responses
-
Referee: Experiments section: No ablation study isolates the effect of the directional temporal ordering (e.g., by randomizing PPG relative to ECG or using symmetric bidirectional reconstruction without delay modeling) while holding masking, architecture, and other factors fixed. This is load-bearing for the central claim, as the reported gains on 15 of 19 tasks (including cardiovascular, sleep, and lab tasks) could arise from generic cross-modal pretraining rather than the physiology-aware timing structure.
Authors: We agree that the manuscript lacks a dedicated ablation that isolates the directional temporal ordering while holding masking, architecture, and other factors fixed. Our current analysis shows that the learned PPG representations reflect the ECG-PPG timing structure, but this does not fully rule out that gains could arise from generic cross-modal pretraining. We will add the requested ablation (including randomized relative timing and symmetric bidirectional reconstruction) in the revised version to directly test causality of the physiology-aware timing mechanism. revision: yes
Circularity Check
No circularity: empirical results on downstream tasks are independent of the pretraining objective definition
full rationale
The paper defines xMAE as a masked cross-modal reconstruction objective that incorporates an external physiological fact (temporal ordering between ECG and PPG due to vascular delay). It then reports measured performance gains on 15 of 19 downstream tasks. This chain does not reduce any claimed result to a fitted parameter renamed as prediction, a self-referential definition, or a load-bearing self-citation. The temporal constraint is imported from physiology rather than derived from the model, and the outperformance numbers are obtained via standard evaluation rather than forced by construction. The derivation is therefore self-contained against external benchmarks.
Assumptions & free parameters
assumptions (1)
- domain assumption Biosignals acquired from different locations on the body often provide temporally ordered views of the same underlying physiological process.
Cite this review
Pith. "Pith review of Physiology-Aware Masked Cross-Modal Reconstruction for Biosignal Representation Learning." pith.science (2026). https://pith.science/paper/2605.00973
@misc{pith2026260500973,
author = {Pith},
title = {Pith review of: Physiology-Aware Masked Cross-Modal Reconstruction for Biosignal Representation Learning},
year = {2026},
howpublished = {\url{https://pith.science/paper/2605.00973}},
note = {Machine review of arXiv:2605.00973}
}
read the original abstract
Biosignals acquired from different locations on the body often provide temporally ordered views of the same underlying physiological process. However, most existing self supervised learning methods treat these signals as interchangeable views, overlooking the directional temporal dynamics that link them. A canonical example is the relationship between electrocardiography (ECG), which captures the electrical activation initiating each heartbeat, and photoplethysmography (PPG), which records the resulting peripheral pulse delayed by vascular dynamics. To capture this structured relationship, we introduce xMAE, a biosignal pretraining framework that leverages masked cross modal reconstruction across temporally ordered biosignals as a training time constraint to encourage physiologically meaningful timing structure in the learned representations. We show that pretraining with xMAE yields representations that outperform both unimodal and multimodal baselines on 15 of 19 downstream tasks, including cardiovascular outcome prediction, abnormal laboratory test detection, sleep staging, and demographic inference, while generalizing across devices, body locations, and acquisition settings. Further analysis suggests that the ECG PPG timing structure is reflected in the learned PPG representations. More broadly, xMAE demonstrates the effectiveness of incorporating temporal structure into multimodal pretraining when signals observe different stages of a shared underlying process. Code is available at https://github.com/hzhou3/xMAE.
Forward citations
Cited by 4 Pith papers
-
GlucoFM: A Dual-Stream Foundation Model for Continuous Glucose Monitoring
GlucoFM decomposes CGM traces into dual state-event streams, pretrains on 109k hours of unlabeled data, and reports superior subject-disjoint performance on seven clinical tasks across four cohorts.
-
HIPNO: Symmetry-Aware Physics-Informed Neural Operators for Noninvasive Hemodynamic Inference
A physics-informed neural operator that quotients out the scale symmetry of the 3-element Windkessel model predicts a vascular decay time constant from noninvasive signals with 32% lower log-scale error than a populat...
-
AURORA: Contextual Orthogonalization for Geometric Representation Learning in Healthcare Foundation Models
AURORA is a representation learning framework that uses contextual orthogonalization and relational alignment to create disentangled, geometrically interpretable latent spaces in healthcare foundation models.
-
MorphologyFM: A Foundation Model for Morphology-Aware Representation Learning from ECG and Pulse Oximetry Waveforms
Morphology-aware masking plus cross-modal ECG–SpO2 pretraining on MIMIC yields stronger transfer than MAE, contrastive, Barlow Twins, and JEPA on several clinical prediction tasks.
Reference graph
Works this paper leans on
-
[1]
Large-scale Training of Foundation Models for Wearable Biosignals
S. Abbaspourazad, O. Elachqar, A. C. Miller, S. Emrani, U. Nallasamy, and I. Shapiro. Large-scale training of foundation models for wearable biosignals. arXiv preprint arXiv:2312.05409,
-
[2]
M. A. Ahmad, C. Eckert, and A. Teredesai. Interpretable machine learning in healthcare. In Proceedings of the 2018 ACM international conference on bioinformatics, computational biology, and health informatics, pages 559--560,
work page 2018
-
[3]
A. F. Ansari, L. Stella, C. Turkmen, X. Zhang, P. Mercado, H. Shen, O. Shchur, S. S. Rangapuram, S. P. Arango, S. Kapoor, et al. Chronos: Learning the language of time series. arXiv preprint arXiv:2403.07815,
- [4]
-
[5]
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
A. Dosovitskiy. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929,
work page Pith review arXiv 2010
-
[6]
Beyond Sensor Data: Foundation Models of Behavioral Data from Wearables Improve Health Predictions
E. Erturk, F. Kamran, S. Abbaspourazad, S. Jewell, H. Sharma, Y. Li, S. Williamson, N. J. Foti, and J. Futoma. Beyond sensor data: Foundation models of behavioral data from wearables improve health predictions. arXiv preprint arXiv:2507.00191,
- [7]
- [8]
Show all 29 references
-
[9]
N. C. Kong, D. Lee, H. Do, D. H. Park, C. Xu, H. Mao, and J. Chung. f-gan: A frequency- domain-constrained generative adversarial network for ppg to ecg synthesis. arXiv preprint arXiv:2406.16896,
-
[10]
S. A. Lee, C. Tanade, H. Zhou, J. Lee, M. Thukral, M. Han, R. Choi, M. S. H. Khan, B. Lu, M. Gwak, et al. Himae: Hierarchical masked autoencoders discover resolution-specific structure in wearable time series. arXiv preprint arXiv:2510.25785,
-
[11]
J. Li, A. Aguirre, J. Moura, C. Liu, L. Zhong, C. Sun, G. Clifford, B. Westover, and S. Hong. An electrocardiogram foundation model built on over 10 million recordings with external evaluation across multiple domains. arXiv preprint arXiv:2410.04133,
-
[12]
Loshchilov and F
I. Loshchilov and F. Hutter. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101,
-
[13]
Mukkamala, J.-O
R. Mukkamala, J.-O. Hahn, O. T. Inan, L. K. Mestha, C.-S. Kim, H. Töreyin, and S. Kyal. Toward ubiq- uitous blood pressure monitoring via pulse transit time: theory and practice. IEEE transactions on biomedical engineering, 62(8):1879--1901,
1901
-
[14]
Narayanswamy, X
G. Narayanswamy, X. Liu, K. Ayush, Y. Yang, X. Xu, S. Liao, J. Garrison, S. Tailor, J. Sunshine, Y. Liu, et al. Scaling wearable foundation models. arXiv preprint arXiv:2410.13638,
-
[15]
G. Nie, G. Tang, Y. Xiao, J. Li, S. Huang, D. Zhang, Q. Zhao, and S. Hong. Anyppg: An ecg-guided ppg foundation model trained on over 100,000 hours of recordings for holistic health profiling. arXiv preprint arXiv:2511.01747,
-
[16]
Pillai, D
A. Pillai, D. Spathis, F. Kawsar, and M. Malekzadeh. Papagei: Open foundation models for optical physiological signals. arXiv preprint arXiv:2410.20542,
-
[17]
Thapa, B
R. Thapa, B. He, M. R. Kjaer, H. Moore, G. Ganjoo, E. Mignot, and J. Zou. Sleepfm: Multi-modal representation learning for sleep across brain activity, ecg and respiratory signals. arXiv preprint arXiv:2405.17766,
-
[18]
K. Wang, J. Yang, A. Shetty, and J. Dunn. Dreamt: Dataset for real-time sleep stage estimation using multisensor wearable technology. PhysioNet https://doi.org/10.13026/62AN-CB28,
-
[19]
W. Whelton. 2017 guideline for the prevention, detection, evaluation, and management of high blood pressure in adults. J Am Coll Cardiol,
2017
-
[20]
M. A. Xu, G. Narayanswamy, K. Ayush, D. Spathis, S. Liao, S. A. Tailor, A. Metwally, A. A. Heydari, Y. Zhang, J. Garrison, et al. Lsm-2: Learning from incomplete wearable sensor data. arXiv preprint arXiv:2506.05321,
-
[21]
H. Zhou, M. M. Rahman, M. B. Morshed, Y. Li, M. S. Islam, L. Zhang, J. Bae, C. Rosa, W. B. Mendes, and J. Kuang. Know your heart better: Multimodal cardiac output monitoring using earbuds. In ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Proces...
2025
-
[22]
Signal Preprocessing Pipeline To facilitate pretraining and evaluation, we follow a standard preprocessing pipeline that ensures high-quality PPG and ECG segments
18 Physiology-Aware Masked Cross-Modal Reconstruction for Biosignal Representation Learning A. Signal Preprocessing Pipeline To facilitate pretraining and evaluation, we follow a standard preprocessing pipeline that ensures high-quality PPG and ECG segments. This preprocessing...
2003
-
[23]
Input We consider paired photoplethysmography (PPG) and electrocardiography (ECG) signals collected synchronously from the same subject. Each input sample consists of a 10-second segment sampled at 100 Hz, yielding sequences 𝑃∈R 𝐿, 𝐸∈R 𝐿, 𝐿= 1000.(5) Curriculum ECG Masking Str...
2020
-
[24]
Learnable positional embeddings are added to encode temporal order
This yields 𝑍∈R 𝑁 ′×𝑑, 𝑁 ′ = ⌊︂𝐿′ 𝑃 ⌋︂ .(8) For fully observed PPG, this results in 𝑁= 25 tokens per segment (length is 40; 40 × 25 = 1000). Learnable positional embeddings are added to encode temporal order. PPG and visible ECG tokens are then processed independently by modal...
2019
-
[25]
PulsePPG (Open-Source Weights) Saha et al
We use the pretrained PPG encoder as provided, and evaluate its representations on our downstream tasks without additional pretraining or task-specific adaptation. PulsePPG (Open-Source Weights) Saha et al. (2025) For this baseline, we adopt the official PulsePPG implementatio...
2025
-
[26]
All training and evaluation are performed on NVIDIA H200 GPUs. C. Evaluation Datasets, Tasks and Protocols In this section, we introduce datasets, tasks, and protocols that are employed for evaluation. C.1. Evaluation Datasets and Tasks In total, we have 19 tasks from 6 datase...
2025
-
[27]
Random Seed We set the random seed to 1 across all tasks and evaluations
We kept the hyperparameters, such as learning rate (1e-5), batch size (2048) same across models. Random Seed We set the random seed to 1 across all tasks and evaluations. D. Justification of Curriculum ECG Masking We provide a justification of our choice on curriculum ECG mask...
-
[28]
E.4. Evidence 2: xMAE Captures the Time Delay Better than Multimodal Baselines Figure 14 evaluates how well different models preserve the physiological time delay between ECG and PPG by comparing the absolute error between the ground-truth delay, Δ𝑡𝑔𝑡, com- puted from real ECG...
2021
-
[29]
Again, these models are trained with different architectures, different sizes, and different pretraining datasets. Yet, xMAE consistently achieves comparable performance on clinically and physiologi- cally grounded tasks, particularly cardiovascular outcomes and laboratory tes...
2024
Reviewed May 9, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.