REVIEW 9 cited by
Reading Your Heart: Learning ECG Words and Sentences via Pre-training ECG Language Model
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Electrocardiogram (ECG) is essential for the clinical diagnosis of arrhythmias and other heart diseases, but deep learning methods based on ECG often face limitations due to the need for high-quality annotations. Although previous ECG self-supervised learning (eSSL) methods have made significant progress in representation learning from unannotated ECG data, they typically treat ECG signals as ordinary time-series data, segmenting the signals using fixed-size and fixed-step time windows, which often ignore the form and rhythm characteristics and latent semantic relationships in ECG signals. In this work, we introduce a novel perspective on ECG signals, treating heartbeats as words and rhythms as sentences. Based on this perspective, we first designed the QRS-Tokenizer, which generates semantically meaningful ECG sentences from the raw ECG signals. Building on these, we then propose HeartLang, a novel self-supervised learning framework for ECG language processing, learning general representations at form and rhythm levels. Additionally, we construct the largest heartbeat-based ECG vocabulary to date, which will further advance the development of ECG language processing. We evaluated HeartLang across six public ECG datasets, where it demonstrated robust competitiveness against other eSSL methods. Our data and code are publicly available at https://github.com/PKUDigitalHealth/HeartLang.
Forward citations
Cited by 9 Pith papers
-
LAEF: A Lead-Agnostic ECG Foundation Model Towards Point-of-Care Diagnostics
A 7M-parameter graph ECG foundation model pre-trained with random lead dropout matches 12-lead models on full input and beats zero-padded baselines on most datasets with 1-2 leads.
-
From Muscle Bursts to Motor Intent: Self-Supervised Token Modeling for Heterogeneous EMG
AEMG pre-trains EMG representations by treating neuromuscular signals as language via a novel tokenizer and cross-device vocabulary, yielding 5.79-9.25% zero-shot LOSO gains and over 90% few-shot performance with 5% t...
-
Learning Cardiac Latent Representations in Vectorcardiogram Space
LVCG is the first self-supervised framework for learning view-invariant latent VCG representations that claims to outperform ECG-space baselines with better robustness and generalization in domain shift settings.
-
How Do Electrocardiogram Models Scale?
Empirical scaling study of ECG models finds SSL scales robustly while ResNets show 1.3-2.5x better parameter efficiency and SSL up to 16x better data efficiency than supervised baselines on out-of-distribution tasks.
-
From Muscle Bursts to Motor Intent: Self-Supervised Token Modeling for Heterogeneous EMG
AEMG learns robust neuromuscular representations from heterogeneous EMG datasets by modeling muscle contractions as tokens and using self-supervised pre-training to improve cross-user gesture recognition while reducin...
-
HeartcareGPT: A Unified Multimodal ECG Suite for Dual Signal-Image Modeling and Understanding
HeartcareGPT proposes Dual Stream Projection Alignment (DSPA) on a structure-aware tokenizer for unified ECG signal-image modeling, supported by Heartcare-400K dataset and Heartcare-Bench.
-
EchoBridge: Long-Tail-Aware ECG-Echocardiography Text Alignment for Echocardiography-Derived Cardiac Findings
EchoBridge’s shared–private ECG–echo-text alignment plus frequency-adaptive prototypes beats strong baselines on classifier-free and cross-center frozen probing, including several low-prevalence valvular findings.
-
From Muscle Bursts to Motor Intent: Self-Supervised Token Modeling for Heterogeneous EMG
AEMG learns reusable neuromuscular representations from eight standardized EMG datasets by modeling contraction events as tokens and using self-supervised pre-training to improve robustness to new users and reduce cal...
-
UniECG: Understanding and Generating ECG in One Unified Model
UniECG combines ECG interpretation and text-to-ECG generation in one model by fine-tuning a language model and aligning its output tokens with a pretrained ECG diffusion generator.
Discussion (0). Sign in to comment.