Pith. sign in

REVIEW 7 cited by

Zero-Shot ECG Classification with Multimodal Learning and Test-time Clinical Knowledge Enhancement

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.06659 v3 pith:IXHYLXFK submitted 2024-03-11 eess.SP cs.AIcs.LG

classification eess.SPcs.AIcs.LG
keywords merlclinicallearningclassificationdataesslknowledgezero-shot
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Electrocardiograms (ECGs) are non-invasive diagnostic tools crucial for detecting cardiac arrhythmic diseases in clinical practice. While ECG Self-supervised Learning (eSSL) methods show promise in representation learning from unannotated ECG data, they often overlook the clinical knowledge that can be found in reports. This oversight and the requirement for annotated samples for downstream tasks limit eSSL's versatility. In this work, we address these issues with the Multimodal ECG Representation Learning (MERL}) framework. Through multimodal learning on ECG records and associated reports, MERL is capable of performing zero-shot ECG classification with text prompts, eliminating the need for training data in downstream tasks. At test time, we propose the Clinical Knowledge Enhanced Prompt Engineering (CKEPE) approach, which uses Large Language Models (LLMs) to exploit external expert-verified clinical knowledge databases, generating more descriptive prompts and reducing hallucinations in LLM-generated content to boost zero-shot classification. Based on MERL, we perform the first benchmark across six public ECG datasets, showing the superior performance of MERL compared against eSSL methods. Notably, MERL achieves an average AUC score of 75.2% in zero-shot classification (without training data), 3.2% higher than linear probed eSSL methods with 10\% annotated training data, averaged across all six datasets. Code and models are available at https://github.com/cheliu-computation/MERL

Discussion (0). Sign in to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. LAEF: A Lead-Agnostic ECG Foundation Model Towards Point-of-Care Diagnostics

    cs.LG 2026-08 conditional novelty 7.0 of 10

    A 7M-parameter graph ECG foundation model pre-trained with random lead dropout matches 12-lead models on full input and beats zero-padded baselines on most datasets with 1-2 leads.

  2. Do ECG Foundation Models Transfer to Rare Cardiac Diseases? Evidence from Brugada Syndrome Detection

    cs.LG 2026-07 conditional novelty 6.5 of 10

    For Brugada syndrome detection, ECG foundation-model pre-training mainly stabilizes optimization rather than encoding transferable clinical knowledge, and fails to improve zero-shot cross-site generalization.

  3. Learning Cardiac Latent Representations in Vectorcardiogram Space

    cs.LG 2026-05 unverdicted novelty 6.0 of 10

    LVCG is the first self-supervised framework for learning view-invariant latent VCG representations that claims to outperform ECG-space baselines with better robustness and generalization in domain shift settings.

  4. How Do Electrocardiogram Models Scale?

    cs.LG 2026-05 conditional novelty 6.0 of 10

    Empirical scaling study of ECG models finds SSL scales robustly while ResNets show 1.3-2.5x better parameter efficiency and SSL up to 16x better data efficiency than supervised baselines on out-of-distribution tasks.

  5. EchoBridge: Long-Tail-Aware ECG-Echocardiography Text Alignment for Echocardiography-Derived Cardiac Findings

    cs.LG 2026-07 conditional novelty 5.5 of 10

    EchoBridge’s shared–private ECG–echo-text alignment plus frequency-adaptive prototypes beats strong baselines on classifier-free and cross-center frozen probing, including several low-prevalence valvular findings.

  6. SPOTR: Spatio-temporal Pooling One-Token Reconstruction for Universal Physiological Signal Self-supervised Learning

    cs.LG 2026-06 unverdicted novelty 5.0 of 10

    SPOTR is a single-token reconstruction self-supervised pretraining method for EEG, iEEG, ECG, and PPG that reports AUC gains of 4.64-21.71% over baselines under linear probing while cutting latency and memory.

  7. Physiology-Aware CNN and Zero-Shot Multimodal LLMs for ECG Image Classification: A Comparative Study

    cs.LG 2026-06 unverdicted novelty 4.0 of 10

    Zero-shot LLMs achieve near-chance ROC-AUC (~0.5) on ECG image classification while CNN models reach 0.92-0.94 internally and 0.85-0.86 externally on PTB-XL.

Pith tools