REVIEW 4 cited by
Teach Multimodal LLMs to Comprehend Electrocardiographic Images
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The electrocardiogram (ECG) is an essential non-invasive diagnostic tool for assessing cardiac conditions. Existing automatic interpretation methods suffer from limited generalizability, focusing on a narrow range of cardiac conditions, and typically depend on raw physiological signals, which may not be readily available in resource-limited settings where only printed or digital ECG images are accessible. Recent advancements in multimodal large language models (MLLMs) present promising opportunities for addressing these challenges. However, the application of MLLMs to ECG image interpretation remains challenging due to the lack of instruction tuning datasets and well-established ECG image benchmarks for quantitative evaluation. To address these challenges, we introduce ECGInstruct, a comprehensive ECG image instruction tuning dataset of over one million samples, covering a wide range of ECG-related tasks from diverse data sources. Using ECGInstruct, we develop PULSE, an MLLM tailored for ECG image comprehension. In addition, we curate ECGBench, a new evaluation benchmark covering four key ECG image interpretation tasks across nine different datasets. Our experiments show that PULSE sets a new state-of-the-art, outperforming general MLLMs with an average accuracy improvement of 15% to 30%. This work highlights the potential of PULSE to enhance ECG interpretation in clinical practice.
Forward citations
Cited by 4 Pith papers
-
From Token to Rhythm: A Multi-Scale Approach for ECG-Language Pretraining
MELP pretrains ECG and text encoders with token-, beat-, and rhythm-level cross-modal supervision and beats prior baselines on several ECG classification benchmarks.
-
UniECG: Understanding and Generating ECG in One Unified Model
UniECG combines ECG interpretation and text-to-ECG generation in one model by fine-tuning a language model and aligning its output tokens with a pretrained ECG diffusion generator.
-
Signal, Image, or Symbolic: Exploring the Best Input Representation for Electrocardiogram-Language Models Through a Unified Framework
A unified benchmark across six ECG datasets and five text-generation metrics finds tokenized symbolic ECG inputs outperform raw signal and image inputs for ECG-language models.
-
Enhancing Explainable Cardiac Diagnosis with Guide-Grounded Multimodal LLMs
Guide-grounded prompting is reported to improve BERTScore of ECG impressions from 0.818 to 0.953, but the supporting tables contain implausible duplicated baseline numbers.
Discussion (0). Sign in to comment.