Pith. sign in

REVIEW 4 cited by

Teach Multimodal LLMs to Comprehend Electrocardiographic Images

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.19008 v1 pith:5YR7DEGC submitted 2024-10-21 eess.IV cs.AIcs.CV

classification eess.IVcs.AIcs.CV
keywords imageinterpretationmllmspulsecardiacchallengesconditionscovering
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The electrocardiogram (ECG) is an essential non-invasive diagnostic tool for assessing cardiac conditions. Existing automatic interpretation methods suffer from limited generalizability, focusing on a narrow range of cardiac conditions, and typically depend on raw physiological signals, which may not be readily available in resource-limited settings where only printed or digital ECG images are accessible. Recent advancements in multimodal large language models (MLLMs) present promising opportunities for addressing these challenges. However, the application of MLLMs to ECG image interpretation remains challenging due to the lack of instruction tuning datasets and well-established ECG image benchmarks for quantitative evaluation. To address these challenges, we introduce ECGInstruct, a comprehensive ECG image instruction tuning dataset of over one million samples, covering a wide range of ECG-related tasks from diverse data sources. Using ECGInstruct, we develop PULSE, an MLLM tailored for ECG image comprehension. In addition, we curate ECGBench, a new evaluation benchmark covering four key ECG image interpretation tasks across nine different datasets. Our experiments show that PULSE sets a new state-of-the-art, outperforming general MLLMs with an average accuracy improvement of 15% to 30%. This work highlights the potential of PULSE to enhance ECG interpretation in clinical practice.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. From Token to Rhythm: A Multi-Scale Approach for ECG-Language Pretraining

    eess.SP 2025-06 conditional novelty 6.0 of 10

    MELP pretrains ECG and text encoders with token-, beat-, and rhythm-level cross-modal supervision and beats prior baselines on several ECG classification benchmarks.

  2. UniECG: Understanding and Generating ECG in One Unified Model

    cs.CL 2025-09 conditional novelty 5.0 of 10

    UniECG combines ECG interpretation and text-to-ECG generation in one model by fine-tuning a language model and aligning its output tokens with a pretrained ECG diffusion generator.

  3. Signal, Image, or Symbolic: Exploring the Best Input Representation for Electrocardiogram-Language Models Through a Unified Framework

    cs.AI 2025-05 conditional novelty 5.0 of 10

    A unified benchmark across six ECG datasets and five text-generation metrics finds tokenized symbolic ECG inputs outperform raw signal and image inputs for ECG-language models.

  4. Enhancing Explainable Cardiac Diagnosis with Guide-Grounded Multimodal LLMs

    cs.AI 2026-07 reject novelty 4.0 of 10

    Guide-grounded prompting is reported to improve BERTScore of ECG impressions from 0.818 to 0.953, but the supporting tables contain implausible duplicated baseline numbers.

Pith tools