REVIEW 3 major objections 8 minor 44 references
Aligning ECGs to standardized echo conclusions with shared–private projections and frequency-aware prototypes yields stronger, transferable ECG representations of structural heart findings—including rare valvular ones.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-31 11:48 UTC pith:EAWA7IBP
load-bearing objection Solid multi-protocol ECG–echo-text method paper with real external cohorts; the headline “classifier-free” gains are partly supervised prototype access, and the label-only control undercuts the low-budget probing story. the 3 major comments →
EchoBridge: Long-Tail-Aware ECG-Echocardiography Text Alignment for Echocardiography-Derived Cardiac Findings
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Paired ECGs and standardized echocardiography conclusions can teach transferable ECG representations of predefined cardiac findings if cross-modal alignment is restricted to complementary shared projections and the shared hypersphere is calibrated with frequency-adaptive prototype margins plus spherical repulsion; the resulting space improves prompt-based and linear access to findings—including low-prevalence valvular ones—across centers.
What carries the argument
Complementary Shared–Private Projection (CSPP) plus Adaptive Prototype Boundary Calibration (APBC): CSPP maps ECG and text into shared and private heads, enforces within-modality orthogonality, and bidirectionally aligns normalized shared vectors; APBC places learnable class prototypes on the sphere with training-frequency-adaptive positive angular margins and Riesz repulsion so tail classes stay compact and prototypes stay separated.
Load-bearing premise
The method assumes that short, template-style conclusions built straight from the same structured finding labels are rich enough cross-modal supervision—not just multi-label training routed through a text encoder—and still work when real free-text reports, templates, or label extraction differ.
What would settle it
Replace the deterministic conclusion templates with real physician free-text echo reports (or a different verbalization template) while holding ECG data, labels, and protocols fixed: if classifier-free and transfer gains over strong ECG–text baselines largely disappear, the claimed cross-modal advantage is mostly label routing rather than echo-text correspondence.
If this is right
- Frozen ECG encoders pretrained this way can support prompt-based screening for pretraining-seen echo findings without training a new classifier.
- Limited target-hospital labels become more useful: linear probes at 1–10% budgets remain competitive under center and prevalence shift.
- Source-only deployment (fixed model and thresholds) can retain ranking and positive-class discrimination on external cohorts for shared findings.
- Tail valvular findings that usually starve contrastive learning get stronger geometric supervision via larger adaptive margins and prototype separation.
- Separating shared vs private ECG factors offers a practical path when imaging storage or governance blocks raw echo–ECG video training.
Where Pith is reading between the lines
- If private branches absorb acquisition and rhythm confounds as intended, similar shared–private splits may help other sparse clinical pairing settings (ECG–report, PPG–ECG) where one modality is compositional text and the other is continuous waveforms.
- Frequency-adaptive angular margins are a general recipe for multi-label medical hyperspheres whenever positive rates span orders of magnitude; the same schedule could be stress-tested outside cardiology.
- Because prompts are definition-only and fixed before test, the setup invites a controlled study of how much lexical overlap with training templates inflates zero-shot scores versus true semantic accessibility.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript proposes EchoBridge, a framework for aligning 12-lead ECGs with standardized echocardiography conclusion-style text to learn ECG representations of echo-derived cardiac findings. It introduces two components: Complementary Shared–Private Projection (CSPP), which maps each modality into shared and private projections with within-modality orthogonality and symmetric contrastive alignment of the shared space; and Adaptive Prototype Boundary Calibration (APBC), which organizes the shared hypersphere with class-specific prototypes, training-frequency-adaptive angular margins (applied to positive pairs only), and spherical Riesz repulsion between prototypes. Pretraining text is constructed by deterministically verbalizing the structured multi-label finding vector. Evaluation covers EchoNext-Mini plus two independent hospital cohorts (PKUPH, SHTMU) under four protocols — prompt-based classifier-free inference, in-domain frozen linear probing at 1%/10%/100% labels, target-domain cross-center probing, and source-only transfer — with bootstrap CIs, matched baselines, a component-wise ablation including a label-only BCE control, and finding-level analyses. EchoBridge reports the best point estimates across essentially all protocols and budgets, with headline prompt-based gains of 7.88/5.61/4.54 points in AUROC/AUPRC/F1 over the strongest baseline.
Significance. If the results hold, the work is a solid, practically relevant contribution to AI-ECG screening for structural heart disease: label-efficient frozen representations that transfer across institutions, with a usable prompt-based inference mode for pretraining-seen findings. Particular strengths that weigh in its favor: (i) four complementary evaluation protocols including a genuinely difficult source-only transfer setting with frozen thresholds; (ii) bootstrap CIs on nearly all main tables; (iii) a matched label-only BCE control in the ablation — an unusually honest inclusion, even though its numbers partly cut against the paper (see major comment 2); (iv) two independent external cohorts with finding-level breakdowns and explicit caution about small positive counts; (v) public code. The methodological novelty (shared–private projection with orthogonality; frequency-adaptive angular margins; spherical Riesz prototype repulsion) is incremental over the recent SGERA/D-BETA line of ECG–text alignment, but the components are cleanly ablated and each contributes. The significance ceiling is set by the supervision confound in the headline protocol and by the deterministic label-to-text pre
major comments (3)
- [Table 1 / §3.4, §E.1, Eqs. 10–15] The headline prompt-based gains (+7.88 AUROC / +5.61 AUPRC / +4.54 F1 over SGERA) compare EchoBridge against baselines that receive no class-level supervision on the seven evaluated findings, while EchoBridge's L_proto (Eq. 15) trains the shared embedding with multi-label BCE against learnable prototypes for exactly those seven classes. The prompt-based protocol (§E.1) then substitutes definition-only text prompts for the learned prototypes and calls the result 'classifier-free.' The authors do disclose the scope ('this protocol measures the accessibility of pretraining-seen finding semantics'), which is creditable, and the Align.+Proto. row of Table 6 (67.18 AUROC) shows the proposed components add something beyond raw prototype supervision. But the abstract's framing still invites a CLIP-style zero-shot reading, and the Table 1 comparison confounds alignment quality with an additional
- [Table 6 / §4.5 (ablation and label-only control)] The paper emphasizes label efficiency at low probing budgets (§4.2: 'gains indicate that finding-related information remains linearly accessible under limited supervision'), but its own matched label-only BCE control reaches 73.79 AUROC at 1% probing — above EchoBridge's 72.77 — and 76.87 vs 76.94 at 10% (a tie). Only at 100% labels does EchoBridge clearly lead (78.79 vs 76.87). So at precisely the budgets where the label-efficiency claim is made, plain multi-label BCE on the same encoder explains the AUROC advantage; the text-alignment machinery contributes on AUPRC/F1 (26.53 vs 26.16; 31.32 vs 31.07 at 1%) but those gaps are small. Two specific requests: (i) Table 6 is the only results table without bootstrap CIs — please add them, since the 1%/10% comparisons are near-ties and their sign matters for the central claim; (ii) add an align-only ablation row (contrastive alignment + CSPP,
- [Tables 3–5 / §4.3–4.4 (cross-center claims vs. CI overlap)] The repeated claim of 'highest point estimates across all budgets and both transfer cohorts' is technically accurate, but several decisive cells have overlapping bootstrap CIs: e.g., Table 5 PKUPH F1 (EchoBridge 27.39 [26.57, 28.48] vs D-BETA 27.13 [26.31, 28.21]) and several 100%-budget cells in Tables 3–4 (e.g., Table 3 100% AUROC 79.48 vs MERL-ECHO 77.88/ECG-CLIP 78.02 — CIs overlap broadly). Since these are the same frozen representations evaluated on the same test sets, a paired bootstrap on per-sample score differences (or at least a statement of which cells are statistically separable) would substantially strengthen the transfer claims. As written, 'exceeds the strongest competing method for each metric by 0.26 F1 points' (§4.4, PKUPH) reads as a significant gain where the CIs suggest equivalence. This is fixable with analysis rather than new experiments.
minor comments (8)
- [§2.2, Eqs. 2–4] The private projections z^p receive gradient only through L_orth (Eq. 4); no alignment, reconstruction, or auxiliary objective uses them. They are essentially unconstrained capacity, and Table 13 shows they nonetheless retain predictive signal (75.67 AUROC under 100% probing) — presumably inherited through the shared encoder. Please state explicitly what the private branch is for mechanistically, and whether the +CSPP ablation isolates the orthogonality term or the mere addition of extra heads (the current ablation changes both at once).
- [§3.2 (text construction)] The pretraining text is a deterministic verbalization of the same structured label vector used for L_proto and for evaluation (§3.2). The manuscript should state plainly that the 'cross-modal' supervision carries no information beyond the label set — the text modality functions as a label-conditioned codebook — and discuss what is gained versus lost relative to free-text reports (e.g., MERL-ECHO). This matters for interpreting the prompt-based protocol, where definition-only prompts partially break the lexical identity between training text and test prompts.
- [§2.3 / Table 9] No sensitivity analysis is provided for the APBC hyperparameters m0=0.20 rad, [m_min, m_max]=[0.05, 0.50], q=2.0, λr=0.05 (Table 9). Given that FAAM/SRR each contribute ~1 AUROC point over CSPP alone (Table 6), even a small sweep (e.g., m0 ∈ {0.1, 0.2, 0.3}, λr ∈ {0.01, 0.05, 0.1}) would clarify robustness.
- [Typesetting of Eqs. 4, 10, 14] In Eq. (4) and Eq. (10) the inner-product notation ⟨·,·⟩ and the squared terms render as broken glyphs in the PDF (¨z-symbols with missing brackets). Please fix the typesetting; the equations are currently hard to parse without the algorithm box.
- [Tables 10–12 (finding-specific CIs)] Several low-prevalence finding-level estimates rest on very few positives (e.g., Table 11: TR P=26, RVE P=35; Table 12: RAE P=28). The authors appropriately flag this in the text, but the per-finding tables should also carry CIs, or at least a footnote quantifying the CI width at these positive counts, since the +8.15 AUPRC gain for PR (Table 10, P=154) is a headline finding-level result.
- [§4.1] The prompt-based absolute numbers are modest (F1 31.83, AUPRC 26.79) and far below 100% linear probing (35.82 / 32.66). A sentence calibrating what 'semantic accessibility' means operationally — and that prompt-based inference is not yet competitive with a probed classifier — would prevent over-reading of Table 1.
- [Front matter / ACM Reference Format] The ACM Reference Format block reads '2018' (copyright year) while the venue is KDD '27 — presumably a template artifact, but please correct before camera-ready. Also, the Generative AI Usage disclosure is appreciated; consider noting in §E.3 that prompt candidates were fixed before any test evaluation (this is stated, but worth cross-referencing from the disclosure).
- [§C.2 vs. main text] Please clarify whether the text encoder (MedCPT) is frozen or fine-tuned during pretraining — Appendix C.2 says 'all parameters are jointly fine-tuned,' but the main text never states this, and it affects interpretation of how much label information is baked into the text branch itself versus the shared projection.
Circularity Check
No derivation circularity: empirical multi-label ECG–text pretraining evaluated on held-out patients, with disclosed pretraining-seen scope.
full rationale
EchoBridge is an empirical representation-learning paper, not a closed-form derivation. CSPP (Eqs. 1–7) and APBC (Eqs. 8–17) are architectural/training objectives; reported gains are macro AUROC/AUPRC/F1 on patient-disjoint test sets under four protocols. Deterministic verbalization of structured finding vectors into conclusion-style text (Sec. 3.2) and class-prototype BCE on the same predefined findings (L_proto, Eqs. 10–15) make prompt-based inference a measure of accessibility of pretraining-seen semantics—which the paper explicitly scopes in §E.1—rather than a quantity forced equal to its inputs by construction. Held-out patients can (and baselines do) score lower; the label-only BCE control in Table 6 is an independent ablation, not a fitted parameter renamed as prediction. Self-citations in related work are contextual and not load-bearing uniqueness claims. No self-definitional identity, fitted-input-as-prediction, or uniqueness-import chain is present. Concerns about confounding prototype supervision with alignment quality or about contribution vs. pure multi-label BCE are experimental-design/contribution issues, not circularity.
Axiom & Free-Parameter Ledger
free parameters (5)
- base angular margin m0 =
0.20 rad
- margin clip bounds (m_min, m_max) =
0.05 rad, 0.50 rad
- Riesz exponent q and weight λr =
q=2.0, λr=0.05
- contrastive temperature τ and logit scale γ=exp(s) =
τ not numerically fixed in main hyperparameter table; s learnable
- optimization and architecture knobs =
as in Table 9 / Sec. 3.3
axioms (5)
- domain assumption Standardized conclusion-style text verbalized from structured multi-label findings preserves clinically salient echo semantics for ECG supervision.
- ad hoc to paper Within-modality orthogonality between shared and private projections reduces harmful directional redundancy without needing semantic labels on private factors.
- ad hoc to paper Training-set positive rate is a sufficient statistic for allocating angular margins under long-tail multi-label imbalance.
- domain assumption Patient-disjoint 7:1:2 splits and dataset-specific ECG–echo matching windows yield valid generalization estimates.
- standard math Cosine geometry on an ℓ2-normalized shared hypersphere is an appropriate common space for ECG–text alignment and prototype classification.
invented entities (3)
-
Complementary Shared–Private Projection (CSPP)
no independent evidence
-
Adaptive Prototype Boundary Calibration (APBC) with FAAM and spherical Riesz repulsion
no independent evidence
-
Cross-modal class prototype matrix P shared by ECG and text branches
no independent evidence
read the original abstract
Standardized echocardiography conclusions provide meaningful supervision for learning ECG representations of echocardiography-derived cardiac findings. Global ECG--text alignment may entangle modality-specific factors, while long-tailed finding distributions provide sparse positive supervision for low-prevalence conditions. We propose EchoBridge with Complementary Shared--Private Projection (CSPP) and Adaptive Prototype Boundary Calibration (APBC). CSPP maps each modality into shared and auxiliary private projections, reduces directional redundancy via within-modality orthogonality, and bidirectionally aligns normalized shared projections. APBC organizes the shared hypersphere with class-specific prototypes, training-frequency-adaptive angular margins, and spherical Riesz repulsion. We evaluate EchoBridge on EchoNext-Mini and independent PKUPH and SHTMU cohorts under four protocols: prompt-based inference without downstream classifier training, in-domain frozen linear probing, target-domain cross-center frozen linear probing, and source-only cross-center transfer, supplemented by finding-specific analyses. EchoBridge improves classifier-free AUROC, AUPRC, and F1 over the strongest baselines by 7.88, 5.61, and 4.54 points, respectively, and achieves the highest point estimates across all in-domain and target-domain probing budgets and both source-only transfer cohorts. Finding-specific analyses show gains for most conditions, including several low-prevalence valvular findings.
Figures
Reference graph
Works this paper leans on
-
[1]
Arya Aminorroaya, Lovedeep S Dhingra, Aline F Pedroso, Sumukh Vasisht Shankar, Andreas Coppi, Akshay Khunte, Murilo Foppa, Luisa CC Brant, Sandhi M Barreto, Antonio Luiz P Ribeiro, et al . 2025. Development and multinational validation of an ensemble deep learning algorithm for detecting and predicting structural heart disease using noisy single-lead elec...
2025
-
[2]
Zachi I Attia, Suraj Kapa, Francisco Lopez-Jimenez, Paul M McKie, Dorothy J Ladewig, Gaurav Satam, Patricia A Pellikka, Maurice Enriquez-Sarano, Peter A Noseworthy, Thomas M Munger, et al. 2019. Screening for cardiac contractile dysfunction using an artificial intelligence–enabled electrocardiogram.Nature medicine25, 1 (2019), 70–74
2019
-
[3]
Chieh-Ju Chao, Jean-Benoit Delbrouck, Mohammad Asadi, Imon Banerjee, Juan M Farina, Francesca Galasso, Ahmed K Mahmoud, Mohammed Tiseer Abbas, Yu- Chiang Wang, Reza Arsanjani, et al . 2025. EchoGraph system for automated quality assessment of echocardiography reports.NPJ Digital Medicine(2025)
2025
-
[4]
Jian Chen, Xiaoru Dong, Wei Wang, Shaorui Zhou, Lequan Yu, and Xiping Hu
-
[5]
Jian Chen, Yipeng Du, Wenhao Yuan, Shuai Wang, Jinfeng Xu, Zewei Liu, Running Zhao, and Edith C. H. Ngai. 2026. SGERA: Stein-Guided ECG-Report Alignment for ECG Representation Learning. InForty-third International Conference on Machine Learning
2026
-
[6]
International Joint Conferences on Artificial Intelligence, 4824–4832
-
[7]
Matthew Christensen, Milos Vukadinovic, Neal Yuan, and David Ouyang. 2024. Vision–language foundation model for echocardiogram interpretation.Nature Medicine30, 5 (2024), 1481–1488
2024
-
[8]
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. 2020. A simple framework for contrastive learning of visual representations. InInterna- tional conference on machine learning. PmLR, 1597–1607
2020
-
[9]
Lovedeep S Dhingra, Arya Aminorroaya, Veer Sangha, Aline F Pedroso, Sumukh Vasisht Shankar, Andreas Coppi, Murilo Foppa, Luisa CC Brant, Sandhi M Barreto, Antonio Luiz P Ribeiro, et al . 2025. Ensemble deep learn- ing algorithm for structural heart disease screening using electrocardiographic images: PRESENT SHD.Journal of the American College of Cardiolo...
2025
-
[10]
Sanghyuk Chun. 2023. Improved probabilistic image-text representations.arXiv preprint arXiv:2305.18171(2023)
Pith/arXiv arXiv 2023
-
[11]
Xiaocheng Fang, Zhengyao Ding, Jieyi Cai, Yujie Xiao, Bo Liu, Jiarui Jin, Haoyu Wang, Guangkun Nie, Shun Huang, Ting Chen, et al. 2026. ECGFlowCMR: Pre- training with ECG-Generated Cine CMR Improves Cardiac Disease Classification and Phenotype Prediction.arXiv preprint arXiv:2601.20904(2026)
Pith/arXiv arXiv 2026
-
[12]
Pierre Elias, Timothy J Poterucha, Vijay Rajaram, Luca Matos Moller, Victor Rodriguez, Shreyas Bhave, Rebecca T Hahn, Geoffrey Tison, Sean A Abreau, Joshua Barrios, et al . 2022. Deep learning electrocardiographic analysis for detection of left-sided valvular heart disease.Journal of the American College of Cardiology80, 6 (2022), 613–626
2022
-
[13]
Goro Fujiki, Satoshi Kodera, Naoto Setoguchi, Kengo Tanabe, Kotaro Miyaji, Shunichi Kushida, Mike Saji, Mamoru Nanasato, Hisataka Maki, Hideo Fujita, et al. 2025. Deep learning-based identification of echocardiographic abnormalities from electrocardiograms.JACC: Asia5, 1_Part_1 (2025), 88–98
2025
-
[14]
Xiaocheng Fang, Jiarui Jin, Haoyu Wang, Che Liu, Jieyi Cai, Yujie Xiao, Guangkun Nie, Bo Liu, Shun Huang, Hongyan Li, et al . 2025. PPGFlowECG: Latent Rec- tified Flow with Cross-Modal Encoding for PPG-Guided ECG Generation and Cardiovascular Disease Detection.arXiv preprint arXiv:2509.19774(2025)
arXiv 2025
-
[15]
John Weston Hughes, Linyuan Jing, Joshua Finer, Dustin Hartzel, Christopher Kelsey, Aaron Long, Daniel Rocha, Jeffrey Ruhl, Timothy Poterucha, and Pierre Elias. 2026. EchoNext-mini: A dataset and baseline AI model for detecting struc- tural heart disease from electrocardiograms.NEJM AI3, 5 (2026), AIdbp2500516
2026
-
[16]
Amirata Ghorbani, David Ouyang, Abubakar Abid, Bryan He, Jonathan H Chen, Robert A Harrington, David H Liang, Euan A Ashley, and James Y Zou. 2020. Deep learning interpretation of echocardiograms.NPJ digital medicine3, 1 (2020), 10
2020
-
[17]
Jiarui Jin, Haoyu Wang, Hongyan Li, Jun Li, Jiahui Pan, and Shenda Hong. 2025. Reading your heart: Learning ecg words and sentences via pre-training ecg language model.arXiv preprint arXiv:2502.10707(2025)
Pith/arXiv arXiv 2025
-
[18]
Manh Pham Hung, Aaqib Saeed, and Dong Ma. 2025. Boosting Masked ECG- Text Auto-Encoders as Discriminative Learners. InForty-second International Conference on Machine Learning
2025
-
[19]
Qiao Jin, Won Kim, Qingyu Chen, Donald C Comeau, Lana Yeganova, W John Wilbur, and Zhiyong Lu. 2023. Medcpt: Contrastive pre-trained transformers with large-scale pubmed search logs for zero-shot biomedical information retrieval. Bioinformatics39, 11 (2023), btad651
2023
-
[20]
Jiarui Jin, Haoyu Wang, Xingliang Wu, Xiaocheng Fang, Xiang Lan, Zihan Wang, Deyun Zhang, Bo Liu, Yingying Zhang, Xian Wu, et al. 2026. ECG-R1: Protocol- Guided and Modality-Agnostic MLLM for Reliable ECG Interpretation.arXiv preprint arXiv:2602.04279(2026)
Pith/arXiv arXiv 2026
-
[21]
Elizabeth Knight, Evangelos K Oikonomou, Arya Aminorroaya, Aline F Pedroso, and Rohan Khera. 2026. Wearable-Echo-FM: an ECG echo foundation model for 1-lead electrocardiography.European Heart Journal-Digital Health7, 4 (2026), ztag049
2026
-
[22]
Georgios A Kaissis, Marcus R Makowski, Daniel Rückert, and Rickmer F Braren
-
[23]
Joon-Myoung Kwon, Soo Youn Lee, Ki-Hyun Jeon, Yeha Lee, Kyung-Hee Kim, Jinsik Park, Byung-Hee Oh, and Myong-Mook Lee. 2020. Deep learning–based algorithm for detecting aortic stenosis using electrocardiography.Journal of the American Heart Association9, 7 (2020), e014717
2020
-
[24]
Sravan Kumar Lalam, Hari Krishna Kunderu, Shayan Ghosh, Harish Kumar, Samir Awasthi, Ashim Prasad, Francisco Lopez-Jimenez, Zachi I Attia, Samuel Asirvatham, Paul Friedman, et al. 2023. Ecg representation learning with multi- modal ehr data.Transactions on Machine Learning Research(2023)
2023
-
[25]
Gloria Hyunjung Kwak, Dana Moukheiber, Mira Moukheiber, Lama Moukheiber, Sulaiman Moukheiber, Neel M Butala, Leo A Celi, and Christina W Chen. 2025. Large open access database of echocardiogram reports in intensive care unit patients.Scientific Data12, 1 (2025), 1153
2025
-
[26]
Che Liu, Cheng Ouyang, Zhongwei Wan, Haozhe Wang, Wenjia Bai, and Rossella Arcucci. 2025. Knowledge-enhanced multimodal ecg representation learning with arbitrary-lead inputs.arXiv preprint arXiv:2502.17900(2025)
Pith/arXiv arXiv 2025
-
[27]
Che Liu, Zhongwei Wan, Sibo Cheng, Mi Zhang, and Rossella Arcucci. 2024. Etp: Learning transferable ecg representations via ecg-text pre-training. InICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 8230–8234
2024
-
[28]
Jun Li, Che Liu, Sibo Cheng, Rossella Arcucci, and Shenda Hong. 2024. Frozen language model helps ecg zero-shot learning. InMedical Imaging with Deep Learning. PMLR, 402–415
2024
-
[29]
Yeongyeon Na, Minje Park, Yunwon Tae, and Sunghoon Joo. 2024. Guiding masked representation learning to capture spatio-temporal relationship of elec- trocardiogram.arXiv preprint arXiv:2402.09450(2024)
Pith/arXiv arXiv 2024
-
[30]
Guangkun Nie, Gongzheng Tang, Yujie Xiao, Jun Li, Shun Huang, Deyun Zhang, Qinghao Zhao, and Shenda Hong. 2025. Anyppg: An ecg-guided ppg foundation model trained on over 100,000 hours of recordings for holistic health profiling. arXiv preprint arXiv:2511.01747(2025)
Pith/arXiv arXiv 2025
-
[31]
Che Liu, Zhongwei Wan, Cheng Ouyang, Anand Shah, Wenjia Bai, and Rossella Arcucci. 2024. Zero-shot ecg classification with multimodal learning and test-time clinical knowledge enhancement.arXiv preprint arXiv:2403.06659(2024)
Pith/arXiv arXiv 2024
-
[32]
Timothy J Poterucha, Linyuan Jing, Ramon Pimentel Ricart, Michael Adjei-Mosi, Joshua Finer, Dustin Hartzel, Christopher Kelsey, Aaron Long, Daniel Rocha, Jeffrey A Ruhl, et al. 2025. Detecting structural heart disease from electrocardio- grams using AI.Nature644, 8075 (2025), 221–230
2025
-
[33]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. InInternational conference on machine learning. PmLR, 8748–8763
2021
-
[34]
David Ouyang, Bryan He, Amirata Ghorbani, Neal Yuan, Joseph Ebinger, Curtis P Langlotz, Paul A Heidenreich, Robert A Harrington, David H Liang, Euan A Ashley, et al. 2020. Video-based AI for beat-to-beat assessment of cardiac function. Nature580, 7802 (2020), 252–256
2020
-
[35]
Milos Vukadinovic, I-Min Chiu, Xiu Tang, Neal Yuan, Tien-Yu Chen, Paul Cheng, Debiao Li, Susan Cheng, Bryan He, and David Ouyang. 2026. Comprehensive echocardiogram evaluation with view primed vision language AI.Nature650, 8103 (2026), 970–977
2026
-
[36]
Wai-Chak Wong, Che Liu, Pierre Elias, John Weston Hughes, Chun-Yu Leung, Xiao-Yan Qian, Hang-Long Li, Yuk-Ming Lau, Chao-Fan Tao, Ali Choo, et al
-
[37]
Alvaro E Ulloa-Cerna, Linyuan Jing, John M Pfeifer, Sushravya Raghunath, Jef- frey A Ruhl, Daniel B Rocha, Joseph B Leader, Noah Zimmerman, Greg Lee, Steven R Steinhubl, et al. 2022. rECHOmmend: an ECG-based machine learning approach for identifying patients at increased risk of undiagnosed structural heart disease detectable by echocardiography.Circulati...
2022
-
[38]
Han Yu, Peikun Guo, and Akane Sano. 2024. Ecg semantic integrator (esi): A foundation ecg model pretrained with llm-enhanced cardiological text.arXiv preprint arXiv:2405.19366(2024)
Pith/arXiv arXiv 2024
-
[39]
Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer. 2023. Sig- moid loss for language image pre-training. InProceedings of the IEEE/CVF inter- national conference on computer vision. 11975–11986
2023
-
[40]
Contrastive Multi-modal Training with Electrocardiography and Natural Language Echocardiography Reports for Zero-shot Prediction of Structural Heart Disease.medRxiv(2025), 2025–09
2025
-
[41]
Xiaoxi Yao, David R Rushlow, Jonathan W Inselman, Rozalina G McCoy, Thomas D Thacher, Emma M Behnken, Matthew E Bernard, Steven L Rosas, Abdulla Akfaly, Artika Misra, et al. 2021. Artificial intelligence–enabled electro- cardiograms for identification of patients with low ejection fraction: a pragmatic, randomized clinical trial.Nature medicine27, 5 (2021...
2021
-
[44]
Xue Zhou, Tianhui Li, Hiromasa Hayama, Keijiro Nakamura, Shing-Hong Liu, Wenxi Chen, and Xin Zhu. 2025. Diagnosis of cardiac conditions from 12-lead KDD ’27, August, 2027, San Jose, CA, USA Fang et al. electrocardiogram through natural language supervision.npj Digital Medicine8, 1 (2025), 697. A Acknowledgments As an informal quality check, five senior ca...
2025
-
[2020]
Secure, privacy-preserving and federated machine learning in medical imaging.Nature Machine Intelligence2, 6 (2020), 305–311
2020
-
[2025]
In34th International Joint Conference on Artificial Intelligence, IJCAI
DERI: Cross-Modal ECG Representation Learning with Deep ECG-Report Interaction. In34th International Joint Conference on Artificial Intelligence, IJCAI
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.