{"id":"229bfc8c-1ba9-49fe-b6f7-5fa6490a75f5","arxiv_id":"2506.03698","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":2.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A narrative review of AI in cardiovascular imaging and signals, summarizing selected CT, MRI, ECG, and ultrasound studies with a brief limitations discussion.","lead":"This preprint reviews recent artificial intelligence applications in cardiovascular disease across CT, MRI, ECG, and ultrasound. It summarizes selected studies and proposes that AI improves diagnostic workflows, while noting that models often cannot validate the correctness of input data.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No significant objection beyond the reader's selection-bias concern; the review's central claim is an unsupported generalization from an unexplained, possibly self-weighted study set.","rationale":"The reader's weakest_assumption identifies the same core issue: the paper lacks a search strategy and inclusion criteria, making the selected studies potentially non-representative. I agree that this is the most load-bearing concern. The paper's strongest claim is a broad superiority assertion; that claim rests entirely on the studies listed in Tables 1–4. Since the tables are not generated by a transparent, reproducible method, the review cannot empirically support 'surpassing human capabilities' as a general statement. The self-citation density in Section 3.4 strengthens the concern because it suggests a convenience sample rather than a systematic literature coverage. I would keep the verdict CONDITIONAL rather than REJECT, since the underlying studies exist and are mostly real, but the review's synthesis needs a methodology section and a more calibrated claim. My concrete test would directly quantify whether the headline claim depends on the suspicious subset.","tokens_in":8944,"tokens_out":615,"duration_ms":7984,"concrete_test":"Reconstruct the selection for Table 4: list each US citation, classify it as self-citation (authors overlapping with the review authors) and as reporting a direct human-comparison benchmark. Then check whether removing all self-citations or all studies lacking a direct human-comparison changes the abstract's claim of surpassing human performance. If the claim survives only with these studies included, the strongest statement is unsupported.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's central claim is that deep learning 'surpasses human capabilities in diagnostic accuracy and workflow efficiency' across CT, MRI, ECG, and US. The evidence base is Tables 1–4, but the review never states a search strategy, inclusion criteria, or quality assessment. Section 3.4 is particularly skewed: of the 16 ultrasound citations, roughly 10 are authored by Pu, Li, Liang, Lu, or their close collaborators (refs 29–33, 35–42). This creates a representation problem: if the selection is not systematic, the aggregated impression of 'surpassing human capabilities' is an artifact of cherry-picking successful papers from one research group, rather than a literature-level conclusion. The paper also makes global claims without quantitative support (e.g., 'surpassing human capabilities' appears in abstract and conclusion without a pooled comparison or effect size). This is a load-bearing premise because a review's validity rests on its selection being representative; the absence of methodology plus the visible self-citation concentration undermines the generalizability of the strongest claim. Additionally, the limitations section (Section 5) only addresses input-data validation and computational cost, not selection bias, publication bias, or heterogeneity across studies.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a narrative review of artificial intelligence (AI) applications in cardiovascular disease research, organized by modality: CT, MRI, ECG, and ultrasound (US). It describes deep learning architectures, surveys a selected set of studies in four tables, summarizes dataset sizes, and discusses limitations centered on input-data validation and computational cost. The paper's central claim, stated in the abstract, introduction, and conclusion, is that deep learning models 'surpass human capabilities in diagnostic accuracy and workflow efficiency' across these imaging and signal modalities. The review concludes by calling for self-validating AI systems and multimodal learning frameworks for future clinical deployment.","tokens_in":9266,"tokens_out":3079,"duration_ms":30075,"significance":"If its central claim were properly supported, this review would be a useful consolidation of recent AI applications across four major cardiovascular diagnostic modalities. The paper has strengths: it covers a broad range of applications, includes study-level limitations in its tables, and identifies important practical concerns such as input-data validation and model deployability. It also cites several high-impact recent studies, including Nature Medicine and Nature Communications papers. However, the evidence presented in the tables does not substantiate the headline claim of surpassing human capabilities; several table entries explicitly report mixed or inferior performance. The absence of a systematic literature search, inclusion criteria, or quality assessment further weakens the generalizability of the conclusions. The review is therefore more of an illustrative overview than a definitive assessment of the field.","major_comments":[{"comment":"The claim that deep learning 'surpasses human capabilities in diagnostic accuracy and workflow efficiency' is not supported by the evidence reported in the manuscript. Table 2 (MRI-2) lists 'low rates of accurate classifications in LVEF' as a limitation; Table 2 (MRI-4) reports an F1 score of 0.931 versus 0.927 for physicians, which is a match rather than a clear surpass, and specifically notes 'lower F1 scores for myocarditis'; Table 1 (CT-2) reports 89% agreement with human observers, not superiority. No pooled effect sizes, meta-analytic comparison, or statistical synthesis is provided. This is a load-bearing assertion because it is the central message of the paper. The authors should temper the claim to reflect the mixed and modality-specific evidence, or provide a systematic quantitative comparison.","section":"Abstract; Section 1; Section 6"},{"comment":"The review does not describe a search strategy, inclusion criteria, study selection process, or quality assessment. Without this information, the studies in Tables 1 through 4 cannot be regarded as a representative sample of the literature, and the aggregate impression of the field may be an artifact of the selection. This is especially problematic for a review whose conclusions are global statements about AI capabilities. The authors should either add a methodology section describing how studies were identified and selected, or explicitly reposition the paper as an illustrative narrative review and avoid literature-level generalizations.","section":"Section 3; Tables 1-4"},{"comment":"The ultrasound subsection relies heavily on the authors' own prior publications. Of the approximately fifteen studies described in this subsection, a large majority are authored by Pu, Liang, Li, Lu, He, Yang, or their close collaborators (references 29-42). In a non-systematic review, this visible concentration creates a representation bias that can inflate the apparent maturity of AI in fetal cardiac ultrasound. The authors should disclose this overlap, include a broader set of independent studies, and temper any claims about the state of the field in this subsection.","section":"Section 3.4"},{"comment":"The limitations section only addresses input-data validation and computational cost, but a review of this scope should also discuss review-level limitations such as publication bias, heterogeneity of datasets and evaluation metrics across the surveyed studies, potential demographic or selection biases (some of which are noted in Table 3, ECG-4), and the lack of external validation in many studies. Without discussing these, the review gives an incomplete picture of the reliability of the evidence it summarizes.","section":"Section 5"}],"minor_comments":[{"comment":"Reference 28 is identical to reference 25 (Hughes et al.), and the in-text citation at [28] appears to be a duplicate. Please correct or remove the redundant reference.","section":"References"},{"comment":"The phrase 'compared to CT MRI and US' should be 'compared to CT, MRI, and US'; there is also a missing comma after 'CT'.","section":"Section 3.4"},{"comment":"In the ECG-3 row, 'inhealthy population' should be 'in healthy population'.","section":"Table 3"},{"comment":"Figure 2 is cited as references [46] and [47], but these references are a chest radiograph database and a medical image segmentation model, respectively. The relationship between the figure and these references is unclear and should be clarified or the citations corrected.","section":"Figure 2"},{"comment":"The sentence 'Transformer originally introduced in the domain of natural language processing (NLP), has progressively demonstrated its versatility' is grammatically incomplete; it should be revised to a full sentence, for example 'Transformers, originally introduced for natural language processing, have progressively demonstrated their versatility in medical image segmentation.'","section":"Section 2"},{"comment":"Figure 1 is described as a summary of selected studies' datasets, but the figure itself is not included in the text available for review. Please ensure the figure is self-explanatory and includes clear labels for modalities and sample-size statistics.","section":"Section 4"}],"recommendation":"major_revision","confidential_remarks":"The manuscript's central claim is not supported by its own evidence, and the review lacks the methodological transparency needed for a literature-level conclusion. The paper may be suitable for a journal that accepts narrative reviews, provided the authors substantially revise the claims, add a selection methodology or clearly frame the scope as illustrative, and disclose the self-citation concentration in the ultrasound section. These issues are fixable within the scope of the manuscript, so I recommend major_revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Yuanlin Mo et al. have written a narrative review of AI across CT, MRI, ECG, and US in cardiovascular disease. It is exactly what it looks like: a broad, modality-by-modality summary with one table per modality and no new data, methods, or systematic synthesis. If you need a quick orientation to some names and results in these four areas, it is usable, but that is the ceiling of its value.\n\nThe paper does a few things fine. The tables give a compact snapshot of study purpose, dataset size, strengths, and limitations. The CT and MRI sections cite some substantial recent work (Miller et al., Wang et al. in Nature Medicine). The limitations section at least names input-data validation and computational cost, though it does little more than name them.\n\nThe soft spots are real and load-bearing. The abstract and introduction claim deep learning 'surpasses human capabilities in diagnostic accuracy and workflow efficiency.' The paper's own tables do not support that: Wang et al. report low rates of accurate LVEF classification, Wang et al. (MRI-4) have lower F1 for myocarditis, and Hughes et al. note demographic biases. That overclaim is not a minor rhetorical slip; it is the paper's headline claim.\n\nThe bigger problem is selection. There is no stated search strategy, inclusion criteria, or quality assessment. Tables 1–4 are therefore uninterpretable as a literature summary, and the imbalance in Section 3.4 makes it worse: most of the ultrasound citations are from the authors' own collaborative circle (Pu, Li, Lu, Liang, He, Tseng, Yang). Of seven entries in Table 4, six are from that group. That is not evidence of a field-wide trend; it is an artifact of the chosen sample. The limitations section does not acknowledge selection bias or publication bias.\n\nThis is not a paper with an underlying result, so the critique is mainly about whether it can serve as a reliable entry point. As written, I would not point a student to it without a warning. The framework is fine; the execution is uncritical. The review could be salvageable if the authors added a methodology section, softened the 'surpassing human capabilities' claims to match the mixed results in their own tables, and broadened the ultrasound coverage. Without that, it is a desk-reject-quality review. I would not send it to peer review in its current form; a serious referee would spend time on a paper whose central claim is unsupported and whose selection is opaque. If the authors resubmit with methodology and corrected claims, it might become a usable survey.","headline":"A readable but uncritical narrative review whose central 'surpassing human capabilities' claim is not supported by its own tables, with an unexplained, self-weighted study selection.","tokens_in":9661,"tokens_out":3504,"would_cite":false,"duration_ms":34593,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This review argues that deep learning on CT, MRI, ECG, and ultrasound can surpass human diagnostic accuracy in cardiovascular disease, and that the main obstacle is models' inability to verify input data.","keywords":["Artificial Intelligence","Cardiovascular Disease","Deep Learning","Cardiac Imaging","Electrocardiography","Computed Tomography","Magnetic Resonance Imaging","Ultrasound"],"falsifier":"A systematic literature search with explicit inclusion criteria that surfaced a substantial body of comparable studies where AI failed to match expert readers—or where pooled sensitivity and specificity were no better than human performance—would falsify the review's general claim. Concretely, re-deriving the four tables from an independent, reproducible search and comparing pooled metrics against human benchmarks would settle it.","tokens_in":8731,"feed_emoji":"🫀","tokens_out":3478,"duration_ms":39155,"temperature":0.7,"pith_summary":"This review paper aims to establish that deep-learning systems applied to four cardiovascular diagnostic modalities—computed tomography, magnetic resonance imaging, electrocardiography, and ultrasound—can automate image and signal analysis with accuracy and efficiency that exceed human experts. The paper surveys selected studies in each modality, reporting metrics such as high agreement with human readers on coronary calcium scoring, expert-level classification of cardiac function on MRI, and strong screening performance for atrial fibrillation and ventricular dysfunction from ECG. It also argues that a critical unresolved challenge remains: most AI models cannot tell whether the input data are correct, so mistakes in input can propagate into diagnostic errors. The intended upshot is that AI is ready to meaningfully reshape cardiovascular diagnostics, but only if validation protocols and self-checking mechanisms are built around it.","feed_headline":"Deep learning beats human experts at heart diagnostics, review claims","feed_subtitle":"Four imaging modalities reviewed: CT, MRI, ECG, and ultrasound; the catch is that models cannot check their own inputs.","key_machinery":"The organizing mechanism is not a single algorithm but a modality-by-modality review structured by four tables, one for CT, MRI, ECG, and ultrasound. Each table entry pairs a model and its dataset size with the study's stated strengths and limitations, and the paper aggregates these selected studies into broad conclusions about deep learning's capabilities. This assembled evidence base, together with a dataset-summary figure showing that ECG studies draw on far larger cohorts than imaging studies, carries the argument that AI-driven diagnostics outperform human readers and that data-validation failure is the main remaining barrier.","core_discovery":"The paper's central claim is that deep-learning architectures—including convolutional neural networks, recurrent neural networks, and generative adversarial networks—enable automated analysis of cardiovascular imaging and physiological signals at a level that surpasses human capabilities in both diagnostic accuracy and workflow efficiency. It supports this by reviewing selected studies: an AI system for congenital heart disease from CT, automated coronary artery calcium scoring that placed 89% of patients into the same risk category as human observers, an end-to-end cardiac magnetic resonance system whose F1 score matched cardiologists with more than a decade of experience, an ECG-based deep network trained on over 1.6 million recordings, an AI-ECG screen for cardiac contractile dysfunction, and an echocardiogram video model that segments the left ventricle with a Dice coefficient of 0.92 and detects heart failure with reduced ejection fraction with an area under the curve of 0.97. Alongside these successes, the paper contends that the field's decisive limitation is that models cannot validate the correctness of their input data, which can propagate diagnostic errors and compromise clinical reliability.","pith_inferences":["If this claim generalizes, the practical bottleneck in cardiovascular AI shifts from improving model accuracy to data-quality assurance, regulatory approval, and prospective testing in diverse patient populations.","The paper's emphasis on input verification suggests a testable design goal: models should be evaluated not only on clean data but on their ability to flag corrupted or mislabeled inputs, since that is where the paper locates the greatest risk.","Because the ECG datasets in the summary are far larger than the imaging datasets, the fastest route to clinical validation may be ECG-based screening followed by confirmatory imaging, rather than attempting to deploy imaging-only AI directly.","A reader could extend the review by comparing hybrid ECG-plus-imaging models against single-modality baselines on the same cohort; the paper implies such multimodal fusion would improve diagnostic confidence but does not test it."],"forward_implications":["AI-enabled ECG analysis could screen for conditions such as atrial fibrillation during normal sinus rhythm and detect left ventricular dysfunction, shifting which patients are referred for expensive imaging.","Automated coronary calcium scoring and end-to-end cardiac MRI interpretation could reduce radiologist workload and make population-level cardiovascular screening faster and cheaper.","Video-based echocardiogram models could provide real-time, beat-to-beat assessment of cardiac function, enabling point-of-care diagnosis of heart failure.","Because models cannot verify input correctness, clinical deployment will require explicit validation protocols or self-checking mechanisms before AI outputs are trusted for patient care.","Hybrid models that combine multimodal data—such as ECG with imaging—and adaptive algorithms are presented as the likely route toward personalized cardiovascular care."],"supporting_citations":[{"why":"Provides the CT-based congenital heart disease classification result that grounds the CT section's accuracy claim.","marker":"[10]"},{"why":"Supplies the automated coronary artery calcium scoring model with 89% agreement with human risk categorization.","marker":"[11]"},{"why":"Supports the claim that AI reduces MRI acquisition time via free-breathing 3D coronary angiography.","marker":"[16]"},{"why":"Shows automated MRI measurements of ejection fraction and mass are highly consistent with expert readings.","marker":"[17]"},{"why":"Provides the end-to-end cardiac MRI system that matches experienced cardiologists in F1 score.","marker":"[19]"},{"why":"Contributes the large-scale 12-lead ECG deep network that anchors the ECG accuracy claims.","marker":"[20]"},{"why":"Supports the claim that AI-enabled ECG can screen for left ventricular dysfunction.","marker":"[24]"},{"why":"Supplies the echocardiogram video model with Dice 0.92 segmentation and heart failure AUC 0.97.","marker":"[43]"},{"why":"Supports the limitation discussion by offering an example of a model that can assess input correctness.","marker":"[44]"}],"fun_headline_variants":["AI heart diagnostics surpass humans but can't verify their own input","Deep learning tops human cardiologists, yet blind to input errors","AI beats experts in heart imaging, but input validation is a gap","Cardiac AI excels at diagnosis, but can't check its own data","Review: AI surpasses humans in heart scans, input trust unresolved"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the studies selected for Tables 1 through 4 fairly represent the broader literature; if the selection overweights successful applications or one group's work, the generalized claim that AI surpasses human diagnostic performance is not established.","fun_headline_variants_meta":{"raw":{"variants":["AI heart diagnostics surpass humans but can't verify their own input","Deep learning tops human cardiologists, yet blind to input errors","AI beats experts in heart imaging, but input validation is a gap","Cardiac AI excels at diagnosis, but can't check its own data","Review: AI surpasses humans in heart scans, input trust unresolved"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000879,"raw_usage":{"total_tokens":3757,"prompt_tokens":862,"completion_tokens":2895,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":478,"completion_tokens_details":{"reasoning_tokens":2805}},"tokens_in":478,"tokens_out":2895,"duration_ms":19884,"temperature":1.0,"reasoning_tokens":2805,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T10:56:17.654596+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A systematic literature search with explicit inclusion criteria that surfaced a substantial body of comparable studies where AI failed to match expert readers—or where pooled sensitivity and specificity were no better than human performance—would falsify the review's general claim. Concretely, re-deriving the four tables from an independent, reproducible search and comparing pooled metrics against human benchmarks would settle it.","supporting_citations":[{"cited_title":"Medical Image Analysis 90, 102953 (2023)","cited_arxiv_id":null,"evidence_quote":"Provides the CT-based congenital heart disease classification result that grounds the CT section's accuracy claim."},{"cited_title":"European radiology 33(1), 321 –329 (2023)","cited_arxiv_id":null,"evidence_quote":"Supplies the automated coronary artery calcium scoring model with 89% agreement with human risk categorization."},{"cited_title":"Magnetic resonance in medicine 86(5), 2837–2852 (2021)","cited_arxiv_id":null,"evidence_quote":"Supports the claim that AI reduces MRI acquisition time via free-breathing 3D coronary angiography."},{"cited_title":"Cardiovascular Imaging 15(3), 413–427 (2022)","cited_arxiv_id":null,"evidence_quote":"Shows automated MRI measurements of ejection fraction and mass are highly consistent with expert readings."},{"cited_title":"Nature Medicine 30(5), 1471 – 1480 (2024)","cited_arxiv_id":null,"evidence_quote":"Provides the end-to-end cardiac MRI system that matches experienced cardiologists in F1 score."},{"cited_title":"Nature communications 11(1), 1760 (2020)","cited_arxiv_id":null,"evidence_quote":"Contributes the large-scale 12-lead ECG deep network that anchors the ECG accuracy claims."},{"cited_title":"Nature medicine 25(1), 70–74 (2019)","cited_arxiv_id":null,"evidence_quote":"Supports the claim that AI-enabled ECG can screen for left ventricular dysfunction."}],"review_version":1}