{"id":"3a6a6247-8ff6-4fb8-99db-63b984198af8","arxiv_id":"2506.08229","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":1.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A literature review asserting AI improves CVD detection and reduces costs, with a quantitative table that cites studies missing from its reference list.","lead":"This paper is a narrative review arguing that machine learning can improve early cardiovascular disease diagnosis and lower healthcare costs. It reports a table of accuracy metrics for several named studies, but provides no data, code, or verifiable systematic review evidence to support its claims.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central empirical evidence in Table 1 is unverifiable: none of the five cited studies appear in the reference list, and the table's citation markers point to unrelated papers. The cost-reduction claim therefore rests on data whose provenance cannot be checked.","rationale":"The reader's weakest assumption—that the five studies in Table 1 exist and that their reported performance metrics are accurate and comparable—is exactly the load-bearing concern. My independent reading of the manuscript confirms the same problem: the reference list contains no matching author-year entries for those five studies, and the in-text citation numbers used in Table 1 correspond to completely unrelated references. This is not a stylistic or cosmetic flaw; it is the empirical foundation for the paper's central claim. Without Table 1, the paper offers only a restatement of background literature and unsupported assertions about cost reduction. The internal contradictions in the study selection counts reinforce the unreliability but do not change the verdict. A targeted external search for the five studies would settle the matter, and if the entries are untraceable, the paper's quantitative conclusion should not be accepted. Since my analysis does not move the reader's verdict, the appropriate recommendation is UNCHANGED, retaining REJECT.","tokens_in":13801,"tokens_out":2915,"duration_ms":32522,"concrete_test":"Independently locate each of the five Table 1 studies (Smith 2019, Johnson 2020, Lee 2021, Kim 2022, Patel 2023) via PubMed, Google Scholar, and Crossref by author names and titles or DOIs; for any study found, verify whether the reported accuracy, precision, recall, and AUC match Table 1 and whether the dataset matches. If any of the five entries cannot be found or the metrics do not match, Table 1 is unverified and the central empirical conclusion lacks support.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim—that integrating AI improves early CVD detection, patient outcomes, and healthcare costs—is supported quantitatively by Results, Table 1, which reports accuracy, precision, recall, and AUC for five ML models attributed to Smith et al. (2019), Johnson et al. (2020), Lee et al. (2021), Kim et al. (2022), and Patel et al. (2023). This table is load-bearing because the abstract, Results, and Discussion all draw their conclusions from these exact numbers. However, none of these five author-year entries appears in the reference list, and the table's citation numbers [1]–[5] point to Johnson et al. 2018, Topol 2019, Esteva et al. 2017, Gulshan et al. 2016, and LeCun et al. 2015—none of which are CVD machine-learning model papers reporting the metrics given. The methodology is also internally inconsistent: the Study Selection Process states 'five major works and twelve other works' and then says five papers were selected, while the Figure 1 description says 20 papers were chosen and 5 selected. No search log, data extraction form, or dataset-level validation is provided. If the Table 1 entries cannot be traced to real, methodologically sound studies, the central quantitative support for the paper's conclusion collapses. The broader literature plausibly supports AI-assisted CVD diagnosis, but this preprint supplies no trustworthy evidence for its specific accuracy and cost claims.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript is presented as a systematic review claiming that integrating artificial intelligence (AI) into early cardiovascular disease (CVD) detection improves diagnostic accuracy, reduces diagnostic time, improves patient outcomes, and lowers healthcare costs. The central quantitative evidence is Table 1, which reports accuracy, precision, recall, and AUC for five machine-learning models attributed to Smith et al. (2019), Johnson et al. (2020), Lee et al. (2021), Kim et al. (2022), and Patel et al. (2023), using datasets such as Framingham, Cleveland, MESA, and UK Biobank. The paper synthesizes these numbers into a broad conclusion that AI and machine learning should be adopted in cardiovascular diagnostics. No code, data extraction forms, or reproducible analysis artifacts are provided.","tokens_in":13984,"tokens_out":4159,"duration_ms":47871,"significance":"If the quantitative claims were properly supported, this paper could serve as a useful scoping review of AI-based CVD diagnostics and their potential economic implications. However, the paper's specific contribution is not established: the Table 1 metrics cannot be traced to the listed references, the study selection process is internally inconsistent, and no cost or patient-outcome data are analyzed. The general idea that AI tools can assist cardiovascular diagnosis is consistent with existing literature, but the manuscript supplies no verifiable new evidence and no reproducible artifacts (no PRISMA checklist, no search log, no data extraction table).","major_comments":[{"comment":"Table 1 is the sole quantitative basis for the abstract, Results, and Discussion claims, but none of the five author-year entries (Smith et al. 2019, Johnson et al. 2020, Lee et al. 2021, Kim et al. 2022, Patel et al. 2023) appears in the reference list. The bracketed citations [1]–[5] in the table point instead to Johnson et al. 2018, Topol 2019, Esteva et al. 2017, Gulshan et al. 2016, and LeCun et al. 2015, none of which is a CVD machine-learning study reporting the displayed accuracy, precision, recall, or AUC values. The central empirical results are therefore unverifiable, and the conclusions drawn from them are unsupported as written.","section":"Results, Table 1"},{"comment":"The Study Selection Process states that the final selection \"were five major works and twelve other works,\" implying 17 included studies, while the Figure 1 description says 150 records were screened, 50 were identified by title/abstract, 20 were chosen for full-text analysis, and 5 papers were selected. The manuscript never reconciles the 17 papers with the 5 selected, and no list of the 12 or 17 records is given. This inconsistency makes the systematic review process non-reproducible and undermines the claim that the review followed a clear selection protocol.","section":"Material and methods, Study Selection Process"},{"comment":"The methods section does not report a search log, exact database query strings, search dates, duplicate-handling procedures, or a risk-of-bias/quality assessment. For a manuscript explicitly labeled a systematic review, these elements are standard and load-bearing: without them, the reader cannot determine whether the five studies in Table 1 were systematically identified or whether any inclusion/exclusion criteria were actually applied. The absence of this documentation is not merely a presentation issue because the paper's conclusions depend on the trustworthiness of the review process.","section":"Material and methods, Data Extraction and Analysis"},{"comment":"The accuracy, precision, recall, and AUC values in Table 1 are diagnostic performance metrics; they do not provide evidence about healthcare expenditures, resource utilization, length of stay, or patient-centered outcomes. The conclusion that AI implementation would \"consequently decrease the overall healthcare expenditures through timely intervention\" is an inferential leap unsupported by any economic analysis, cost data, or before-after comparison in the manuscript. No cost outcomes were extracted or reported, despite the title and abstract promising a connection to healthcare costs.","section":"Discussion and Conclusion"},{"comment":"Table 1 reports point estimates for each model without sample sizes, confidence intervals, or external validation details. Two rows (Smith et al. and Kim et al.) use the Framingham Heart Study dataset with different models, but no statistical comparison or paired testing is provided. The text's claim that \"neural networks and boosting algorithms to be superior\" is therefore not supported by the table as presented; the differences could easily be within sampling variability.","section":"Results"}],"minor_comments":[{"comment":"The abstract contains grammatical and typographical errors, including \"lee time consuming\" and \"the diagnosis of CVDs become more accurate,\" which should be corrected for clarity.","section":"Abstract"},{"comment":"The in-text citation \"[81]\" for Attia et al. (2019) is inconsistent with the reference list, which contains only 50 entries; citation numbers are used erratically throughout the manuscript.","section":"Literature review"},{"comment":"Figure 1 is described in the text but the actual flowchart is not included; the manuscript should either display the figure or remove the reference to it.","section":"Figure 1"},{"comment":"The Discussion contains the typo \"mdml\" in the phrase \"benefits that would accrue from using mdml,\" which should read \"ML\".","section":"Discussion"},{"comment":"Several author-year citations in the text (e.g., Nichols et al. 2014, Krittanawong et al. 2017, Bozkurt et al. 2021) do not have matching reference list entries under those names, making it difficult to trace the cited literature.","section":"References"},{"comment":"The search strategy should specify the exact query strings, database-specific search dates, and how the 2010–2024 range was applied, since the description is too vague to replicate.","section":"Material and methods, Search Strategy"}],"recommendation":"reject","confidential_remarks":"I concur with the stress-test assessment: the load-bearing evidence in Table 1 is unverifiable, and the author-year entries appear nowhere in the reference list. This is not a local citation error that can be fixed by renumbering; the provenance of the numeric values is absent. Unless the authors can supply the full bibliographic details and data extraction records for the five studies, the quantitative core of the paper cannot be repaired within a routine revision. I therefore recommend rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's the short version: this paper is not ready for peer review. Its only quantitative support — Table 1, reporting accuracy, precision, recall, and AUC for five models — is unverifiable. None of the five author-year entries (Smith 2019, Johnson 2020, Lee 2021, Kim 2022, Patel 2023) appears in the reference list, and the table's citation markers [1]–[5] point to Johnson et al. 2018, Topol 2019, Esteva et al. 2017, Gulshan et al. 2016, and LeCun et al. 2015 — none of which is a CVD machine-learning study reporting those metrics. The abstract and conclusion lean directly on these numbers, so the load-bearing support for the main claim collapses.\n\nWhat the paper does do okay: it assembles the standard talking points about AI in cardiovascular diagnosis, and its general direction — AI may improve early detection and reduce costs — is consistent with a broad, credible literature. The discussion cites real work (Hannun et al. 2019, Attia et al. 2019, Weng et al. 2017). So the topic is real, and the narrative is not crazy.\n\nThe soft spots go beyond the missing citations. The methods section self-contradicts: it says 'five major works and twelve other works' were selected, then says five papers were selected, and Figure 1's description says 20 papers were chosen and 5 selected. There is no search log, no data extraction form, no sample sizes, no error bars, and no cost model. The cost-reduction claim is made with zero economic analysis. These are not cosmetic issues; they make the review's evidence base untrustworthy.\n\nWho is this for? Possibly a reader who wants a very high-level overview of AI in cardiology and doesn't care about rigor. But as a systematic review — which is what the paper claims to be — it fails basic standards of verifiability and internal consistency. I would not put it in front of referees. Desk reject is the right call. If the authors rework it, they need to actually include the primary studies with correct citations, report extraction details, and either provide a real cost model or drop the cost claim entirely.","headline":"Table 1's five 'studies' don't exist in the reference list, so the paper's quantitative case for cost savings is untraceable — desk reject.","tokens_in":14613,"tokens_out":1948,"would_cite":false,"duration_ms":21350,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This systematic review argues that adding AI to cardiovascular diagnostics makes early detection faster and more accurate, and that this would lower healthcare costs and improve patient outcomes.","keywords":["artificial intelligence","cardiovascular disease","early detection","machine learning","deep learning","healthcare costs","patient outcomes","systematic review"],"falsifier":"Re-run the five models on their named datasets under a single preprocessing and validation protocol; if the reported accuracy, precision, recall, and AUC values do not reappear, or if the underlying studies cannot be located, the paper's claim that AI diagnostics outperform conventional care is falsified.","tokens_in":13540,"feed_emoji":"🫀","tokens_out":11320,"duration_ms":114811,"temperature":0.7,"pith_summary":"This systematic review argues that adding artificial intelligence to cardiovascular diagnostics makes early detection of heart disease faster and more accurate than conventional clinical evaluation, and that the shift would lower healthcare spending while improving patient outcomes. The paper's evidence is a table of five machine-learning studies reporting accuracies between 0.85 and 0.92 across four public datasets. On that basis, the authors conclude that AI-assisted screening should become a routine complement to traditional diagnosis, catching disease before expensive late-stage treatment is needed.","feed_headline":"AI catches heart disease earlier and cuts costs, review claims","feed_subtitle":"Review of five machine-learning studies says earlier detection means timely intervention and lower spending.","key_machinery":"The central object carrying the argument is Table 1, which lists five machine-learning models, their datasets, and their reported accuracy, precision, recall, and AUC values. The table does the work of demonstrating that AI-based diagnostics perform well across different algorithms and datasets; the paper's conclusion about cost savings and better outcomes is inferred directly from these numbers, with the high and consistent metrics standing in for evidence that AI would improve diagnostic acuity in practice.","core_discovery":"The paper's central claim is that machine-learning models—logistic regression, support vector machines, neural networks, random forests, and gradient boosting—can predict cardiovascular risk from routine clinical data with high accuracy, precision, recall, and AUC, and that this diagnostic acuity translates into earlier intervention, better patient outcomes, and reduced overall healthcare expenditures. The authors present this as a systematic synthesis of five studies, with the neural network on the MESA dataset reported as the best performer (accuracy 0.92, AUC 0.94) and gradient boosting on UK Biobank close behind (accuracy 0.89, AUC 0.91). They further claim that the consistency of results across Framingham Heart Study, Cleveland Heart Disease, MESA, and UK Biobank shows the models are transferable across patient populations, making them suitable for integration into varied clinical settings.","pith_inferences":["A testable extension would be to run the five model/dataset combinations under a single preprocessing and validation protocol; comparable performance would turn the table into a much stronger claim than the current narrative.","The cost-reduction argument is plausible but under-specified; the actual savings would depend on the alert threshold, since a high-recall model generates more false positives and more follow-up visits.","The review implicitly calls for prospective, externally validated trials comparing AI-assisted screening with usual care on hard outcomes such as mortality, not just AUC—an inference from the paper's own stated limitations."],"forward_implications":["If the claim holds, AI-assisted ECG interpretation could become a standard screening step in primary care, detecting arrhythmias and silent dysfunction earlier than current practice.","Clinical decision-support tools trained on electronic health records could flag high-risk patients for preventive treatment before they progress to late-stage disease.","Hospitals adopting these models would spend less on expensive terminal and emergency cardiovascular care, offsetting the cost of algorithm development and integration.","The reported transferability across Framingham, Cleveland, MESA, and UK Biobank suggests the models could generalize to diverse patient populations, supporting adoption beyond a single institution.","Widespread use of AI monitoring, including wearable devices, would shift cardiovascular care from reactive treatment toward preventive, continuous assessment."],"supporting_citations":[{"why":"Table 1's first row; supplies the logistic-regression performance on the Framingham Heart Study that anchors the review's accuracy claims.","marker":"Smith et al. (2019) [1]"},{"why":"Table 1's second row; provides SVM metrics on the Cleveland Heart Disease dataset, extending evidence to a second population.","marker":"Johnson et al. (2020) [2]"},{"why":"Table 1's third row; reports the best-performing neural network on MESA, the strongest single evidence point for deep learning.","marker":"Lee et al. (2021) [3]"},{"why":"Table 1's fourth row; reports random-forest metrics on Framingham, supporting the ensemble-learning conclusion.","marker":"Kim et al. (2022) [4]"},{"why":"Table 1's fifth row; reports gradient-boosting metrics on UK Biobank, supporting the boosting-algorithm conclusion.","marker":"Patel et al. (2023) [5]"},{"why":"Cited in the literature review as matching cardiologist-level arrhythmia detection; underpins the claim that AI equals expert performance.","marker":"Hannun et al. (2019) [1]"},{"why":"Cited for using machine learning on ECG to diagnose cardiac contractile dysfunction; supports early detection from routine signals.","marker":"Attia et al. (2019) [81]"}],"fun_headline_variants":["AI predicts heart disease early, cuts costs in review of 5 studies","Machine learning finds heart disease earlier, says systematic review","AI models detect heart risk accurately, cutting costs per review","Early AI detection of heart disease lowers costs, review finds","Review: AI improves early heart disease detection and cuts costs"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central claim rests on the five studies in Table 1 being real and on their reported accuracy, precision, recall, and AUC values being accurate and comparable; if those numbers are wrong or cannot be reproduced, the conclusion that AI improves early detection and reduces costs is unsupported.","fun_headline_variants_meta":{"raw":{"variants":["AI predicts heart disease early, cuts costs in review of 5 studies","Machine learning finds heart disease earlier, says systematic review","AI models detect heart risk accurately, cutting costs per review","Early AI detection of heart disease lowers costs, review finds","Review: AI improves early heart disease detection and cuts costs"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000607,"raw_usage":{"total_tokens":2741,"prompt_tokens":767,"completion_tokens":1974,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":383,"completion_tokens_details":{"reasoning_tokens":1891}},"tokens_in":383,"tokens_out":1974,"duration_ms":13787,"temperature":1.0,"reasoning_tokens":1891,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T05:16:30.338109+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the five models on their named datasets under a single preprocessing and validation protocol; if the reported accuracy, precision, recall, and AUC values do not reappear, or if the underlying studies cannot be located, the paper's claim that AI diagnostics outperform conventional care is falsified.","supporting_citations":[],"review_version":1}