{"id":"7b0b5371-9254-4043-a47a-1b5c3387e6a5","arxiv_id":"2501.04493","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A systematic review of machine learning for congenital heart disease that compiles datasets, algorithms, and reported metrics, but its tables contain citation and accuracy errors.","lead":"This paper reviews 74 studies that apply machine learning to diagnose congenital heart disease, summarizing datasets, algorithms, and reported accuracy. It is useful as a starting map, but its central tables contain citation and metric errors that need correction before the synthesis can be trusted.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table 6's data extraction is internally inconsistent—duplicate citation numbers, a misassigned dataset, and an implausible Dice score—so the central claim of a reliable 74-paper synthesis is not supported as printed; a corrected extraction audit is required.","rationale":"The reader's weakest assumption is exactly the load-bearing point: the review's reliability rests on Tables 5 and 6 faithfully reporting the method, dataset, and metrics of the 74 primary papers. The internal evidence confirms this assumption fails. The duplicate use of reference [40] for two distinct rows, the repeated references [51] and [128], the impossible Dice score of 1, and the misattribution of the Shanxi dataset to an ECG paper are concrete, checkable data-extraction errors. These are not merely cosmetic; Section 5.1's conventional-ML narrative and Section 6.3.4's DL-versus-TML comparison both draw directly on Table 6. The paper does have some genuine value: it assembles a broad list of CHD datasets and modalities, uses a PRISMA-style screening flow, and explicitly acknowledges the lack of clinical co-authors. Those strengths do not rescue the central synthesis, because a one-stop reference inventory is only useful if the rows in its summary tables are trustworthy. An independent re-extraction audit would settle the matter; if the errors are isolated to a handful of rows, a conditional acceptance with mandatory corrections could be appropriate, but as printed the evidence base is compromised. Therefore I agree with the reader's REJECT verdict and would not adjust it.","tokens_in":35084,"tokens_out":3222,"duration_ms":31346,"concrete_test":"Perform an independent row-by-row re-extraction audit of Table 6: for each of the 74 cited papers, locate the original article and record its method, data type, dataset, and metric values, then compare with the table. At minimum, verify three flagged rows: (a) the Dice value reported by Diller et al. [118] in BMC Medical Imaging 2020; (b) whether the 'Arnaout et al.' and 'Rima et al.' rows both trace to the same Nature Medicine 2021 paper; and (c) whether Section 5.1's Shanxi dataset claim traces to reference [45] or to Luo et al. [83]. If any of these mismatches persists, Tables 5 and 6 cannot serve as the evidence base, and the paper should remain rejected until the extraction is corrected and re-audited.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that it provides a reliable, systematic synthesis of 74 CHD-ML studies, including datasets, algorithms, and reported metrics. This claim requires Tables 5 and 6 to be faithful extractions from the cited primary papers. Internal evidence shows they are not. The same reference number [40] is used for two different 2023 rows in Table 6, 'Arnaout et al.' and 'Rima et al.,' although reference [40] is the 2021 Nature Medicine paper by Arnaout et al.; reference [51] appears twice with different methods; and reference [128] is assigned to three distinct rows. A Dice score of 1 is reported for Diller et al. [118], which is not credible for a real cardiac segmentation result without explicit explanation, and it appears with no dataset listed. Section 5.1 attributes the Shanxi maternal-risk dataset to reference [45], but reference [45] is an ECG LSTM paper by Liu and Kim (EMBC 2018); the Shanxi dataset belongs to Luo et al. [83]. Because every bibliometric trend, modality comparison, and algorithm comparison in Sections 3–6 is built on these tables, the review's central synthesis is unreliable as printed. This is an internal-consistency failure, not a disagreement with field consensus.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper claims to present a systematic review and meta-analysis of 432 references on machine learning for congenital heart disease (CHD) recognition between 2018 and 2024, with an in-depth review of 74 primary studies. It organizes the field by diagnostic modality (PCG, X-ray, echo, pulse oximetry, cardiac MRI, ECG, CT, and non-clinical data), lists reported datasets in Table 5, summarizes methods and evaluation metrics in Table 6, and discusses challenges, research gaps, and future directions. The central claim is that this is the first systematic literature review to provide a comprehensive inventory of CHD-ML datasets and a comparison of ML/DL algorithms.","tokens_in":35432,"tokens_out":10644,"duration_ms":83341,"significance":"If the data extraction were reliable, the paper would be a useful one-stop inventory for the CHD-ML community: it assembles 22 datasets across 8 modalities, tabulates 74 primary studies with algorithms and metrics, follows a PRISMA-style protocol, and includes a candid limitations section (Section 6.4). The paper's main strengths are its breadth of coverage and its explicit attempt to catalog public datasets, a dimension that previous reviews summarized in Table 1 apparently omit. However, the paper's value is entirely contingent on the fidelity of its tables and bibliometric counts; the internal inconsistencies documented below mean that, as printed, the synthesis does not yet support the abstract's claims. The authors also deserve credit for acknowledging the lack of medical co-authorship and the resulting limits on clinical depth.","major_comments":[{"comment":"The reported corpus size is inconsistent: the abstract says 432 references, Section 1.1 says 422 research publications, Section 2.2 says 432 articles, and the PRISMA flow diagram (Fig. 2) shows n=432 and then lists 'Excluded 422' next to the 'Paper for full text review (n=74)' box. The arithmetic of the screening stages (4065 to 2240 to 1920 to 1266 to 980 to 854 to 432 to 74) should be spelled out with correct exclusion counts. If the corpus is 422, the abstract and all bibliometric figures must be updated; if it is 432, Section 1.1 must be corrected.","section":"§1.1, §2.2, Fig. 2"},{"comment":"Table 6 contains multiple duplicate citation labels that make the extraction unreliable. Citation [40] labels both 'Arnaout et al.' and 'Rima et al.' in the 2023 block, although reference [40] is the 2021 Nature Medicine paper by Arnaout et al.; citation [51] is used for two different method/dataset combinations in the same block; and citation [128] is assigned to three rows with different data types (Echo/CT/MRI, non-clinical, and CT). Since Section 6.3.4 derives its TML-vs-DL comparison from Table 6, the table must be re-verified against the primary sources and corrected before the review's conclusions can be evaluated.","section":"Table 6"},{"comment":"The Table 6 row for Diller et al. [118] reports DS c = 1 with no dataset listed. A perfect Dice coefficient on real cardiac MRI is implausible, and Section 5.2.2 describes the same work as reporting a 'segmentation accuracy discrepancy of less than 1%' between synthetic- and real-trained U-Nets. The table should either state the exact metric and dataset or be corrected to the value reported in the original paper.","section":"Table 6, Diller et al. [118]"},{"comment":"Section 5.1 misattributes the Shanxi maternal-risk dataset to reference [45]. The text says 'The work proposed in [45] used a dataset collected from Shanxi Province, China' and describes SVM/RF/LR classification of maternal risk factors. Reference [45] is Liu and Kim's ECG-LSTM paper, which Section 5.2.3 correctly describes as an ECG heartbeat classification study. The Shanxi dataset belongs to Luo et al. [83], as the 2018 row of Table 6 itself indicates. This error directly affects the claimed mapping between algorithms and non-clinical data.","section":"§5.1"},{"comment":"The paper repeatedly assigns the same citation numbers to unrelated studies. Section 5.2.8 attributes to Khan et al. [136] a graph-matching whole-heart CT segmentation method that corresponds to Xu et al. [46], while reference [136] is the authors' SKIMA 2023 paper. Table 5 lists HSS [61] with age range 21-88 years although Section 4.1.3 describes the same dataset as 170 babies. Table 5 also attributes the CirCor DigiScope dataset to reference [60], which in the reference list is an echocardiography-mentoring paper, not the CirCor dataset. Because the dataset inventory is a main claimed contribution, these citation-to-content mismatches must be systematically repaired.","section":"§4, §5.2.8, Table 5"},{"comment":"The abstract states the paper conducts 'a meta-analysis of 432 references,' but the Methods section describes a PRISMA-compliant systematic review with qualitative synthesis only; no effect-size pooling or quantitative meta-analytic estimate is presented anywhere in Sections 5-7. The term 'meta-analysis' should be removed from the abstract unless actual pooling (e.g., of reported accuracy or Dice metrics) is added.","section":"Abstract, §2"}],"minor_comments":[{"comment":"The typo 'congential' appears in the search keywords in Section 2.1, in Figure 2, and in Table 4; please correct to 'congenital' throughout.","section":"§2.1, Fig. 2, Table 4"},{"comment":"The DICOM radiograph dataset is cited as [64] in the text and [78] in Table 5; neither reference appears to correspond to a publicly available 828-radiograph CHD dataset. Please add the correct citation or reconcile the two entries.","section":"§4.2.1, Table 5"},{"comment":"The frequencies in Table 4 are internally inconsistent with the prose: the text says 'congenital heart disease' appears 40 times as a trigram, but the following paragraph says it appears 35 times; please reconcile these counts.","section":"§3.5, Table 4"},{"comment":"Section 6.2.4 says 'only 51% datasets are publicly available,' while Table 5 lists 22 datasets with 11 marked '✓' (50%); please either correct the percentage or justify the rounding, and define whether 'publicly available' includes restricted-access datasets.","section":"§6.2.4, Table 5"},{"comment":"Table 1's row numbering jumps from 3 to 5, omitting S. No. 4; please renumber the rows.","section":"Table 1"}],"recommendation":"major_revision","confidential_remarks":"The number and pattern of citation misassignments in Tables 5 and 6 go beyond ordinary copy-editing slips and suggest that the table extraction was not checked against the reference list. I recommend requesting, as a revision condition, a supplementary audit table that lists, for each of the 74 primary studies, the verified reference number, dataset, method, and reported metrics. The paper may become publishable if that audit is performed and the central trends remain unchanged."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis survey has the right aim but the execution does not support the claim. The central tables that the whole synthesis rests on are not reliable. I would not cite it as a reference.\n\nWhat is genuinely useful: the paper maps the CHD-ML landscape better than earlier reviews, brings together 22 datasets across modalities in Table 5, and organizes algorithms by family (CNN, GAN, RNN, hybrid, U-Net, YOLO) in Section 5. That structure is a helpful starting point for someone new to the field. The PRISMA flow and the explicit limitations section (no medical co-authors, limited clinical depth) are honest. The bibliometric part (year counts, countries, funding) is straightforward and largely descriptive.\n\nThe soft spots are not minor. The abstract says 432 references, Section 1.1 says 422, and Section 2.1 says 432. More seriously, Table 6 has multiple duplicate and misassigned citation numbers: [40] is given to both Arnaout and Rima; [128] appears three times under two author spellings; [51] appears twice with different methods. Diller et al. is reported with a Dice score of 1 and no dataset, which is not credible. Section 5.1 attributes the Shanxi maternal-risk dataset to reference [45] (an ECG paper by Liu and Kim), when it actually belongs to Luo et al. [83], and Table 6 even labels Luo's row as PCG/ZCHSound. These errors are load-bearing because every algorithm comparison and trend in Sections 3–6 is built on those tables.\n\nThe paper is salvageable: a careful audit of the 74 extracted rows, reconciling citation numbers and re-reading each primary paper, would fix most of the problems. But as printed, the synthesis is unreliable. It might be a useful cautionary example for a reading group on SLR data extraction, and entry-level researchers could use the dataset list as a pointer to the primary literature. I would desk-reject it in current form if I were the editor, but would welcome a resubmission after a full data-extraction audit.","headline":"The survey has the right aim but the central tables don't hold up; I wouldn't cite it as printed, though a careful audit could make it useful.","tokens_in":35855,"tokens_out":6202,"would_cite":false,"duration_ms":53678,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This systematic review sets out to be the first comprehensive map of machine learning in congenital heart disease diagnosis: it aggregates 432 Scopus-indexed papers, reviews 74 in depth, and tabulates 22 datasets across eight diagnostic…","keywords":["congenital heart disease","machine learning","deep learning","systematic literature review","medical imaging datasets","echocardiography","electrocardiogram","diagnostic modalities"],"falsifier":"Take any ten rows of Table 6, open the cited primary papers, and check whether the dataset name, data type, method, and metric values match. If more than two of the ten rows misreport the source, the review's per-study performance comparisons and its trend claims lose their evidentiary basis.","tokens_in":34851,"feed_emoji":"🫀","tokens_out":5693,"duration_ms":49529,"temperature":0.7,"pith_summary":"This paper is a systematic literature review of machine learning (ML) and deep learning (DL) applied to congenital heart disease (CHD) diagnosis. It aggregates 432 Scopus-indexed papers published between 2018 and 2024, selects 74 for detailed review, and compiles the datasets, algorithms, evaluation metrics, and reported results into two large tables. The authors claim this is the first review in the field to offer a complete public dataset inventory, spanning heart sounds, X-rays, echocardiography, MRI, ECG, CT, pulse oximetry, and non-clinical data. They report that DL methods dominate (81% of the 74 studies), echocardiography is the most-used modality, only about half of the datasets are publicly available, and the main barriers are data scarcity, annotation cost, and privacy. If the compilation is accurate, it gives new researchers a single reference for choosing datasets and methods, while highlighting gaps such as the unexplored use of vision transformers and synthetic data.","feed_headline":"Review maps 74 machine-learning studies on congenital heart disease","feed_subtitle":"Catalogs 22 datasets across eight diagnostic modalities and flags the gaps that remain.","key_machinery":"The load-bearing mechanism is the two extraction tables. Table 5 catalogs 22 datasets across eight modalities, giving availability, age range, and CHD types; Table 6 records each of the 74 studies' ML method, data type, dataset, and reported evaluation metrics. These tables ground every trend claim in the paper: the DL-to-classical ratio, the modality distribution, the accuracy comparisons, and the public-data share. The screening funnel described in the methods (from 4065 initial records down to 74 full-text reviews) is the selection machinery that produces these tables.","core_discovery":"The central discovery is the landscape itself: CHD-ML research is small, recent, and fragmented. Across 74 studies, DL outnumbers classical ML by 60 to 14; CNNs are the most common architecture; echo is the leading modality at 26%, followed by CMR and ECG at 18% each; and only 51% of the reported datasets are public. The paper's contribution is a structured inventory in Tables 5 and 6 that maps each study to its modality, algorithm, dataset, and evaluation metrics, alongside a bibliometric profile of contributing countries, journals, publishers, and funding patterns.","pith_inferences":["Because the table rows show internal inconsistencies (one reference assigned to two different studies, a reported Dice of exactly 1, and a dataset attributed to the wrong source paper), the per-study metric columns should be treated as unreliable even if the dataset inventory and broad trend claims survive.","If only half of the datasets are public, the reported state-of-the-art accuracies are likely optimistic; a community benchmark with standardized splits and hidden test sets would be a natural extension that the paper points toward but does not propose.","The review's own evidence that GAN-generated synthetic images train segmenters within about 1% of real-data performance suggests that synthetic data, rather than larger real collection, is a promising route to break the annotated-data bottleneck.","A machine-readable version of Tables 5 and 6, with stable dataset identifiers and links to each primary study, would convert this static review into a living map; without such a resource, the inventory will age quickly as new datasets and models appear."],"forward_implications":["A researcher entering CHD-ML can use Table 5 to locate a public dataset for echo, MRI, ECG, CT, or heart-sound work without re-searching the literature.","The 81% DL share and the reported metric ranges indicate that hand-crafted feature approaches are largely superseded, though the paper notes direct comparisons across studies are blocked by dataset heterogeneity.","The claim that only 51% of datasets are public identifies reproducibility as a structural problem of the field, not a per-project oversight.","The modality distribution (echo 26%, CMR and ECG 18% each) points to where annotation bottlenecks and clinical validation efforts will concentrate.","The review's stated gaps, such as vision transformers, synthetic data, and interdisciplinary collaboration, define concrete next projects for the community."],"supporting_citations":[{"why":"Supplies the largest non-clinical EHR dataset (567,498 pediatric patients) and the hierarchical LR/NLP baseline reported in the hybrid-models section.","marker":"[38]"},{"why":"One of the most-cited CHD-ML studies; in Table 6 the marker is reused for two different 2023 entries, anchoring the prenatal echo classification row.","marker":"[40]"},{"why":"Defines the CirCor DigiScope dataset used for heart-murmur classification and appears in Table 6 with CNN results.","marker":"[41]"},{"why":"Provides the ZCHSound pediatric heart-sound dataset and the RF/KNN binary and multi-class F1 results that Table 6 reports.","marker":"[55]"},{"why":"Supplies the DICOM X-ray CHD dataset and a ResNet18 classifier with reported accuracy and ROC, anchoring the X-ray modality row.","marker":"[65]"},{"why":"The multi-view multi-modal TTE study (1,932 children) that supports the claim that echo is the gold standard and the most-used modality.","marker":"[79]"},{"why":"Introduces the CHDdECG dataset (65,869 pediatric ECGs) and the CNN/human-concept method, backing the ECG modality row.","marker":"[81]"},{"why":"The Shanxi Province maternal-risk dataset and its SVM/RF/LR comparison; Section 5.1 misattributes this dataset to [45], which is why extraction accuracy matters.","marker":"[83]"},{"why":"The ImageCHD 3D CT dataset (110 patients, 16 CHD variants) anchoring the CT modality row.","marker":"[98]"},{"why":"The GAN-generated synthetic CMR study with a reported Dice of 1, illustrating the synthetic-data approach and a suspect table entry.","marker":"[118]"}],"fun_headline_variants":["74-study review maps ML landscape in congenital heart disease","Deep learning dominates: 60 of 74 CHD-ML studies","Survey catalogs 22 CHD datasets, only half public","CHD machine learning: small, recent, fragmented field"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that Tables 5 and 6 faithfully transcribe each of the 74 studies' dataset, method, and metrics, and that the Scopus-based screening captured the relevant literature; internal evidence, including one reference used for two different works, a perfect Dice score of 1, and a dataset attributed to the wrong source paper, shows that premise is not fully satisfied.","fun_headline_variants_meta":{"raw":{"variants":["74-study review maps ML landscape in congenital heart disease","Deep learning dominates: 60 of 74 CHD-ML studies","Survey catalogs 22 CHD datasets, only half public","CHD machine learning: small, recent, fragmented field"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000197,"raw_usage":{"total_tokens":1299,"prompt_tokens":816,"completion_tokens":483,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":432,"completion_tokens_details":{"reasoning_tokens":413}},"tokens_in":432,"tokens_out":483,"duration_ms":5189,"temperature":1.0,"reasoning_tokens":413,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T21:30:48.020229+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take any ten rows of Table 6, open the cited primary papers, and check whether the dataset name, data type, method, and metric values match. If more than two of the ten rows misreport the source, the review's per-study performance comparisons and its trend claims lose their evidentiary basis.","supporting_citations":[{"cited_title":"”Congenital heart disease detection by pediatric electrocardiogram based deep learning integrated with human concepts.” Nature Communications 15.1 (2024): 976","cited_arxiv_id":null,"evidence_quote":"Introduces the CHDdECG dataset (65,869 pediatric ECGs) and the CNN/human-concept method, backing the ECG modality row."},{"cited_title":"”Predicting congenital heart defects: A comparison of three data mining methods.” PloS one 12.5 (2017): e0177811","cited_arxiv_id":null,"evidence_quote":"The Shanxi Province maternal-risk dataset and its SVM/RF/LR comparison; Section 5.1 misattributes this dataset to [45], which is why extraction accuracy matters."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The ImageCHD 3D CT dataset (110 patients, 16 CHD variants) anchoring the CT modality row."},{"cited_title":"& Orwat, S","cited_arxiv_id":null,"evidence_quote":"The GAN-generated synthetic CMR study with a reported Dice of 1, illustrating the synthetic-data approach and a suspect table entry."}],"review_version":1}