{"id":"e165b449-aaef-407f-9e4b-840dc92bc08c","arxiv_id":"2411.18451","paper_version":1,"verdict":"UNVERDICTED","confidence":"HIGH","novelty_score":0.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A review of ten published MI detection methods for wearables, listing their reported accuracies and hardware specs without adding new results.","lead":"This paper reviews ten prior studies on detecting heart attacks (myocardial infarction) from wearable ECG devices. It summarizes their reported accuracies and hardware designs, but does not present any new experiments or data.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The accuracy values in Table I are not comparable across studies because class definitions, datasets, and metrics differ; the review's 'promising candidates' claim relies on this invalid comparison.","rationale":"The reader's weakest assumption correctly identified the comparability of accuracy metrics as the critical unsupported premise. I agree with that assessment and find no additional load-bearing concern beyond it. The central claim is a qualitative statement that the reviewed methods are 'promising candidates' for future wearable integration. Even if the accuracy figures are not directly comparable, each study does report some positive result on its own terms, so the qualitative 'promising' claim is not false; it is simply not quantitatively well-grounded. The paper is a narrative review with no new results, so verifying its synthesis is inherently difficult. Thus the appropriate verdict remains UNVERDICTED. A concrete test would be to audit the source studies and re-table the results with full task specifications; this would either expose the incomparability or, if the tasks happen to align, support the comparison. I therefore do not change the reader's verdict.","tokens_in":4635,"tokens_out":7686,"duration_ms":65686,"concrete_test":"Audit the primary sources for all entries in Table I and construct a normalized comparison table with columns: reference, dataset (name, number of subjects/beats), class definition (binary vs. multi-class, number of classes), evaluation protocol (beat- vs. patient-level split, cross-validation scheme), and the exact metric reported (accuracy vs. detection rate vs. sensitivity/specificity). Then re-examine whether any row in Table I is directly comparable to any other. If the tasks and metrics differ across rows, then the review's comparative statements ('outperformed previous works', 'robust performance metrics') are unsupported and the table should be replaced with a qualitative summary or explicitly caveated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim in Section IV that the methods in [1,3,8,10] are 'promising candidates' for wearable MI detection is supported primarily by the accuracy figures in Table I. That table presents a single accuracy per study without noting that the underlying tasks are not commensurate. For example, [5] reports 83.26% for a hierarchical classifier distinguishing five beat classes (normal, MI, bundle branch block, premature ventricular contraction, other), [6] reports 99.72% for an MLP on derived vectorcardiography, and [8] reports a 73% 'detection rate' with a 5% false alarm rate on the EDB database, which is not a multi-class accuracy at all. The evaluation protocols also differ (10-fold in [2], 5-fold in [7], K-fold in [9]) and the datasets are not shared. Section IV nevertheless uses these numbers to assert that [1] 'outperformed previous works' and that the listed methods have 'robust performance metrics.' If the accuracies are not measured on a common task, the table cannot support any comparative or integrative conclusion about which methods are promising.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper is a literature review of myocardial infarction (MI) detection and classification methods intended for wearable devices. It surveys preprocessing and feature extraction techniques (morphological filtering, wavelet decomposition, Pan-Tompkins, Hilbert curve mapping), classification algorithms (SVM, CNN/BCNN, MLP, decision trees, dendritic networks), and hardware implementations (SmartCardia, ARM microcontrollers, FPGA and ASIC designs). It presents a comparative table of seven studies with reported accuracies and concludes that the methods in references [1, 3, 8, 10] are particularly promising for wearable integration.","tokens_in":4817,"tokens_out":3342,"duration_ms":30019,"significance":"A well-executed review in this area would be valuable given the rapid growth of wearable ECG monitoring and the need to translate algorithmic advances into practical, energy-constrained devices. The paper does assemble a relevant set of primary works and draws attention to hardware/energy considerations that are often underemphasized. However, the central claim that the surveyed methods are 'promising candidates' is supported mainly by a cross-study accuracy comparison in Table I that is not statistically or methodologically valid. If the paper were revised to present the surveyed studies descriptively rather than comparatively, and to temper its conclusions accordingly, it could serve as a useful starting point for researchers. As it stands, the integrative conclusion is not supported by the evidence presented.","major_comments":[{"comment":"The accuracy values in Table I are not directly comparable, and the comparative conclusions in Section IV rest on this invalid comparison. For example, reference [5] reports 83.26% for a hierarchical classifier distinguishing five beat classes, reference [6] reports 99.72% for an MLP on derived vectorcardiography, and reference [8] reports a 73% detection rate with a 5% false alarm rate on the EDB database, which is not a multi-class accuracy at all. Evaluation protocols also differ (10-fold in [2], 5-fold in [7], K-fold in [9]), and the datasets and preprocessing pipelines differ across all rows. Without a common evaluation task, Table I cannot support the Section IV statement that the methods in [1, 3, 8, 10] 'showcase robust performance metrics encompassing accuracy, sensitivity, and specificity,' nor the claim that [1] 'outperformed previous works.' The table should be restructured to include a column for dataset, classification task, metric definition, and validation method, and the comparative and integrative claims in Section IV should be softened or removed.","section":"Table I and Section IV"},{"comment":"The statement that the method in [1] 'outperformed previous works' is not supported by any statistical or experimental comparison reported in the manuscript. The 90% accuracy is reported from a single study with a specific hardware target and energy constraints; no previous work is compared under the same conditions. This comparative claim should be replaced with a description of what the original study reported, or a proper within-study comparison should be cited.","section":"Section IV"},{"comment":"Table I lists seven rows but cites only references [1], [2], [3], [5], [6], [7], and [8]. References [9] and [10] are discussed at length in Sections II and III but are absent from the table, and reference [4] is never cited in the body of the paper. This inconsistency undermines the claim that the table provides a 'comparison of different classification methods.' The table should include all methods described in the text (with their appropriate metrics) or explicitly justify the exclusions, and reference [4] should either be cited in the text or removed from the reference list.","section":"Table I and Reference List"},{"comment":"The conclusion that the methods in [1, 3, 8, 10] exhibit 'robust performance metrics encompassing accuracy, sensitivity, and specificity' is inaccurate for reference [8], which reports only a detection rate and a false alarm rate. Sensitivity and specificity for [10] are given in the text but not in Table I. This mismatch between the table, the narrative, and the cited metrics needs to be corrected so that the claims match the evidence.","section":"Section IV"}],"minor_comments":[{"comment":"The phrase 'passings' in Section I should be replaced with 'deaths', and 'ministrations' is an unusual word choice that likely should be 'measurements' or 'operations'.","section":"Abstract and Section I"},{"comment":"There are several typographical errors in Table I: 'Hierachiacal classifiation' should read 'Hierarchical classification', 'Classififer' should read 'Classifier', and '98. 95%' should read '98.95%' without a space.","section":"Table I"},{"comment":"The sentence 'achieving an the average sensitivity of 86.18%' contains a grammatical error and should be corrected to 'achieving an average sensitivity of 86.18%'.","section":"Section IV"},{"comment":"Reference [6] is incomplete: it lacks the conference or journal name, volume, pages, or DOI. Please provide the full bibliographic details.","section":"References"},{"comment":"The phrase 'K-fold' in Section IV and elsewhere should be specified as 'K-fold cross-validation' for clarity, and the value of K used in reference [9] should be reported if known.","section":"Section II-B"}],"recommendation":"major_revision","confidential_remarks":"The paper is a broad survey with a useful collection of references, but the core comparative analysis is not methodologically sound, and the reference list is not fully integrated into the text. The errors are fixable within the scope of a major revision, so I do not recommend rejection. However, the authors should be strongly encouraged to restructure Table I with a common set of evaluation descriptors and to rewrite the conclusions to avoid unsupported comparative claims. The missing citation of reference [4] suggests the reference list was assembled with insufficient care."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The one thing to know: this is a student-level literature review with a single table that appears to compare ten MI-detection studies, but the table omits three cited references, and the accuracies come from incompatible tasks, so the table can't support the paper's main claim that the listed methods are 'promising candidates' for wearable use.\n\nWhat the paper does well: it's a compact orientation to a genuine application niche. The authors read and summarized ten real papers, and the structure—preprocessing, classification models, hardware, evaluation—is sensible. If you're new to wearable ECG and want a quick list of the main approaches (SVM cascade, binary CNN, MLP on vectorcardiography, VLSI classifier), this gives you that.\n\nThe soft spots are real. Table I lists seven rows but the text cites ten references; [4], [9], and [10] are missing from the table. More importantly, the accuracy column mixes non-comparable numbers: [5] reports a multi-class accuracy for five beat types, [6] is a binary MLP accuracy on derived VCG, [8] reports a detection rate with a false alarm rate—not an accuracy at all. The cross-validation schemes differ (10-fold, 5-fold, K-fold), and the datasets aren't shared. Section IV then uses this table to say that [1] 'outperformed previous works' and that [1,3,8,10] show 'robust performance metrics.' That conclusion doesn't follow from the table. There are also numerous editorial slips ('Hierachiacal', 'the an average'), and the references themselves are inconsistently formatted—[6] has no venue details.\n\nNone of this makes the underlying sources bad; it makes the review's synthesis invalid. A reader cannot use this table to decide which method is best or 'promising' without going back to the original papers and normalizing the experimental setups.\n\nWho it's for: a raw beginner looking for a pointer to the literature, and even then only with heavy caveats. I wouldn't use it as a citable source, and I wouldn't send it to peer review in its current state. My recommendation: desk reject, with an encouragement to the authors to either turn it into a proper systematic review (with inclusion criteria, a PRISMA-style flow, and a normalized comparison) or narrow it to a short narrative review that avoids quantitative comparisons across incompatible studies. It's the kind of paper that could be made useful with real revision, but right now the load-bearing comparison is broken.","headline":"A weak review whose only comparative table is invalid because the studies measure different tasks with different metrics; useful only as a starting pointer, not as evidence.","tokens_in":5314,"tokens_out":2919,"would_cite":false,"duration_ms":26679,"reading_group":"no","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Wearable ECG detectors reach 99.72% accuracy, review finds","keywords":["myocardial infarction","ECG classification","wearable devices","machine learning","deep learning","convolutional neural networks","VLSI","review"],"falsifier":"Run the methods cited in Table I on a single shared ECG dataset with the same folds, labels, and preprocessing, and compare their accuracies and energy use; if the rankings or the 73-99.72% range do not reproduce, the review's conclusion that these methods are equally ready for wearables would not be supported.","tokens_in":4471,"feed_emoji":"💓","tokens_out":5542,"duration_ms":46300,"temperature":0.7,"pith_summary":"This review asks whether machine-learning and signal-processing techniques for detecting myocardial infarction (MI) from ECG signals have matured enough to run on wearable devices. It brings together studies spanning traditional morphological filtering and wavelet decomposition, support vector machines, convolutional and binary neural networks, multilayer perceptrons, a convolutional dendrite model, and VLSI implementations, reporting accuracies from 73% to 99.72%. The paper's central contention is that these methods, especially low-power designs, are candidates for integration into wearable healthcare devices, enabling continuous monitoring and earlier intervention. A sympathetic reader would care because if that contention holds, early MI detection no longer requires hospital-grade equipment.","feed_headline":"Wearable ECG detectors reach 99.72% accuracy, review finds","feed_subtitle":"Survey of MI classifiers shows 73-99.72% accuracy, with low-power chips enabling continuous monitoring.","key_machinery":"The load-bearing object is Table I, the review's comparison table of seven classification methods with their reported accuracies, together with the hardware section's energy and area figures. Table I is what lets the authors claim maturity and readiness: it aggregates results from independent studies that use different feature sets (morphological features, wavelets, entropies, Hilbert-curve images), different classifiers (SVM, BCNN, MLP, CDD, statistical), and different low-power targets (STM32L151, EFM32 Leopard Gecko, an FPGA, and an ASIC in 180 nm). The review's argument runs through this table: because such a spread of approaches reaches high accuracy on wearable-grade hardware, integration into wearables is a realistic next step.","core_discovery":"The paper's central claim is that the reviewed MI classification techniques form a mature toolbox for wearable deployment: a two-level SVM classifier (90%), binary CNNs (91.22% and 90.29%), a hierarchical classifier (83.26%), an MLP on derived vectorcardiography (99.72%), a convolutional dendrite net (98.95%), and a statistical method (73%) all report usable accuracy, with some designs cutting energy by a factor of three and a VLSI classifier reaching 86.18% sensitivity and 96.5% specificity at 5.12 microwatts. The review argues that these results show real-time, energy-aware MI detection on low-power microcontrollers is achievable, and that wearable devices could provide early diagnosis that prevents irreversible heart damage. The claim is presented through a comparative synthesis rather than a new experiment.","pith_inferences":["Because the reviewed studies use different datasets and validation schemes (10-fold versus 5-fold, different class labels), the ordering of accuracies in Table I is likely not stable; a head-to-head benchmark on a shared database would give a more dependable ranking.","The success of Hilbert-curve encoding in the CDD-net study suggests a natural next step: combining image-based ECG representations with binary neural networks to cut memory further on wearables.","If the reported energy figures scale as claimed, the same low-power designs could be extended to detect other arrhythmias, since the feature-extraction and classification pipeline is not MI-specific."],"forward_implications":["Continuous ECG monitoring on a wearable could catch MI early enough to summon care before irreversible muscle damage occurs.","Low-power designs in the reviewed studies, including a 5.12 microwatt VLSI classifier, suggest battery-powered devices can run MI detection without frequent recharging.","Energy-aware architectures such as binary CNNs and hierarchical classifiers make real-time on-device classification feasible on 32-bit microcontrollers.","The spread of reported accuracies (73% to 99.72%) indicates that even simple statistical and signal-processing methods may be usable when computational budget is tight."],"supporting_citations":[{"why":"Two-level SVM classifier at 90% accuracy; the review's first demonstration of real-time MI detection on a wearable.","marker":"[1]"},{"why":"Binary CNN with 91.22% accuracy and energy-aware neural architecture search; supplies the deep-learning-on-low-power result.","marker":"[2]"},{"why":"Binary CNN at 90.29% accuracy with sensitivity and specificity across 10 folds; second BCNN data point.","marker":"[3]"},{"why":"Hierarchical classifier at 83.26% using wavelets, fractal dimension and entropies; broadens feature set beyond standard parameters.","marker":"[5]"},{"why":"MLP classifier at 99.72% based on derived vectorcardiography; the highest reported accuracy in Table I.","marker":"[6]"},{"why":"Convolutional dendrite net at 98.95% using Hilbert-curve ECG image encoding; the image-transformation approach.","marker":"[7]"},{"why":"Statistical method with 73% detection rate and 5% false alarms on the EDB database; the baseline traditional approach.","marker":"[8]"},{"why":"VLSI stage classifier with 86.18% sensitivity and 96.5% specificity at 5.12 microwatts; the hardware-implementation result.","marker":"[10]"}],"fun_headline_variants":["Wearable ECG hits 99.72% accuracy for MI detection","Low-power chips enable real-time heart attack detection","Review: Wearable MI detection reaches 99.72% accuracy","Continuous monitoring on wearables catches heart attacks early"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The review assumes that accuracy numbers reported by different studies are directly comparable, even though the studies use different datasets, preprocessing pipelines, class definitions, and cross-validation schemes.","fun_headline_variants_meta":{"raw":{"variants":["Wearable ECG hits 99.72% accuracy for MI detection","Low-power chips enable real-time heart attack detection","Review: Wearable MI detection reaches 99.72% accuracy","Continuous monitoring on wearables catches heart attacks early"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000396,"raw_usage":{"total_tokens":2029,"prompt_tokens":853,"completion_tokens":1176,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":469,"completion_tokens_details":{"reasoning_tokens":1108}},"tokens_in":469,"tokens_out":1176,"duration_ms":9389,"temperature":1.0,"reasoning_tokens":1108,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T11:11:02.991927+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the methods cited in Table I on a single shared ECG dataset with the same folds, labels, and preprocessing, and compare their accuracies and energy use; if the rankings or the 73-99.72% range do not reproduce, the review's conclusion that these methods are equally ready for wearables would not be supported.","supporting_citations":[{"cited_title":"Real-time classification technique for early detection and prevention of myocardial infarction on wearable devices,","cited_arxiv_id":null,"evidence_quote":"Two-level SVM classifier at 90% accuracy; the review's first demonstration of real-time MI detection on a wearable."},{"cited_title":"Energy-Aware Design Methodology for Myocardial Infarction Detection on Low-Power Wearable Devices,","cited_arxiv_id":null,"evidence_quote":"Binary CNN with 91.22% accuracy and energy-aware neural architecture search; supplies the deep-learning-on-low-power result."},{"cited_title":"Real-Time Event-Driven Classification Technique for Early Detection and Prevention of Myocardial Infarction on Wearable Systems,","cited_arxiv_id":null,"evidence_quote":"Hierarchical classifier at 83.26% using wavelets, fractal dimension and entropies; broadens feature set beyond standard parameters."},{"cited_title":"A Hybrid System for Myocardial Infarction Classification with Derived Vectorcardiography","cited_arxiv_id":null,"evidence_quote":"MLP classifier at 99.72% based on derived vectorcardiography; the highest reported accuracy in Table I."},{"cited_title":"Early detection of Myocardial Infarction using WBAN","cited_arxiv_id":null,"evidence_quote":"Statistical method with 73% detection rate and 5% false alarms on the EDB database; the baseline traditional approach."}],"review_version":1}