{"id":"bd6f65d4-ba9d-493d-8a49-6d4d330c3178","arxiv_id":"2601.00014","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"DeepHHF, trained on day-long single-lead Holter ECGs from 40,174 patients, predicted incident heart failure within five years with AUROC 0.80 and external AUROC 0.81.","lead":"Deep learning on full 24-hour single-lead Holter ECGs predicted 5-year heart failure diagnosis with AUROC 0.80 on held-out data, beating 30-second ECG windows and the PCP-HF clinical score. The work is an opportunistic-screening proof: a routine Holter could flag high-risk patients for preventive follow-up.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'incident HF' label may include prevalent but undocumented HF; the paper's own Discussion admits this could inflate the AUROC, so DeepHHF's 0.80 may partly reflect detection of existing disease rather than 5-year risk prediction.","rationale":"The reader's CONDITIONAL verdict is appropriate. I agree with the reader's weakest assumption: the acknowledged prevalent-HF mechanism is the single most load-bearing threat to the central claim. If the proposed test lands, the 'prediction' framing fails; if not, the claim is strengthened. Secondary concerns (PCP-HF subset mismatch, small external validation, code availability) are real but less central to the core claim of learning incident HF risk from 24-hour ECG. I therefore recommend no change to the reader's verdict.","tokens_in":22056,"tokens_out":6820,"duration_ms":73580,"concrete_test":"Using the Leumit EMR, stratify the 4,461-exam test set by the presence of pre-Holter markers of prevalent HF: any dispensed loop diuretic, MRA, or SGLT2i, or any echocardiogram within 12 months before the Holter. Recompute DeepHHF's AUROC (with bootstrap CI) in the stratum lacking all such markers. If the clean-stratum AUROC is materially lower than the reported 0.80 (e.g., ≤0.75), or loses significance against PCP-HF on the same stratum, the headline is inflated by prevalent-HF detection rather than incident prediction.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing assumption is that the endpoint is truly incident HF. The Methods define the endpoint as the first documented HF ICD-9 code and exclude only recordings with a HF diagnosis documented before the Holter. The Discussion explicitly concedes: \"patients newly diagnosed with HF may be recorded in the EMR several months after their initial diagnosis. Consequently, undiagnosed prevalent HF cases with active symptoms and treatments at the time of Holter recording could influence the ECG data and potentially inflate model performance.\" Because Holter monitoring is performed for cardiac indications (syncope, palpitations, arrhythmia follow-up), a nontrivial fraction of 'negative-at-baseline' patients may already have HF physiology. If so, the 0.80 AUROC partially measures detection of prevalent disease, not 5-year prediction. The label-verification analyses (medications, echocardiography) confirm that documented HF diagnoses are real, but they do not establish that HF was absent at the time of the recording. The performance decline from 0.81 (0-2 y before diagnosis) to 0.77 (4-5 y) is consistent with progressive disease but also with prevalent cases whose diagnosis was delayed by years; the paper's 'consistency within first two years' argument does not rule out inflation.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents DeepHHF, a deep-learning model that takes a full 24-hour single-lead Holter ECG recording as input and outputs a five-year heart-failure risk score. The model is trained and evaluated on the Technion-Leumit Holter ECG (TLHE) dataset, a large retrospective Israeli HMO cohort, with a temporally held-out test set (January–April 2018), patient-level split, and an external zero-shot cohort from Rambam. The authors report AUROC 0.80 on the internal test set, outperforming a 30-second-window encoder (0.77) and the PCP-HF clinical score (0.74), with AUROC 0.81 on the external cohort. Additional analyses include subgroup performance by time-to-diagnosis, Kaplan-Meier survival curves, risk-group stratification, and gradient-attention-rollout explainability. The paper's central claim is that five-year incident heart-failure risk is learnable from day-long single-lead ECG, with the full 24-hour recording adding value over short segments.","tokens_in":22266,"tokens_out":4426,"duration_ms":47630,"significance":"If the central claim holds, this is a substantial contribution: it is, to my knowledge, the first demonstration that raw 24-hour single-lead Holter ECG predicts incident heart failure at five years, and it is supported by several strong methodological components — a large real-world dataset, a temporally held-out test set with complete follow-up, patient-level stratification, bootstrapped confidence intervals, zero-shot external validation, and extensive label-verification analyses. The publicly released model and reproducible architecture details are also strengths. However, the manuscript has three load-bearing concerns that must be addressed before the claim can be accepted at face value: the endpoint may include prevalent but undocumented heart failure, the PCP-HF comparison is not on the same test subset, and the training-set labeling may include censored non-HF examples. These are fixable with additional analyses or re-analysis, and the underlying dataset and model are sufficiently valuable that the paper merits major revision rather than rejection.","major_comments":[{"comment":"The paper's own Discussion concedes that 'undiagnosed prevalent HF cases with active symptoms and treatments at the time of Holter recording could influence the ECG data and potentially inflate model performance.' This is load-bearing for the claim of incident-risk prediction. The exclusion criteria only remove patients with a pre-Holter documented HF diagnosis; label-verification analyses (prescriptions, echocardiography) establish that documented diagnoses are real but cannot establish absence of HF at Holter time. The argument that AUROC remains consistent within 0–2 years does not rule out inflation, because diagnoses delayed by up to two years would still fall in that interval. Please provide a sensitivity analysis that excludes patients with HF medications, loop diuretics, or abnormal echocardiography before the Holter, or otherwise demonstrate that the 0.80 AUROC is not substantia","section":"Methods — Class definition; Discussion paragraph"},{"comment":"The PCP-HF AUROC is computed on the 1,917 test-set examinations for which all PCP-HF covariates were available, while the DeepHHF AUROC is computed on the full 4,461-examination test set. These are different patient subsets, so the reported significant difference (p<0.05) is confounded by case mix. The manuscript should either re-evaluate DeepHHF on the same 1,917-exam subset or evaluate PCP-HF on the full test set using imputation (e.g., multiple imputation or mean/median imputation with sensitivity analysis). Without this, the claim that DeepHHF outperforms PCP-HF is not supported by the displayed analysis.","section":"Figure 3b; Methods — PCP-HF score computation"},{"comment":"The text states that the test set (January–April 2018) ensures complete 5-year follow-up, but the training and validation sets appear to include Holter recordings from 2010 through June 2023, while EMR data end in April 2024. If any training/validation recording was made after April 2019, a 5-year follow-up window is not available, and labeling such recordings as non-HF because no diagnosis appears before the data cutoff creates censored negative labels. Please clarify whether training/validation recordings were restricted to those with at least five years of observable follow-up; if not, the model training used systematically mislabeled negatives, and the analysis should be repeated with a training set restricted to complete observations or with a time-to-event formulation.","section":"Methods — Dataset split and preprocessing; Results — Study cohort"}],"minor_comments":[{"comment":"The keyword 'circardian' appears to be a typo for 'circadian.'","section":"Keywords; Discussion"},{"comment":"The text refers to 'Figure 1a' when discussing subgroup performance by time-to-diagnosis; the correct reference appears to be Figure 4a. Please check all cross-references.","section":"Discussion, paragraph 8"},{"comment":"The 95% CI for AUROC is obtained by bootstrapping 250 positive and 250 negative examples per iteration, which is not the usual full test-set bootstrap. Please report CIs at the full test-set level and, if appropriate, use DeLong's test or paired bootstrap for classifier comparison.","section":"Performance measures and statistical analysis"},{"comment":"The external cohort AUROC of 0.81 is reported without a confidence interval. Given only 29 positive cases, the CI is likely wide; please include it or at least an exact binomial CI.","section":"External Validation"},{"comment":"The sentence 'The optimization attempted to maximize the validation set AUROC score' is clear, but the hyperparameter list is long and the final selected values are not reported. Providing the final hyperparameters in a table or Extended Data would aid reproducibility.","section":"Methods — Deep learning model"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is technically substantial and the dataset alone is a major asset. My recommendation of major revision is driven by the three load-bearing issues above, all of which seem addressable with additional analyses rather than being fatal. The endpoint-prevalence concern is the deepest: the authors themselves acknowledge it, and their response in the discussion (performance consistency in years 0–2) is insufficient to rule out delayed-diagnosis inflation. The PCP-HF comparison is straightforward to fix. The training-set censoring issue, if confirmed, would require re-running the model with a properly censored training set; this is more work but within scope. I would not recommend rejection unless a re-analysis shows the AUROC collapses. I also suggest the editor ask the authors to make the final hyperparameters and a cohort flow diagram with exact recording counts by year available."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First the bottom line: this is the first paper I know that trains on full 24-hour single-lead Holter ECG for 5-year heart failure risk, and it does it on a large new dataset with a mostly careful evaluation. The central AUROC of 0.80 on a temporally held-out test set, with patient-level splitting, bootstrapped CIs, and a zero-shot external cohort, is credible. The cleanest result is the internal comparison: 24-hour input beats 30-second windows (0.80 vs 0.77), which supports the premise that longer recordings add something.\n\nThe dataset itself is a real contribution. TLHE with ~57k recordings from ~40k patients and 20 years of EMR follow-up will be useful to the community. The label verification analyses—medication prescription patterns, echocardiography timing, repeated diagnosis checks—are more thorough than most EHR-based ML papers.\n\nNow the soft spots, in proportion. The most load-bearing one is the endpoint. The label is the first documented HF ICD-9 code, and only recordings with a documented prior HF diagnosis are excluded. The paper itself concedes in the Discussion that 'undiagnosed prevalent HF cases with active symptoms and treatments at the time of Holter recording could influence the ECG data and potentially inflate model performance.' That is not a hypothetical: these are Holter recordings taken because of syncope, palpitations, or known arrhythmia, so a nontrivial fraction of 'negative at baseline' patients may already have HF physiology. The performance gradient—0.81 for events within 2 years, 0.77 for 4–5 years—is consistent with delayed diagnosis of prevalent disease. The paper's 'consistency within the first two years' argument does not actually rule out inflation. So I would read the 0.80 as partly detection of existing disease, not purely 5-year prediction. That said, this is not fatal: a screen that detects subclinical HF in a Holter population still has clinical value. But the framing should be adjusted.\n\nOther issues are smaller. The PCP-HF comparison is run on 1,917 test exams for which all PCP-HF variables were available, and the paper does not report DeepHHF's AUROC on that same subset, so the head-to-head is not fully matched. External validation has only 29 events and no confidence intervals. And the trained model is promised 'at URL upon publication,' which is not actually available now despite the Code Availability section implying it is.\n\nOverall: this deserves a serious referee. The right reviewers should push for a sensitivity analysis excluding or re-labeling events in the first 1–2 years, a matched PCP-HF comparison, CIs for the external cohort, and a real code release. If those are addressed, it becomes a solid paper.\n\nWho is it for? People working on AI-ECG risk prediction, wearable ECG, and anyone building EHR-based labels for incident disease. I'd bring it to a reading group as a case study in label validity, and I'd cite it if I worked in this area—with a caveat.\n\nRecommendation: send to peer review. Not desk reject. But expect major revision.","headline":"A strong first demonstration of 24-hour Holter ECG for HF risk, but the incident-HF label likely captures some prevalent disease—the paper admits as much—so the 0.80 AUROC should be read as partly detection, not pure prediction.","tokens_in":22893,"tokens_out":3166,"would_cite":true,"duration_ms":28632,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A single 24-hour ECG recording can predict a patient's risk of being diagnosed with heart failure within five years, according to a deep learning model that reads the full day-long signal.","keywords":["heart failure prediction","24-hour Holter ECG","deep learning","single-lead ECG","risk stratification","explainability","circadian variation","incident heart failure"],"falsifier":"Apply DeepHHF to a cohort in which heart failure status at the time of the Holter is independently adjudicated by symptoms, BNP level, and echocardiography. If the AUROC stays near 0.80 among patients confirmed free of heart failure at baseline, prediction is validated; if performance mostly disappears when patients with baseline HF medications or echo evidence are excluded, the result reflects detection bias. Additionally, compute AUROC only for recordings 4–5 years before the documented diagnosis: if it falls to near 0.5, the five-year-horizon claim collapses.","tokens_in":21866,"feed_emoji":"🫀","tokens_out":3038,"duration_ms":33930,"temperature":0.7,"pith_summary":"The paper claims that a deep learning model processing an entire 24-hour single-lead Holter ECG can predict whether a patient will be diagnosed with heart failure within five years, achieving an AUROC of 0.80. The day-long recording matters: a model restricted to 30-second clips scored 0.77, and the authors argue the gap shows that paroxysmal arrhythmias and circadian patterns carry predictive information. If correct, routine Holter exams done for other indications could opportunistically flag high-risk patients for preventive follow-up, such as BNP testing or echocardiography, without extra cost or invasive procedures. The model also stratifies prognosis: patients flagged as high-risk have fourfold higher odds of death and twofold higher odds of hospitalization or death than lower-risk patients.","feed_headline":"24-hour ECG predicts heart failure five years out","feed_subtitle":"Single-lead Holter deep learning scores 0.80 AUROC, beating 30-second clips and a clinical score.","key_machinery":"A two-stage deep learning architecture: first, a convolutional encoder (based on EnCodec blocks) is trained on randomly sampled 30-second windows to extract compact latent features; second, the frozen encoder converts 720 fixed-interval windows spanning the full 24-hour recording into a sequence, which a transformer sequential head integrates into a single HF risk score. Gradient attention rollout then traces which parts of the recording drove the prediction. This machinery enables the model to use the entire day-long signal rather than a short snapshot, and the transformer's sequential integration is what the authors credit for capturing circadian and paroxysmal information.","core_discovery":"DeepHHF, trained and validated on the Technion-Leumit Holter ECG dataset (57,575 recordings from 40,174 patients), learns to predict incident heart failure within five years from raw 24-hour single-lead ECG. The full-recording model reaches AUROC 0.80 on the held-out test set, outperforming both a 30-second-window encoder (0.77) and the PCP-HF clinical score (0.74), and achieves AUROC 0.81 in a zero-shot external cohort. Explainability via gradient attention rollout shows the model concentrates on daytime hours and on ectopic beats—premature ventricular contractions and supraventricular ectopy—consistent with known arrhythmia burdens preceding heart failure. The authors conclude that day-lon","pith_inferences":["A testable extension of this work is whether a compact summary of a Holter recording (for example, hourly ectopic beat counts and heart rate variability) can approach the same performance as the full signal, which would make the approach feasible on wearable patches with limited storage.","If undiagnosed prevalent heart failure partly drives the 0.80 AUROC, then a truly asymptomatic screening population would likely show lower performance; the clinical value depends on whether the model detects subclinical disease before symptoms, which this retrospective design cannot fully separate.","The daytime attention peak (8 AM–3 PM) could be a circadian signature; a direct experiment would be to retrain the model on nighttime-only segments and measure the AUROC drop, isolating the timing contribution from overall arrhythmia burden.","Because the endpoint relies on ICD-9 codes, linking the model to echocardiogram-confirmed heart failure with ejection fraction subtypes would test whether DeepHHF generalizes across HFpEF and HFrEF, which the authors aim to cover."],"forward_implications":["Routine 24-hour Holter recordings, ordered for palpitations or syncope, can be repurposed to estimate five-year heart failure risk without additional tests.","DeepHHF outperforms the guideline-recommended PCP-HF clinical score on the test set, and adding clinical variables pushes AUROC to 0.82.","High-risk patients identified by the model have fourfold odds of all-cause mortality and twofold odds of hospitalization or death versus low/moderate-risk patients.","Zero-shot performance on an external cohort (AUROC 0.81) suggests the model transfers beyond the training health system.","Using DeepHHF to select patients for preventive interventions reduces the number needed to screen to prevent one major cardiovascular hospitalization, from 61 in the overall test set to 21 in the high-risk subgroup."],"fun_headline_variants":["Five-year heart failure risk from a day-long ECG","AI reads 24-hour ECG to foresee heart failure","Full-day ECG beats short clips for heart failure prediction","Explainable AI spots arrhythmia clues in day-long ECG","A single Holter day predicts heart failure risk"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing assumption is that the first documented ICD-9 heart failure code in the electronic medical record within five years is a valid and timely proxy for incident heart failure; if many patients already have undocumented or prevalent heart failure at the time of the Holter recording, the model may be detecting existing disease rather than predicting future onset.","fun_headline_variants_meta":{"raw":{"variants":["Five-year heart failure risk from a day-long ECG","AI reads 24-hour ECG to foresee heart failure","Full-day ECG beats short clips for heart failure prediction","Explainable AI spots arrhythmia clues in day-long ECG","A single Holter day predicts heart failure risk"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000629,"raw_usage":{"total_tokens":2755,"prompt_tokens":766,"completion_tokens":1989,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":510,"completion_tokens_details":{"reasoning_tokens":1913}},"tokens_in":510,"tokens_out":1989,"duration_ms":14355,"temperature":1.0,"reasoning_tokens":1913,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T14:57:18.207525+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Apply DeepHHF to a cohort in which heart failure status at the time of the Holter is independently adjudicated by symptoms, BNP level, and echocardiography. If the AUROC stays near 0.80 among patients confirmed free of heart failure at baseline, prediction is validated; if performance mostly disappears when patients with baseline HF medications or echo evidence are excluded, the result reflects detection bias. Additionally, compute AUROC only for recordings 4–5 years before the documented diagnosis: if it falls to near 0.5, the five-year-horizon claim collapses.","supporting_citations":[],"review_version":1}