{"id":"58cdb59c-8f1d-4db9-92af-179cb4aebfda","arxiv_id":"2607.03009","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.5,"correctness_risk":"low","formal_verification":"none","parameter_count":4,"one_line_summary":"For Brugada syndrome detection, ECG foundation-model pre-training mainly stabilizes optimization rather than encoding transferable clinical knowledge, and fails to improve zero-shot cross-site generalization.","lead":"ECG foundation models help high-capacity nets train on tiny Brugada datasets but do not transfer clinically meaningful features better than compact models trained from scratch. Architecture and site match matter more than pre-training for this rare disease.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified that would overturn the central claim under the paper's own framing.","rationale":"The paper's strongest claim is an empirical negative result carefully bounded to Brugada on two cohorts. The experimental design (paired FS/LP/FT, patient-level pooling, FDR-controlled DeLong, fixed test sets under ablation, bidirectional zero-shot) isolates pre-training from architecture more cleanly than prior single-model rare-disease FM papers. The non-significant between-model gaps on full data and the non-replication of data-efficiency on HUCA are the load-bearing evidence; they survive the acknowledged limitations. The reader's weakest assumption is correctly identified but is not load-bearing for the claim as stated: even if Brugada were present in some pre-training corpora, the cross-site collapse and the architecture-driven ranking would still indicate that any such exposure did not produce transferable clinical features under the tested protocols. No hidden assumption in the statistical pipeline or the LLRD scheme appears capable of flipping the mechanical-vs-semantic interpretation. Therefore the CONDITIONAL verdict (reproducibility gated on private data + open generalization beyond Brugada) already reflects the right degree of caution; no adjustment is warranted.","tokens_in":36924,"tokens_out":568,"duration_ms":4919,"concrete_test":"Independently re-run the between-model DeLong comparisons of Table 4 (BrSwiss-100%) and the HUCA within-site table using only the three compact architectures that already converge from scratch (ECG-CPC, MERL-ResNet, ECGFounder); if the best FT still fails to significantly beat the best FS at FDR-adjusted p<0.05 on both cohorts, the mechanical-vs-semantic claim remains intact for the architectures that matter most.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (pre-training is mechanical rather than semantic for Brugada detection) is tightly scoped to the nine public FMs, two adult case-control cohorts, and the three adaptation strategies. The within-model vs between-model decomposition, the non-replication of the BrSwiss-3% data-efficiency gain on HUCA, and the near-chance zero-shot cross-site results jointly support that conclusion without an internal contradiction. The reader's weakest assumption (representativeness of Brugada + unverified absence from every pre-training corpus) is real but is already treated as a limitation by the authors (Methods 3.1, Discussion Limitations) and does not reverse the reported pattern on the data that were actually evaluated. Sampling-rate mismatch and private BrSwiss data affect absolute transfer numbers but do not create an alternative explanation that would make the observed pre-training gains semantic rather than optimization-related.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The manuscript systematically benchmarks nine publicly available ECG foundation models for Brugada syndrome detection on an internal Swiss cohort (BrSwiss; 294 patients, 87 cases) and an independent external Spanish cohort (HUCA; 363 patients, 76 cases). Each model is evaluated under from-scratch training, linear probing, and full fine-tuning, with matched-architecture supervised baselines, a 3% training-window ablation, and zero-shot cross-site transfer. The central claim is that, for this rare phenotype, FM pre-training is largely mechanical (stabilizing optimization of high-capacity models that fail to converge from scratch) rather than semantic (encoding transferable clinical knowledge): the best fine-tuned FM only marginally exceeds the strongest from-scratch baseline on full BrSwiss (ECG-CPC 0.962 vs 0.932, p=0.091), the BrSwiss-3% data-efficiency gain does not replicate on HUCA, and all strategies collapse toward chance under zero-shot cross-site transfer. Architectural suitability, not pre-training per se, is identified as the most consistent performance driver.","tokens_in":37208,"tokens_out":1620,"duration_ms":28690,"significance":"If the result holds under the paper’s stated scope, it is a timely and useful reality check for the ECG foundation-model literature, which has largely evaluated common, well-labeled conditions. The design strengths are real and should be credited: one-to-one matched architecture from-scratch baselines, patient-level AUC with bootstrap CIs, DeLong tests with FDR control across large comparison blocks, early stopping keyed to validation AUC to equalize effective training exposure under heterogeneous native input windows, and an independent external cohort released after all pre-training cutoffs. These choices make the within-model versus between-model decomposition and the non-replication of data-efficiency gains more persuasive than typical single-model transfer reports. The work does not claim a universal negative about FMs; it scopes a concrete failure mode (rare phenotype outside pre-training coverage, site shift) that the field needs to confront when arguing for clinical deployment of large ECG FMs.","major_comments":[{"comment":"Title, abstract, and conclusion frame the result as evidence about transfer to “rare cardiac diseases,” but all empirical support is for a single adult case–control phenotype (Brugada) on two specialized cohorts. Methods 3.1 and the Limitations section already note this; the load-bearing issue is that the abstract’s closing sentence and the paper title still invite a broader reading than the data justify. Either narrow the title/abstract claims to Brugada (or “a rare ECG phenotype”) or add a short, explicit boundary statement that generalization to other rare cardiovascular conditions is untested and may differ when the phenotype is better represented in pre-training corpora.","section":null},{"comment":"The cross-site claim (“pre-training does not confer cross-site robustness”) is central (Results §4.2 Cross-site generalization; Discussion theme 3; Figure 4; Tables 8–10). Limitations correctly notes the native-rate mismatch (BrSwiss 1000 Hz vs HUCA 100 Hz) and that up-sampling cannot restore content above the original Nyquist frequency. This is not a minor preprocessing detail: several FMs were pre-trained at 250–500 Hz, so HUCA inputs are spectrally truncated relative to their pre-training distribution. Without a controlled same-rate ablation (e.g., all pipelines forced to 100 Hz on both sites, or BrSwiss downsampled before transfer), it remains ambiguous how much of the near-chance collapse is population/protocol shift versus spectral domain shift. A same-rate control, or at minimum a quantitative sensitivity analysis, is needed before the “no robustness from pre-training” conclusion","section":null},{"comment":"The interpretive slogan “mechanical rather than semantic” (Abstract Conclusion; Discussion first theme; Conclusions) is stronger than the direct evidence. Within-model FT–FS gains for large transformers that fail from scratch, non-significant between-model gaps on full BrSwiss and HUCA, non-replication of BrSwiss-3% efficiency on HUCA, and chance-level zero-shot transfer jointly support that pre-training does not deliver a practical, transferable clinical advantage over the best architecture trained from scratch. They do not, by themselves, prove that pre-trained representations contain no Brugada-relevant structure—only that any such structure is insufficient, non-transferable across sites, or dominated by optimization effects under the tested adaptation protocols. Soften or operationalize the language (e.g., “pre-training primarily stabilizes optimization and does not yield transferabl","section":null}],"minor_comments":[{"comment":"Secondary metrics (Appendix Tables 11–16) are reported at a fixed threshold of 0.5. The Clinical utility paragraph correctly notes that Brugada screening needs sensitivity-oriented operating points; consider adding one sensitivity-at-fixed-specificity (or Youden/validation-tuned) column for the top models so readers can see clinical operating characteristics without re-thresholding themselves.","section":null},{"comment":"Patient-level aggregation uses mean probability pooling over segments (Appendix C). Brugada type-1 pattern can be intermittent (spontaneous vs drug/fever-induced). A brief sensitivity check with max pooling (or fraction of positive segments) would clarify whether relative FM vs FS rankings are robust to aggregation choice.","section":null},{"comment":"Table 1 and Methods 3.3 assert that BrSwiss was never in any pre-training corpus and that HUCA’s public release postdates all pre-training. For completeness, a short note on whether any of the public pre-training sources (MIMIC-IV-ECG, CODE, PTB-XL, etc.) are known to contain labeled or unlabeled Brugada-pattern ECGs would strengthen the “outside pre-training coverage” premise used in the Discussion.","section":null},{"comment":"Figure 2–4 error bars and asterisks are dense; ensure the caption explicitly states that asterisks are within-model (LP/FT vs same-architecture FS) and that between-model claims are only in the text. A small legend or footnote would reduce misreading.","section":null},{"comment":"Saliency appendix (Figures 5–7) usefully cautions that IG magnitude scales with logit scale and that maps are per-panel normalized. Consider moving one sentence of that caveat into the main-text Results or Discussion so readers who skip the appendix do not over-interpret visual saturation differences between FS and FT.","section":null},{"comment":"Minor typography: abstract and body occasionally drop spaces after periods or in compound terms (e.g., “Foundationmodels”, “rarediseases”); a pass for spacing and consistent hyphenation of “fine-tuning” / “from-scratch” would help.","section":null}],"recommendation":"minor_revision","confidential_remarks":"The empirical design is stronger than average for ECG FM transfer papers; I would not reject on novelty grounds. Private BrSwiss limits full external re-execution, but HUCA is public and the model list is public, so the core pattern is partially checkable. Scope-of-claim language (title/abstract vs single disease) and the sampling-rate confound on cross-site transfer are the only issues I would insist on before acceptance; both are fixable without new multi-disease cohorts."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The punchline is simple and useful. On Brugada detection, large-scale ECG foundation-model pre-training mostly stabilizes training for models that would not converge from scratch; it does not reliably hand you semantic, transferable clinical features. When the architecture already works supervised (ECG-CPC, MERL-ResNet, ECGFounder), fine-tuning the public FM only marginally beats the same net trained from random weights, and the best-pipeline gap on full BrSwiss is not significant (0.962 vs 0.932, p=0.091). Data-efficiency gains on BrSwiss-3% do not replicate on HUCA. Zero-shot cross-site transfer collapses toward chance for everyone.\n\nWhat is actually new is the design, not a single flashy number. Nine public FMs treated as deployable packages, matched from-scratch baselines of identical architecture, linear probe vs full fine-tune, a controlled window ablation, and bidirectional zero-shot transfer on an independent public cohort. Prior rare-disease FM papers mostly showed one custom model beating one CNN. This is the first systematic head-to-head that isolates pre-training from capacity for a rare phenotype. Patient-level AUC, bootstrap CIs, DeLong with FDR, early stopping to equalize exposure, and explicit within- vs between-model comparisons are done carefully. The conclusion is scoped tightly to what they measured.\n\nSoft spots are real but proportionate and mostly flagged by the authors. BrSwiss is private, so full reproduction needs that cohort. Only one rare disease; sampling-rate mismatch (1000 Hz vs 100 Hz) almost certainly hurts absolute transfer numbers; secondary metrics sit at a fixed 0.5 threshold that is not clinical. The claim that Brugada is absent from every pre-training corpus is asserted rather than exhaustively audited. None of that invents an alternative story in which the observed gains become semantic rather than mechanical on the data they actually ran.\n\nThis is for people building or buying ECG FMs, rare-disease digital cardiology, and anyone who still assumes “more pre-training data ⇒ better rare-disease transfer.” Math and stats look solid; citations cover the FM landscape and the thin rare-disease prior work without padding. I would send it to peer review and bring it to a medical-AI reading group. Worth citing when you need a concrete counterweight to optimistic FM transfer claims in low-prevalence ECG settings.","headline":"Clean negative result: for Brugada, public ECG FMs mostly buy optimization stability, not transferable clinical knowledge—architecture and site alignment dominate.","tokens_in":37792,"tokens_out":586,"would_cite":true,"duration_ms":9078,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"ECG foundation models do not transfer clinically meaningful knowledge for Brugada syndrome; pre-training mainly stabilizes large models that fail from scratch.","keywords":["Electrocardiogram","Foundation Models","Self-Supervised Learning","Rare Cardiac Diseases","Brugada Syndrome","Transfer Learning","Cross-site Generalization"],"falsifier":"On a third independent Brugada (or comparable rare ECG) cohort of similar size, show that the best fine-tuned foundation model significantly and reproducibly beats the best architecture-matched from-scratch baseline both under matched training-set size and under zero-shot cross-site transfer.","tokens_in":37853,"feed_emoji":"🫀","tokens_out":868,"duration_ms":16282,"temperature":0.7,"pith_summary":"This paper tests whether large ECG foundation models, pre-trained on vast unlabeled heart recordings, actually learn transferable clinical knowledge for rare diseases—or whether their benefits are mostly easier training. Using Brugada syndrome as a hard rare-disease case, the authors evaluate nine public ECG foundation models against identical architectures trained from scratch on two independent clinical cohorts, under linear probing, full fine-tuning, severe data reduction, and zero-shot hospital-to-hospital transfer. Pre-training rescues high-capacity models that otherwise collapse, but compact models that already work from scratch gain little; the best fine-tuned foundation model only marginally exceeds the strongest supervised baseline. Data-efficiency wins on one reduced cohort do not hold on the second site, and every strategy collapses toward chance under cross-site transfer. The finding challenges the assumption that scale alone encodes rare-phenotype knowledge and elevates architecture and domain alignment over pre-training status.","feed_headline":"ECG foundation models fail to transfer rare-disease knowledge","feed_subtitle":"Pre-training mainly stabilizes big models; architecture and site matter more for Brugada detection.","key_machinery":"Architecture-matched from-scratch baselines for each of nine public ECG foundation models, compared under linear probing and full fine-tuning, plus a 3% window-per-patient data ablation and zero-shot cross-site transfer between the BrSwiss and HUCA cohorts, which isolates the contribution of pre-trained weights from architecture, data volume, and site shift.","core_discovery":"For Brugada syndrome detection, foundation-model pre-training is mechanical rather than semantic: it stabilizes optimization for high-capacity architectures that cannot converge from scratch, but does not deliver statistically reliable gains over the best architecture-matched supervised baselines, does not produce a data-efficiency advantage that replicates across independent cohorts, and does not improve zero-shot cross-site robustness. Architectural suitability, not pre-training, is the most consistent predictor of performance.","pith_inferences":["Similar mechanical-versus-semantic patterns may appear for other low-prevalence ECG phenotypes absent from public pre-training corpora.","Acquisition-protocol and sampling-rate mismatch may erase any representation advantage faster than pre-training can create one.","Rare-disease ECG benchmarks may need mandatory architecture-matched from-scratch controls before foundation-model benefit is claimed.","If Brugada-like patterns were deliberately added to pre-training corpora, a reopened between-model gap would test whether semantic coverage was the missing ingredient."],"forward_implications":["Practitioners cannot assume off-the-shelf ECG foundation models will help rare cardiac phenotypes; architecture choice may matter more than pre-training.","Label-efficiency claims for ECG foundation models need multi-site replication before they are trusted in rare-disease settings.","Reliable cross-site rare-disease ECG performance will require explicit domain adaptation, not reliance on large-scale pre-training alone.","Compact models that train well from scratch can nearly match fine-tuned foundation models when limited labeled rare-disease data exist.","Pre-training strategies aimed at rare disease should deliberately cover clinical variability and multi-site domain shifts rather than scale alone."],"fun_headline_variants":["ECG FMs give optimization help not rare-disease knowledge transfer","Pre-training stabilizes big ECG models but skips Brugada semantics","Architecture outranks pre-training for Brugada ECG detection","FM pre-training fails to yield replicable rare-phenotype gains","Zero-shot and data-efficiency edges absent for ECG FMs on Brugada"],"cache_read_input_tokens":32896,"weakest_assumption_plain":"The claim depends on Brugada syndrome in these two adult case-control cohorts being a fair rare-disease stress test, and on those cases not having been present in any of the models’ pre-training data.","fun_headline_variants_meta":{"raw":{"variants":["ECG FMs give optimization help not rare-disease knowledge transfer","Pre-training stabilizes big ECG models but skips Brugada semantics","Architecture outranks pre-training for Brugada ECG detection","FM pre-training fails to yield replicable rare-phenotype gains","Zero-shot and data-efficiency edges absent for ECG FMs on Brugada"]},"model":"grok-4.5","effort":"low","cost_usd":0.004968,"raw_usage":{"total_tokens":1502,"prompt_tokens":915,"num_sources_used":0,"completion_tokens":93,"cost_in_usd_ticks":49680000,"prompt_tokens_details":{"text_tokens":915,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":494,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":915,"tokens_out":93,"duration_ms":5117,"temperature":1.0,"reasoning_tokens":494,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-12T05:28:27.699141+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"On a third independent Brugada (or comparable rare ECG) cohort of similar size, show that the best fine-tuned foundation model significantly and reproducibly beats the best architecture-matched from-scratch baseline both under matched training-set size and under zero-shot cross-site transfer.","supporting_citations":[],"review_version":1}