{"id":"cb822358-79e5-44e9-82e3-e68520965522","arxiv_id":"2508.00963","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":4.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"For ECG classification, fusing time and time-frequency features helps, but adding frequency features is redundant, so more modalities are not always better.","lead":"This paper tests whether merging more data representations (raw time, time-frequency images, frequency sequences) improves ECG classification. It reports that a two-way fusion of time and time-frequency signals outperforms single-domain models, but adding a third frequency domain gives no further benefit.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The manuscript body is an unrelated paper (material-behavior discovery), so the central ECG multimodal claim has no supporting methods or results in the submitted text.","rationale":"The reader's weakest_assumption (encoder capacity/tuning) is plausible but secondary; the reader's rationale already notes the content mismatch and marks UNVERDICTED. I partially agree: the same overall verdict is justified, but the load-bearing concern is not the comparability of encoders—it is the fact that the body of the manuscript contains zero ECG experiments. This concern cannot be resolved by inspecting architecture choices; it requires recovering the correct manuscript. Under the stress-test rule that all manuscript text is in-scope, the unrelated full text is decisive. The paper also offers no code, no formal verification, and no parameter-free derivation, so there is no independent support to lean on. I recommend the verdict remain UNCHANGED (UNVERDICTED) rather than moving to REJECT: the mismatch makes the submission unassessable, and a desk-reject determination would need editorial confirmation of the arXiv record.","tokens_in":11765,"tokens_out":3218,"duration_ms":38178,"concrete_test":"Retrieve the actual PDF/source of arXiv:2508.00963 and grep for tokens: 'ECG', 'electrocardiogram', 'Hybrid 1', 'Hybrid 2', '2D-CNN', '1D-CNN', 'Transformer', 'bootstrap', 'Bayesian'. Also compare the first-page abstract with the arXiv metadata. If none of the ECG/fusion terms appear in the body, the manuscript cannot support the abstract's claim; no amount of re-analysis of the model comparison can proceed until the correct full text is provided. If they do appear, then this concern dissolves and the reader's original unverifiability should be revisited with the actual methods.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that Hybrid 1 beats the 2D-CNN baseline on ECG while Hybrid 2 adds nothing, proving complementary domains matter more than modality count. The full text supplied for arXiv:2508.00963 is titled 'Enhancing material behavior discovery using embedding-oriented physically-guided neural networks with internal variables' and concerns nonlinear diffusion, POD/autoencoder decoders, and transfer learning. It contains no ECG dataset, no 1D-CNN/2D-CNN/Transformer encoders, no Hybrid 1 or Hybrid 2 fusion, and no bootstrap or Bayesian inference results. Thus the p-values and Bayesian probabilities cited in the abstract cannot be traced to any experiment in the body. The load-bearing premise of the paper—that a controlled ECG comparison was run and analyzed—is entirely missing. This is not a subtle modeling assumption like encoder capacity; it is the absence of the experiment itself. Consequently, the central claim is unsupported by the manuscript as submitted.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript, as submitted, consists of an abstract claiming a multimodal deep learning study for ECG classification and a full text that is an unrelated paper on embedding-oriented Physically-Guided Neural Networks with Internal Variables for material behavior discovery. The abstract describes three unimodal encoders (1D-CNN, 2D-CNN, 1D-CNN-Transformer), two hybrid fusions, bootstrap and Bayesian analyses, and a proposed \"Complementary Feature Domains\" framework. The full text contains no ECG dataset, no 1D-CNN/2D-CNN/Transformer encoders, no Hybrid 1 or Hybrid 2 architectures, and no bootstrap or Bayesian inference results. The claimed framework is not defined anywhere in the submitted text. Consequently, the central claim that Hybrid 1 outperforms the 2D-CNN baseline due to domain complementarity while Hybrid 2 adds redundancy cannot be checked against any evidence in the manuscript.","tokens_in":11948,"tokens_out":3908,"duration_ms":46346,"significance":"If the ECG experiments described in the abstract were fully reported, the empirical finding that fusing time and time-frequency domains helps while adding a redundant frequency domain does not help could be a useful, if modest, contribution to multimodal biomedical signal classification. The proposed theoretical framework, however, has no formal statement in the submitted text, so its significance cannot currently be assessed. The full text is a separate, more developed SciML paper with its own open-source repository, but that content is not the paper claimed by the abstract and does not bear on the ECG claim. As submitted, the manuscript is not a coherent article and the central contribution is unverified.","major_comments":[{"comment":"The full text (Sections 1–6) is titled \"Enhancing material behavior discovery using embedding-oriented Physically-Guided Neural Networks with Internal Variables\" and addresses a nonlinear diffusion problem with spectral, POD, and autoencoder decoders. It contains no ECG classification experiment, no 1D-CNN, 2D-CNN, or Transformer encoders, no Hybrid 1 or Hybrid 2 architectures, and no bootstrap or Bayesian results. The p-values and Bayesian probabilities cited in the abstract therefore have no supporting experiment in the manuscript. This is a load-bearing absence: the central claim that Hybrid 1 outperforms the 2D-CNN baseline is unsupported by the submitted text and cannot be fixed by local edits.","section":"Abstract vs. full text"},{"comment":"The abstract asserts a \"mathematically quantifiable framework\" named \"Complementary Feature Domains in Multimodal ECG Deep Learning\" and invokes \"intrinsic information-theoretic complementarity,\" but the submitted text provides no definition, equations, or formal criterion for complementarity. As written, the framework risks being circular: if complementarity is measured by the same performance differences it is invoked to explain, then the conclusion is definitional rather than explanatory. The authors need to state an independent measure of complementarity—for example, an information-theoretic quantity computed from the learned representations—and show how Hybrid 1's advantage follows from that measure.","section":"Abstract (framework definition)"},{"comment":"Even taken solely on its own terms, the abstract reports p-values and Bayesian probabilities without effect sizes, confidence intervals, dataset description, train/validation split details, model capacity matching, or multiple-testing corrections. Because the full text provides none of these details, the reported statistical evidence cannot be independently checked. A complete experimental section with architecture specifications, hyperparameters, capacity-matched baselines, and full result tables is required before the claimed findings can be evaluated.","section":"Abstract (statistical evidence)"}],"minor_comments":[{"comment":"The abstract and full text have different titles and clearly describe different research areas; the manuscript must be made internally consistent or resubmitted with the correct body.","section":"Title and authorship consistency"},{"comment":"Phrases such as \"paradigm-shifting\" and \"rigorously evaluated\" in the abstract are not supported by the submitted content and should be replaced with specific, measurable claims.","section":"Terminology"},{"comment":"The full text points to a GitHub repository for the PGNNIV study, but no repository, dataset, or code is provided for the ECG experiments claimed in the abstract; such artifacts are necessary for reproducibility.","section":"Reproducibility"}],"recommendation":"reject","confidential_remarks":"The mismatch between the abstract and the full text is complete rather than a subtle technical flaw. The editor may wish to verify whether the correct manuscript was uploaded; in its current state, the paper cannot be reviewed as a scientific contribution because the claimed experiments and framework are absent."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things you should know up front. The abstract promises a controlled ECG study with five models and bootstrap/Bayesian evidence. The body attached to this arXiv number is a different paper entirely—a PGNNIV study on material behavior discovery. No ECG data, no 1D/2D CNN or Transformer encoders, no Hybrid 1 or Hybrid 2, no statistical results. The stress-test note is right, and it is not a subtle modeling concern; it is the absence of the experiment that the abstract claims to report.\n\nThat said, the abstract itself is not silly. The idea that adding a modality can hurt if it is redundant, and that complementarity matters more than count, is a known theme in multimodal learning. The proposed comparison—time, time-frequency, and frequency encoders, plus two fusions—is a reasonable way to test it. The statistical machinery (bootstrap p-values, Bayesian probabilities) signals an attempt at rigor, though none of it is visible.\n\nThe soft spots are not in the abstract's rhetoric; they are in the manuscript. First, the content mismatch is fatal for review: you cannot verify a single number. Second, the 'Complementary Feature Domains' framework is named as a 'mathematically quantifiable framework' but no definition is given even in the abstract, so the circularity risk is real—complementarity may end up being whatever the performance differences say it is. Third, the abstract's 'paradigm-shifting' and 'redefines a fundamental principle' language is out of proportion to any evidence shown. The reader's concern about encoder capacity is secondary; you cannot even assess capacity without architecture details.\n\nWho is this for? If the actual ECG paper exists and is as described, it would interest people working on multimodal biomedical signal classification. But this submission, as it stands, should be bounced. The authors should be asked to resubmit with the correct full text, and with the framework defined and the experiments reproducible. I would not send this to referees; desk reject and invite a corrected resubmission.","headline":"The abstract is a plausible ECG multimodal paper, but the body is an unrelated material-property paper, so there is no experiment to review.","tokens_in":12405,"tokens_out":2039,"would_cite":false,"duration_ms":23812,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that for ECG classification, fusing time-domain and time-frequency-domain models beats any single-domain baseline, while adding the frequency domain as a third modality produces no further gain.","keywords":["multimodal deep learning","ECG classification","complementary feature domains","time-frequency analysis","representational redundancy","hybrid neural networks","Bayesian inference","ablation study"],"falsifier":"Train Hybrid 2 with a frequency encoder of strictly greater capacity, or with matched parameter count and optimization schedule, and check whether the gap with Hybrid 1 closes; if it does, the claimed redundancy is an artifact of the encoder rather than the domain. A complementary check is to compute the mutual information between the learned features of the time-frequency and frequency encoders: high mutual information with no accuracy gain would support redundancy, while low mutual information with no gain would undermine it.","tokens_in":11581,"feed_emoji":"🫀","tokens_out":7201,"duration_ms":76738,"temperature":0.7,"pith_summary":"This paper sets out to show that adding more input domains to a multimodal deep learning model does not automatically improve ECG classification; what matters is whether the domains carry complementary information. The authors build five models—three single-domain encoders and two fusions—and compare them on ECG classification with bootstrapping and Bayesian inference. They find that the two-domain hybrid fusing raw time with time-frequency representations consistently outperforms the best single-domain baseline, while adding a third frequency-domain stream gives no further gain and sometimes a small loss. The intended upshot is a design principle: optimal multimodal performance comes from information-theoretic complementarity between fused domains, not from the number of modalities.","feed_headline":"ECG fusion: two complementary domains beat three redundant ones","feed_subtitle":"Hybrid 1 beats the 2D-CNN baseline on every metric; the extra frequency stream is redundant.","key_machinery":"The argument runs through five architectures: a 1D-CNN on raw time signals, a 2D-CNN on time-frequency representations, a 1D-CNN-Transformer (an attention-based sequence model) on frequency spectra, Hybrid 1 (1D-CNN + 2D-CNN), and Hybrid 2 (1D-CNN + 2D-CNN + Transformer). The load-bearing device is the comparison between Hybrid 1 and Hybrid 2, evaluated with bootstrapping and Bayesian inference so that the difference is assessed as a distribution rather than a point estimate. The named framework, 'Complementary Feature Domains in Multimodal ECG Deep Learning,' is the mathematical account the paper offers for why the time and time-frequency domains are synergistic while the frequency domain is redundant. The ablation study is what connects the performance gap to representational redundancy rather than to the extra parameters of Hybrid 2.","core_discovery":"The central claim is that complementarity, not modality count, determines the value of multimodal fusion for biomedical signal classification. On ECG data, the time-domain 1D-CNN and the time-frequency 2D-CNN are complementary: their fusion (Hybrid 1) beats the 2D-CNN baseline on every reported metric, with p-values below 0.05 and Bayesian probabilities above 0.90. The frequency-domain 1D-CNN-Transformer does not add complementary information; when it is appended to form Hybrid 2, performance does not improve and can slightly decline. The paper attributes this to representational redundancy between the frequency and time-frequency domains, and reports a targeted ablation study supporting that explanation. It generalizes the finding into a proposed framework, 'Complementary Feature Domains in Multimodal ECG Deep Learning,' intended to quantify which domain combinations are ideal.","pith_inferences":["A direct test of the paper's reasoning would be to hold one encoder fixed and vary the second domain, e.g., replace the frequency spectrum with heart-rate variability features, to see whether complementarity rather than the specific domain drives the gain.","The complementarity criterion could be measured directly: compute mutual information between latent representations of candidate encoders before training and test whether that score predicts which fusion performs best.","The same 'redundant third domain' pattern likely appears in other physiological signals where time, spectral, and time-frequency views are routinely fused, such as EEG or PPG.","If the framework is right, the practical cost of multimodal systems can be cut by pruning redundant streams before training, without sacrificing accuracy."],"forward_implications":["Multimodal ECG models should be built by pairing domains with complementary information, not by stacking every available representation.","A three-domain hybrid with redundant domains can underperform a two-domain hybrid despite seeing strictly more input data.","Bootstrapping and Bayesian inference provide a usable protocol for deciding whether an added modality earns its place in a biomedical classifier.","The 'Complementary Feature Domains' principle gives a quantitative way to rank candidate domain combinations before training.","Future multimodal fusion studies should report whether each added domain improves accuracy beyond the best single domain, rather than only comparing fused models against weaker baselines."],"supporting_citations":[],"fun_headline_variants":["ECG fusion: complementarity, not count, drives gains","Two complementary domains beat three in ECG model","Extra frequency stream adds no value to ECG fusion","Multimodal ECG: fusion benefits from domain complementarity","Redundant modalities: why two beat three in ECG"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The comparison assumes the three single-domain encoders are equally well tuned and of comparable learning capacity, so the frequency encoder's lack of benefit reflects redundancy in the data rather than a weaker model.","fun_headline_variants_meta":{"raw":{"variants":["ECG fusion: complementarity, not count, drives gains","Two complementary domains beat three in ECG model","Extra frequency stream adds no value to ECG fusion","Multimodal ECG: fusion benefits from domain complementarity","Redundant modalities: why two beat three in ECG"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00032,"raw_usage":{"total_tokens":1851,"prompt_tokens":1040,"completion_tokens":811,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":656,"completion_tokens_details":{"reasoning_tokens":735}},"tokens_in":656,"tokens_out":811,"duration_ms":10820,"temperature":1.0,"reasoning_tokens":735,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T06:00:45.946484+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train Hybrid 2 with a frequency encoder of strictly greater capacity, or with matched parameter count and optimization schedule, and check whether the gap with Hybrid 1 closes; if it does, the claimed redundancy is an artifact of the encoder rather than the domain. A complementary check is to compute the mutual information between the learned features of the time-frequency and frequency encoders: high mutual information with no accuracy gain would support redundancy, while low mutual information with no gain would undermine it.","supporting_citations":[],"review_version":1}