{"id":"84080494-ba46-4832-b38b-abe7ea5db2b1","arxiv_id":"2608.09053","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":1,"one_line_summary":"LuminaECG, a 2B ECG agent trained with clinically structured supervision, beat larger zero-shot models and junior human readers on ECG interpretation, with reports that retain prognostic signal on an unseen cohort.","lead":"A 2B vision-language model trained on grid-rendered 12-lead ECGs with color-coded P, QRS, and T waves and measurement-grounded reports outperformed much larger zero-shot models on ECG measurement, diagnosis, and cross-country transfer. The result, if confirmed, suggests that how ECG training data are structured can matter more than model size.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central mechanism claim is not secured: the paper never states that rendered ECG images contain no printed text/device metadata, the promised shortcut ablation is absent, and Appendix B.2 shows non-visual template content in outputs.","rationale":"The paper's strongest evidence—CODE-test macro-F1 0.840 and 0.64–0.84 external transfer with zero training contact—does not by itself establish the mechanism. The claim that the model reads from the wave grid and color overlays depends on an unstated assumption about image content. If printed text or device measurements appear in the rendered images, the model could be solving an OCR/template task, and the human-tier comparison would not support the doctor-grounding narrative. Section 3 explicitly promises a test of whether the model reads 'from these priors or from correlated shortcuts,' but no such test is reported; the only ablation (pretrain-2B) removes all supervision and cannot separate image-content shortcuts from the visible-prior ladder. Appendix B.2's '0.005–150 Hz diagnostic bandwidth' and '60 Hz notch' outputs are non-visual, template-like details, consistent with either memorized gpt-4o-generated report structure or the presence of such text in the images. Neither possibility alone disproves the mechanism, but both make the central attribution insecure. The concrete test is straightforward, and the authors can meet the condition by disclosing image content and running the promised ablation; therefore the conditional verdict stands rather than being strengthened or reversed.","tokens_in":18317,"tokens_out":6571,"duration_ms":63701,"concrete_test":"Perform the shortcut ablation promised in Section 3 on a held-out set from the same pipeline: for the same records, render (1) the current full image, (2) the grid with the color overlay removed, (3) the waveform blanked but any printed header/measurement text retained, and (4) the waveform retained with all text cropped; include a text-only condition that supplies the 303-field numeric panel and device measurements with no image. Measure CODE-test macro-F1 and HR/PR/QRS MAE in each arm. If blanking the waveform leaves performance largely unchanged, or text-only matches the image condition, the central 'measurement-grounded visual reading' claim fails; if performance collapses without the waveform, the concern is resolved. If the production images already have no text, the authors should state this explicitly and condition (3) becomes vacuous.","verdict_should_be":"UNCHANGED","load_bearing_attack":"LuminaECG's central claim is that doctor-grounded visual supervision—the calibrated grid plus color-coded P/QRS/T overlays—is what enables a 2B model to surpass junior-reader tiers and transfer across continents. For that attribution to hold, the rendered image must contain only the waveform and overlay. If it also contains a printed header, numeric measurements, or the device's interpretation, the model could be reading text, and the strong benchmark numbers would not show measurement-grounded visual reading at all. The paper never states that printed text is removed. Section 3 promises to 'test below whether the model reads from these priors or from correlated shortcuts,' but no such ablation appears in the main text or appendices; Table A1's pretrain-2B row is a no-SFT baseline, not a test of grid/overlay/text content. Appendix B.2 adds direct evidence of non-visual template content: LuminaECG reports a '0.005 to 150 Hz diagnostic bandwidth' and a '60 Hz notch' from static grid images. These parameters cannot be read off the waveform, and they are characteristic of device metadata. This is not proof of text reading, since gpt-4o-generated training reports also contain a signal-quality section, but it shows the model reproduces clinically plausible non-visual content and makes the absence of a shortcut ablation decisive. Until the image content is disclosed and the promised ablation runs, the headline mechanism claim is an unverified interpretation of the numbers.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"LuminaECG is a 2B vision-language model fine-tuned on 12-lead ECG images rendered on grid paper with color-coded P/QRS/T segment overlays, paired with structured eight-section reports generated by gpt-4o from a numeric measurement panel. The paper evaluates the model on held-out MIMIC-IV-ECG report generation and numeric measurement, critical-finding recall on PTB-XL, diagnostic performance on CODE-test against human reader tiers, cross-cohort transfer to Brazil, China, and Germany, and emergent prognostic value in the Sami-Trop cohort. The main claims are that doctor-grounded visual supervision, rather than model scale, yields measurement-grounded ECG interpretation that surpasses junior-reader tiers and transfers across populations.","tokens_in":18496,"tokens_out":8036,"duration_ms":71060,"significance":"If the central attribution holds, this is an important contribution to medical multimodal LLM design: a compact 2B model trained with explicit grid, color-coded wave segments, and structured measurement-grounded reports outperforms much larger zero-shot systems on ECG measurement and diagnosis, and transfers across continents without retraining. The use of external benchmarks (CODE-test, PTB-XL, Chapman, Sami-Trop) with no training contact is a strong point, as is the out-of-fold prognostic analysis and the public release of model weights and a project page. The paper also provides a rare comparison against human reader tiers and critical-finding recall, going beyond aggregate F1. However, the specific attribution of the gains to the visual priors is not yet isolated from potential textual shortcuts, and the internal MIMIC evaluation is partly circular with respect to the training supervision. These concerns are fixable with additional disclosure and ablations.","major_comments":[{"comment":"The central attribution claim is not yet supported because the rendered ECG image content is not disclosed and the promised shortcut ablation is absent. Section 5.1 describes the grid, 6×2 layout, and color overlay, but never states whether printed headers, numeric measurements, or device-interpretation text are removed from the rendered images. If text is present, the model could read the diagnosis and measurements rather than measure the waveform, and the benchmark gains would not demonstrate measurement-grounded visual reading. The paper explicitly promises to test 'whether the model reads from these priors or from correlated shortcuts' (end of §2), yet no such experiment appears; the only ablation-style row, pretrain-2B in Table A1, is a no-SFT baseline and does not vary the image content. Appendix B.2 adds evidence of non-visual template content: the model outputs '0.005–150 Hz diagnostic bandwidth with 60 Hz notch' and '0.0005–150 Hz diagnostic bandwidth' from static grid images, values that cannot be inferred from the waveform. This makes the missing content disclosure and ablation decisive for the mechanism claim. Please disclose the exact rendering content (including whether any text or metadata is present) and add ablations that remove the grid, remove the color overlay, and add or remove printed text to isolate which visual components drive the gains.","section":"§5.1 (Image rendering) and §2"},{"comment":"The report-generation and numeric-measurement evaluations on MIMIC-IV-ECG are partly circular with respect to the training supervision. The reference reports appear to be the same gpt-4o-generated reports produced from the numeric panel used to construct training examples (§5.1), and the numeric references are derived from the same delineator pipeline. High BLEU/ROUGE/CIDEr and low MAE therefore measure agreement with the model's own training distribution rather than independent clinical quality. Please state explicitly whether the evaluation references are the same as those used in training, and provide an independent validation, such as a small sample of reports scored by cardiologists or measurements compared with manually annotated intervals, to support the claim of 'measurement-grounded ECG reporting' (Finding 1).","section":"§3.1, §5.1 (MIMIC-IV-ECG evaluation)"},{"comment":"The prognostic analysis does not pre-specify the report-derived feature set. With 104 deaths in Sami-Trop, selecting features such as the major-arrhythmia flag and prolonged QTc from the full report-derived panel can inflate the reported ΔC and hazard ratios. The paper should report the full candidate feature list, the selection procedure, and whether the bootstrap confidence interval accounts for selection. The out-of-fold design is a strength, but outcome-based feature selection would still bias the estimate; a pre-registered feature set or a sensitivity analysis entering all features simultaneously would strengthen Finding 5.","section":"§3.6 (Emergent prognostic value)"}],"minor_comments":[{"comment":"The sentence about ECG-R1 is internally inconsistent: recovering 1.1% of critical findings implies missing 98.9%, but the text says 'missing 81%'. Please clarify what the 81% refers to.","section":"§3.3"},{"comment":"The row label 'pretrain-2B (ablation, no SFT)' is misleading because this row is a no-fine-tuning baseline, not an ablation of the grid, overlay, or textual priors. Renaming it would avoid confusion with the promised shortcut ablation.","section":"Table A1"},{"comment":"'0.0005–150 Hz diagnostic bandwidth' in the fourth case is likely a typo for '0.05–150 Hz'; please correct it.","section":"Appendix B.2"},{"comment":"The term 'doctor-grounded' overstates the supervision: the training reports are written by gpt-4o from a numeric panel rather than by cardiologists. Please clarify the extent of direct clinician involvement in constructing the supervision data and consider softening the terminology accordingly.","section":"§5.1"},{"comment":"The comparison with human reader tiers relies on published benchmark values from the original CODE study. Please state explicitly that the reader protocol in that study (annotation set, number of readers, adjudication) matches the protocol used here, or provide a precise citation to the reader study.","section":"§3.4 and Table A2"}],"recommendation":"major_revision","confidential_remarks":"This is a strong empirical study with a clear central hypothesis, and the external benchmarks are a genuine strength. The main risks are the missing disclosure of rendered-image content and the absent shortcut ablation; if the images contain printed text or device metadata, the headline mechanism claim would be substantially weakened. The MIMIC-IV-ECG report-generation and measurement evaluations should also be framed as internal-distribution comparisons. I encourage the authors to add the disclosure and ablations in revision; with those, the paper could be acceptable. The release of model weights and the project page is appreciated."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Honestly, the reader's conditional verdict is close to mine. This is the first ECG-MLLM pipeline I know that puts grid rendering, color-coded P/QRS/T overlays, and measurement-grounded supervised reports into a single training example, and the external results are worth taking seriously. A 2B model that beats junior-reader tiers on CODE-test, keeps 0.64–0.84 macro-F1 across Brazil, China, and Germany, and produces a report structure with an independent prognostic signal (delta-C +0.086 on a Chagas cohort it never saw) is an interesting result. The external labels are genuinely independent, so the diagnostic claims are not circular. The critical-finding recall gap (50.8% versus <10.3% for frontier zero-shot models) is stark and worth checking.\n\nThe soft spots are real. The paper's central mechanism claim is that the model reads the calibrated grid and the colored wave overlays. That is not actually demonstrated, because the paper never discloses what the rendered images contain. If they carry printed headers or device interpretation, the model could be reading text. Appendix B.2 makes this concrete: LuminaECG reports a '0.005 to 150 Hz diagnostic bandwidth' and a '60 Hz notch' from a static grid image. Those cannot be inferred from pixels; they look like template content from the training reports. The paper promises to 'test below whether the model reads from these priors or from correlated shortcuts,' and that ablation simply does not appear. Table A1's pretrain-2B row is a no-SFT baseline, not a content ablation.\n\nTwo smaller issues: the MIMIC report and measurement evaluation uses references generated by the same automated pipeline that built the supervision, so those numbers demonstrate pipeline consistency, not clinical fidelity. And the CODE-test comparison reports point estimates without confidence intervals, so the claim of finishing above the ED-resident tier is a point estimate with no uncertainty attached.\n\nNone of this sinks the paper. The external diagnostic results stand on independent labels, and the cross-cohort transfer is credible. But the headline mechanism—measurement-grounded visual reading—is overclaimed relative to the evidence. A serious referee should ask for the exact rendering disclosure, the promised grid/overlay/text ablation, and confidence intervals for the human-reader comparison. I would send this out for review; the method is novel and the required fixes are addressable.","headline":"A novel, credible ECG-MLLM training pipeline whose 'measurement-grounded visual reading' mechanism is not yet proven — send to serious peer review, but demand the rendering disclosure and the promised ablation.","tokens_in":19163,"tokens_out":4591,"would_cite":true,"duration_ms":39393,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A compact, doctor-grounded ECG model can read above junior human readers and transfer across continents, suggesting supervision design matters more than model scale.","keywords":["ECG interpretation","vision-language model","doctor-grounded supervision","measurement grounding","critical-finding recall","cross-cohort generalization","structured report generation","low-rank fine-tuning"],"falsifier":"Render a held-out set of ECGs at 50 mm/s instead of 25 mm/s with the same clinical content and check whether reported PR, QRS, and QT intervals rescale correctly; if they do not, the model is reciting values rather than measuring the grid. In addition, count how often outputs contain phrases that cannot be visually derived from a static grid image—such as the '0.005–150 Hz diagnostic bandwidth with 60 Hz notch' reported in Appendix B.2—to test whether template text leaks into allegedly measurement-grounded reports.","tokens_in":18009,"feed_emoji":"🫀","tokens_out":12030,"duration_ms":101247,"temperature":0.7,"pith_summary":"This paper tries to establish that the way ECG training data is built—not just the size of the model—decides whether an AI agent can interpret an electrocardiogram the way a clinician does. The authors build LUMINAECG, a two-billion-parameter vision-language agent fine-tuned on twelve-lead ECGs rendered on standard grid paper, with P, QRS, and T waves colour-delineated and each report tied to waveform-derived measurements. On the CODE-test human-reader benchmark it reaches a macro-F1 of 0.840—a balanced average across six red-flag diagnoses—above the medical-student tier (0.817) and emergency-resident tier (0.829). Across three unseen external cohorts it keeps macro-F1 between 0.64 and 0.84, while a conventional waveform classifier trained on the same data falls to 0.05–0.47. The wider claim is that supervision preserving how clinicians localize, measure, and connect evidence can substitute for model scale in medical AI.","feed_headline":"A 2B ECG agent reads above medical-student and ER-resident tiers","feed_subtitle":"Grid-rendered tracings and doctor-style reports lift a compact model to junior-reader level and across continents.","key_machinery":"The load-bearing mechanism is a visible-prior ladder built into every training example: a twelve-lead ECG rendered on standard electrocardiographic grid paper at 25 mm/s and 10 mm/mV so that time and voltage become measurable pixel distances; a colour overlay marking P (green), QRS (red), and T (blue) wave segments; and an eight-section structured report generated under a grounding protocol that permits only measurements and device-determined diagnoses. Every number asserted in a report originates from a 303-field automatic measurement panel derived from waveform delineation, so interval, rate, and axis claims are anchored to quantities a clinician could redraw from the tracing. Low-rank supervised fine-tuning of a general two-billion-parameter vision-language backbone is the only trainable component, which isolates the effect of data design from the effect of model scale.","core_discovery":"The central discovery is that cardiologist-grounded priors—visible waveform anatomy, a calibrated grid for measurement, and reports that attach every conclusion to a measured quantity—convert ECG interpretation from label imitation into measurement-grounded reading. A two-billion-parameter model, fine-tuned only through low-rank adaptation and without architectural changes, improves heart-rate measurement error to 0.43 bpm (versus 6.45 bpm for the strongest zero-shot frontier system), recovers 50.8% of critical findings on an unseen PTB-XL cohort while every compared frontier zero-shot system stayed below 10.3%, and reaches a junior-reader tier on CODE-test, exceeding the medical-student and emergency-resident tiers on five of six diagnoses. Without retraining it keeps macro-F1 between 0.644 and 0.840 across Brazilian, Chinese, and German cohorts, where a ResNet1d-101 classifier trained on the same source data collapses to 0.049–0.470. The paper also finds that the structure of its generated reports carries emergent prognostic value on a zero-contact Chagas cardiomyopathy cohort, improving an age-and-sex Cox model's concordance index by $\\Delta C=+0.086$, even though the model never saw mortality outcomes.","pith_inferences":["A direct ablation—removing the colour overlay, distorting the grid scale, or withholding device-derived metadata strings—would isolate how much of the gain comes from true visual measurement versus memorized report structure; the paper does not report such ablations.","The same supervision recipe is a natural candidate for other measurement-intensive visual diagnostics, such as echocardiography or retinal imaging, where calibrated scales and structured reporting conventions exist.","The emergent prognostic association suggests structured diagnostic transcription encodes physiology beyond label content; a testable extension would corrupt the numeric measurements in training reports and see whether the mortality signal degrades.","CODE-test junior-reader alignment is a benchmark result, not a deployment licence: prospective validation, local reader panels, and adjustment for comorbidities would be required before clinical use."],"forward_implications":["Medical multimodal agents for measurement-intensive diagnostics can reach clinically useful performance without frontier-scale parameter counts, redirecting effort from model size to supervision design.","Doctor-grounded supervision behaves as a transportable prior: an agent trained in one health system keeps diagnostic quality on unseen populations and devices, addressing a known failure mode of closed-set ECG classifiers.","Because generated reports carry independent prognostic signal, the report itself becomes a clinical artifact: readable, checkable, and risk-relevant, not just a post-hoc explanation of a label.","Quantitative grounding gives clinicians an audit trail—heart-rate, PR, and QRS errors of 0.43 bpm, 8.97 ms, and 4.48 ms—so the agent's numbers can be verified against the tracing the same way a colleague's reading would be.","Critical-finding recall of 50.8% versus below 10.3% for frontier zero-shot systems suggests that safety-relevant ECG misses are governed more by training-data grounding than by model size."],"supporting_citations":[{"why":"Defines the grid standardization (25 mm/s, 10 mm/mV) that gives rendered ECG images a fixed time-voltage scale for measurement.","marker":"Kligfield et al., 2007"},{"why":"Provides the wavelet-based delineation that locates P/QRS/T boundaries and feeds the automatic measurement panel underlying every grounded report.","marker":"Martínez et al., 2004"},{"why":"Supplies CODE-test, the human-reader benchmark with expertise tiers, and the Ribeiro-2020 DNN baseline to which LUMINAECG is compared.","marker":"Ribeiro et al., 2020"},{"why":"Supplies the two-billion-parameter vision-language backbone used to test whether supervision can substitute for scale.","marker":"Qwen Team, 2025"},{"why":"Provides low-rank adaptation, the only trained component in the fine-tuning setup.","marker":"Hu et al., 2022"},{"why":"Supplies PTB-XL, the unseen German cohort used for critical-finding recall and cross-cohort transfer.","marker":"Wagner et al., 2020"},{"why":"Provides ECG-R1, the ECG-specialist baseline whose 1.1% in-domain critical-finding recall anchors the claim that supervision design, not specialist pretraining, is the bottleneck.","marker":"Jin et al., 2026"}],"fun_headline_variants":["2B ECG agent beats zero-shot frontier models by grounding in waveform measurements","Compact ECG model reads like a cardiologist, beats zero-shot giants","ECG agent transfers across continents without retraining, keeps F1 0.64-0.84","Grid-rendered tracings lift 2B ECG agent to junior-reader level","Doctor-grounded priors let 2B ECG model hit junior-reader tier"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central claim depends on LUMINAECG's reports being produced by reading the waveform, grid, and colour-coded wave segments in the rendered image, rather than by reproducing memorized template language or device metadata that may have leaked into the supervision.","fun_headline_variants_meta":{"raw":{"variants":["2B ECG agent beats zero-shot frontier models by grounding in waveform measurements","Compact ECG model reads like a cardiologist, beats zero-shot giants","ECG agent transfers across continents without retraining, keeps F1 0.64-0.84","Grid-rendered tracings lift 2B ECG agent to junior-reader level","Doctor-grounded priors let 2B ECG model hit junior-reader tier"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00155,"raw_usage":{"total_tokens":6241,"prompt_tokens":1036,"completion_tokens":5205,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":652,"completion_tokens_details":{"reasoning_tokens":5097}},"tokens_in":652,"tokens_out":5205,"duration_ms":35571,"temperature":1.0,"reasoning_tokens":5097,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T04:16:32.816233+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Render a held-out set of ECGs at 50 mm/s instead of 25 mm/s with the same clinical content and check whether reported PR, QRS, and QT intervals rescale correctly; if they do not, the model is reciting values rather than measuring the grid. In addition, count how often outputs contain phrases that cannot be visually derived from a static grid image—such as the '0.005–150 Hz diagnostic bandwidth with 60 Hz notch' reported in Appendix B.2—to test whether template text leaks into allegedly measurement-grounded reports.","supporting_citations":[{"cited_title":"Qwen3-VL : A frontier vision-language family","cited_arxiv_id":null,"evidence_quote":"Supplies the two-billion-parameter vision-language backbone used to test whether supervision can substitute for scale."},{"cited_title":"PTB-XL , a large publicly available electrocardiography dataset","cited_arxiv_id":null,"evidence_quote":"Supplies PTB-XL, the unseen German cohort used for critical-finding recall and cross-cohort transfer."}],"review_version":1}