{"id":"71914dc9-355b-4918-906c-9946df4b361a","arxiv_id":"2607.09948","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"AI for cardiac amyloidosis is most mature for bone-scintigraphy detection and SPECT/CT tracer-burden quantification; subtype classification, prognosis, and treatment-response monitoring remain early-stage and need multimodal, longitudinally validated systems.","lead":"This narrative review organizes AI work in cardiac amyloidosis by clinical task and finds a clear maturity gradient: scintigraphy detection and SPECT/CT quantification are nearest translation, while subtype, prognosis, and treatment-response models lag. It matters because it shows why high AUC alone does not equal clinical readiness and maps the evidence still missing for real workflows.","discovery_kind":"review","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified","rationale":"The reader correctly identifies the only structural vulnerability of a narrative review—selection and qualitative ranking without formal RoB or meta-analysis—and correctly judges that it does not break the claim. The maturity gradient is not an opaque judgment call; it is grounded in concrete, publicly checkable differences in cohort scale, external validation, label independence, and output interpretability that the paper and its supplementary tables document. No internal inconsistency, circular reference standard, or unsupported leap from discrimination to clinical readiness appears in the strongest claim. Therefore no verdict adjustment is warranted: ACCEPT remains appropriate, and the disclosed narrative limitation is already priced into the reader’s confidence and novelty scores.","tokens_in":32972,"tokens_out":468,"duration_ms":5274,"concrete_test":"Spot-check the two flagship anchors against their primary publications: confirm Spielvogel 2024 external AUCs and multi-tracer design, and Miller 2024 CPA/VOI/TBR AUCs plus age-adjusted HRs for CV death/HF hospitalization. If either primary source fails to support the numbers or outcome linkage claimed in the review, the maturity ranking would need revision; otherwise the claim stands.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper’s central claim is a task-based maturity ranking of published CA AI work, not a new empirical result. The ranking rests on transparent, repeatedly cited anchors: Spielvogel et al. (2024) multi-center multi-tracer scintigraphy detection (n≈16k, external AUCs 0.925–1.0) and Miller et al. (2024/2026) AI-enabled SPECT/CT quantification with interpretable, outcome-linked TBR/VOI/CPA biomarkers. Subtype, prognosis, and treatment-response sections correctly flag small cohorts, retrospective enrichment, and incomplete external validation. The Methods section explicitly discloses the narrative (non-systematic) design and absence of formal risk-of-bias scoring. That disclosure, together with the supplementary task tables that list cohort size, validation level, and label type for each study, makes the reader’s weakest-assumption concern real but non-load-bearing: it does not reverse or materially distort the maturity gradient the authors assert.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"This narrative review synthesizes machine-learning and deep-learning applications in cardiac amyloidosis by clinical task (screening, detection, quantification, prognosis, treatment-response monitoring) rather than by input modality. It argues that binary detection of abnormal cardiac uptake on bone scintigraphy and AI-assisted SPECT/CT quantification of myocardial tracer burden are closest to clinical translation, supported by large externally validated cohorts and interpretable, outcome-linked biomarkers (TBR, VOI, CPA), whereas subtype-aware ATTR-versus-AL classification, prognostic stratification, and treatment-response monitoring remain early-stage because of small cohorts, enriched retrospective designs, heterogeneous labels, incomplete external validation, and uncertain calibration at realistic prevalence. High discrimination alone is judged insufficient; clinically useful AI must perform against relevant mimics, avoid circular reference standards, and improve decisions within established multimodal pathways. The authors therefore call for multimodal, subtype-aware, longitudinally validated systems that support rather than replace current diagnostic algorithms.","tokens_in":33162,"tokens_out":706,"duration_ms":6797,"significance":"If the maturity ranking holds, the review supplies a useful organizing principle for a rapidly expanding literature and a concrete translational roadmap. The task-based framing correctly explains why apparently similar models need different cohorts, labels, metrics, and thresholds, and the ranking is anchored in the strongest primary evidence (Spielvogel et al. 2024 multicenter multi-tracer scintigraphy detection; Miller et al. 2024/2026 AI-enabled SPECT/CT quantification with outcome-linked biomarkers). Explicit disclosure of the narrative design, absence of formal risk-of-bias scoring, and the detailed supplementary task tables that list cohort size, validation level, and label type for each study make the appraisal transparent and usable by both methodologists and clinicians. The work is therefore a timely synthesis that can guide dataset design, evaluation standards, and workflow integration priorities in CA AI.","major_comments":[],"minor_comments":[{"comment":"Methods: state the approximate search window (last date of literature search) so readers can judge currency of the synthesis.","section":null},{"comment":"Section 4.2.1 and Supplementary Table B: Mo et al. is correctly flagged as a preprint; ensure the main text consistently labels it as emerging/preprint evidence rather than peer-reviewed validation.","section":null},{"comment":"Figure 4 caption and Section 5.1: a brief note that latent deep-learning features can still be paired with post-hoc explanations (saliency, SHAP) would avoid implying that only nuclear biomarkers are clinically usable.","section":null},{"comment":"Supplementary tables: a few entries (e.g., TRACE-AI, Miller 2026 longitudinal status) use cautious language about multicenter status; align main-text citations with the same caution for consistency.","section":null},{"comment":"Minor typographical inconsistencies (spacing around hyphens in “task-based,” occasional double spaces) can be cleaned in production.","section":null}],"recommendation":"accept","confidential_remarks":"The corresponding author’s disclosed consulting relationship with Pressiant Health is appropriately stated and does not appear to drive the maturity ranking, which rests on independent multicenter primary studies. Fit for a medical-physics / cardiovascular-imaging journal is good; the narrative design is a limitation but is disclosed and does not undermine the central claim."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"This is a clean narrative review that reorganizes the CA AI literature by clinical task (screening, detection, quantification, prognosis, treatment response) instead of by modality. That framing is the real contribution. It makes clear why an ECG screener and a SPECT/CT quantifier need different cohorts, labels, metrics, and deployment thresholds, and it lands a maturity gradient that matches the evidence: large multi-center multi-tracer scintigraphy detection (Spielvogel ~16k patients, external AUCs 0.925–1.0) and AI-enabled SPECT/CT burden markers (Miller TBR/VOI/CPA, outcome-linked) sit closest to translation; subtype, prognosis, and response work remain small, retrospective, and under-validated.\n\nWhat it does well: the clinical background is accurate, the supplementary task tables list cohort size, validation level, and label type, and the authors repeatedly insist that high AUC is not enough without realistic prevalence, mimics, calibration, and workflow fit. The roadmap (multimodal, subtype-aware, longitudinal, staged vs fused) is practical rather than hand-wavy. Citation pattern is dense and appropriate; competing-interest disclosure is present and does not drive the argument.\n\nSoft spots are real but proportionate. It is explicitly non-systematic—no formal risk-of-bias scoring or pooled synthesis—so the maturity ranking is a transparent expert judgment, not a meta-analytic result. A few very recent or preprint items are included as “emerging,” which is fine if labeled. No new model, dataset, or experiment; impact is field-reorientation within CA/nuclear-cardiology AI, not a broader medical-AI shift.\n\nWho it is for: people building or reviewing CA AI, nuclear cardiology groups, and editors who want a map of what is ready versus premature. I would bring it to reading group as a shared map of the literature, cite it when I need the task-based maturity claim, and send it to peer review without hesitation. Accept as a critical synthesis with the narrative limitation already disclosed.","headline":"Solid task-based narrative review that correctly ranks scintigraphy detection and SPECT/CT quantification as nearest to translation; useful synthesis, not a new result.","tokens_in":33748,"tokens_out":513,"would_cite":true,"duration_ms":6542,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Binary detection and AI-assisted quantification on bone scintigraphy and SPECT/CT are closest to clinical translation for cardiac amyloidosis; subtype, prognosis, and treatment-response models remain early-stage, and high discrimination alo","keywords":["Cardiac amyloidosis","machine learning","multimodal imaging","nuclear cardiology","quantitative imaging biomarkers","SPECT/CT","ATTR","AL amyloidosis"],"falsifier":"A large prospective multicenter study that either (a) shows subtype-aware or treatment-response AI matching scintigraphy detection on external validation, calibration at realistic prevalence, and change in referral or outcomes, or (b) shows AI SPECT/CT quantification failing to improve management decisions or outcome linkage when deployed against established staging and biomarkers.","tokens_in":33874,"feed_emoji":"❤️","tokens_out":1078,"duration_ms":17406,"temperature":0.7,"pith_summary":"Cardiac amyloidosis is underdiagnosed because it looks like more common heart diseases, and correct care requires separating transthyretin from light-chain disease with multimodal evidence. This narrative review organizes AI studies by clinical task—screening, detection, quantification, prognosis, and treatment-response monitoring—rather than by imaging modality, showing that each task demands different cohorts, labels, metrics, and deployment thresholds. The evidence forms a clear maturity gradient: binary uptake detection and AI-enabled SPECT/CT tracer-burden measures sit nearest clinical use, backed by large externally validated cohorts and interpretable, outcome-linked biomarkers. Subtype-aware classification, prognostic stratification, and treatment-response monitoring stay early, limited by small cohorts, enriched retrospective designs, mixed labels, incomplete external validation, and uncertain calibration at real-world prevalence. The authors argue the field should stop adding single-modality classifiers and build multimodal, subtype-aware, longitudinally validated systems that support—not replace—established diagnostic pathways.","feed_headline":"Scintigraphy AI is nearest clinic for cardiac amyloidosis","feed_subtitle":"Detection and tracer-burden measures lead; subtype and response models lag, so high AUC is not enough.","key_machinery":"Task-based synthesis of the literature (screening, detection, quantification, prognosis, treatment-response monitoring). Grouping by clinical job rather than input modality shows why outwardly similar AI models need different cohorts, reference standards, evaluation metrics, and implementation thresholds, and why maturity differs by task.","core_discovery":"Across AI applications in cardiac amyloidosis, a maturity gradient exists: binary detection of abnormal cardiac uptake and AI-assisted quantification of myocardial tracer burden on bone scintigraphy and SPECT/CT are closest to clinical translation, while subtype-aware classification, prognostic risk models, and treatment-response monitoring remain early-stage. High discrimination scores alone do not establish clinical usefulness; models must hold against relevant mimics, across subgroups, with independent reference standards, and improve real referral or management decisions inside existing workflows.","pith_inferences":["If SPECT/CT quantification becomes a de facto monitoring standard, trial endpoints for ATTR disease-modifying drugs may shift from biomarkers and functional class toward volumetric tracer-burden change.","The same task-based maturity filter could be reused for other rare infiltrative or phenotypically overlapping cardiomyopathies where single-modality AUC papers currently outpace multimodal validation.","Equity gaps already visible in AI-ECG subgroup performance (for example ethnicity, LVH, conduction disease) will likely reappear in any deployed screening program unless post-deployment monitoring is required from day one.","Opportunistic CT and pre-TAVI AI case-finding may become the practical bridge that brings nuclear quantification into earlier care pathways before dedicated amyloid imaging is ordered."],"forward_implications":["Scintigraphy uptake-detection models and SPECT/CT burden biomarkers (TBR, volume of involvement, cardiac pyrophosphate activity) should be prioritized for prospective workflow trials and regulatory planning first.","Screening and detection claims must report PPV, calibration, and performance against realistic mimics (HFpEF, hypertensive LVH, HCM, aortic stenosis), not only AUC on case-control cohorts.","Subtype, prognosis, and monitoring models need larger subtype-confirmed, outcome-linked, multicenter datasets before they can guide therapy choice or trial endpoints.","Staged multimodal pipelines (for example ECG then echo) and direct multimodal fusion should be compared head-to-head for yield, false positives, and time to diagnosis.","Future systems must be designed to support existing non-biopsy ATTR algorithms and monoclonal-protein testing, not replace them."],"fun_headline_variants":["Scintigraphy AI leads cardiac amyloidosis translation pathway","Detection, tracer burden closest; subtype models still early","Maturity gradient: scintigraphy AI nearer clinic than rest","Binary uptake AI ready-er than subtype or response models","High AUC alone fails; scintigraphy detection leads CA AI"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The authors’ narrative, non-systematic selection of studies and qualitative ranking of which tasks are “closest to translation” versus “early-stage” accurately reflect the true state of the field.","fun_headline_variants_meta":{"raw":{"variants":["Scintigraphy AI leads cardiac amyloidosis translation pathway","Detection, tracer burden closest; subtype models still early","Maturity gradient: scintigraphy AI nearer clinic than rest","Binary uptake AI ready-er than subtype or response models","High AUC alone fails; scintigraphy detection leads CA AI"]},"model":"grok-4.5","effort":"low","cost_usd":0.003222,"raw_usage":{"total_tokens":1117,"prompt_tokens":822,"num_sources_used":0,"completion_tokens":62,"cost_in_usd_ticks":32220000,"prompt_tokens_details":{"text_tokens":822,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":233,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":822,"tokens_out":62,"duration_ms":2405,"temperature":1.0,"reasoning_tokens":233,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-14T14:22:18.515819+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"A large prospective multicenter study that either (a) shows subtype-aware or treatment-response AI matching scintigraphy detection on external validation, calibration at realistic prevalence, and change in referral or outcomes, or (b) shows AI SPECT/CT quantification failing to improve management decisions or outcome linkage when deployed against established staging and biomarkers.","supporting_citations":[],"review_version":1}